Ukrainian TTS (text-to-speech) using Coqui TTS

Last update: Dec 26, 2022

Overview

title	emoji	colorFrom	colorTo	sdk	app_file	pinned
Ukrainian TTS	🐸	green	green	gradio	app.py	false

Ukrainian TTS 📢 🤖

Ukrainian TTS (text-to-speech) using Coqui TTS.

Trained on M-AILABS Ukrainian dataset using sumska voice.

Link to online demo -> https://huggingface.co/spaces/robinhad/ukrainian-tts

Support

If you like my work, please support -> SUPPORT LINK

Example

test.mp4

How to use :

pip install -r requirements.txt.
Download model from "Releases" tab.
Launch as one-time command:

tts --text "Text for TTS" \
    --model_path path/to/model.pth.tar \
    --config_path path/to/config.json \
    --out_path folder/to/save/output.wav

or alternatively launch web server using:

tts-server --model_path path/to/model.pth.tar \
    --config_path path/to/config.json

How to train:

Refer to "Nervous beginner guide" in Coqui TTS docs.
Instead of provided config.json use one from this repo.

Attribution

Code for app.py taken from https://huggingface.co/spaces/julien-c/coqui

Comments

Error with file: speakers.pth

FileNotFoundError: [Errno 2] No such file or directory: '/home/user/Soft/Python/mamba1/TTS/vits_mykyta_latest-September-12-2022_12+38AM-829e2c24/speakers.pth'

opened by akirsoft 4

doc: fix examples in README

Problem

The one-time snippet does not work as is and complains that the speaker is not defined

 > initialization of speaker-embedding layers.
 > Text: Перевірка мікрофона
 > Text splitted to sentences.
['Перевірка мікрофона']
Traceback (most recent call last):
  File "/home/serg/.local/bin/tts", line 8, in <module>
    sys.exit(main())
  File "/home/serg/.local/lib/python3.8/site-packages/TTS/bin/synthesize.py", line 350, in main
    wav = synthesizer.tts(
  File "/home/serg/.local/lib/python3.8/site-packages/TTS/utils/synthesizer.py", line 228, in tts
    raise ValueError(
ValueError:  [!] Look like you use a multi-speaker model. You need to define either a `speaker_name` or a `speaker_wav` to use a multi-speaker model.

Also it speakers.pth should be downloaded.

Fix

Just a few documentation changes:

make instructions on what to download from Releases more precise
add --speaker_id argument with one of the speakers

opened by seriar 2

One vowel words in the end of the sentence aren't stressed

Input:


Бобер на березі з бобренятами бублики пік.

Боронила борона по боронованому полю.

Ішов Прокіп, кипів окріп, прийшов Прокіп - кипить окріп, як при Прокопі, так і при Прокопі і при Прокопенятах.

Сидить Прокоп — кипить окроп, Пішов Прокоп — кипить окроп. Як при Прокопові кипів окроп, Так і без Прокопа кипить окроп.

Result:


Боб+ер н+а березі з бобрен+ятами б+ублики пік.

Борон+ила борон+а п+о борон+ованому п+олю.

Іш+ов Пр+окіп, кип+ів окр+іп, прийш+ов Пр+окіп - кип+ить окр+іп, +як пр+и Пр+окопі, т+ак +і пр+и Пр+окопі +і пр+и Прокопенятах.

Сид+ить Прок+оп — кип+ить окроп, Піш+ов Прок+оп — кип+ить окроп. +Як пр+и Пр+окопові кип+ів окроп, Т+ак +і б+ез Пр+окопа кип+ить окроп.```

opened by robinhad 0

Error import StressOption

Traceback (most recent call last): File "/home/user/Soft/Python/mamba1/test.py", line 1, in from ukrainian_tts.tts import TTS, Voices, StressOption ImportError: cannot import name 'StressOption' from 'ukrainian_tts.tts'

opened by akirsoft 0

Vits improvements

vitsArgs = VitsArgs(
    # hifi V3
    resblock_type_decoder = '2',
    upsample_rates_decoder = [8,8,4],
    upsample_kernel_sizes_decoder = [16,16,8],
    upsample_initial_channel_decoder = 256,
    resblock_kernel_sizes_decoder = [3,5,7],
    resblock_dilation_sizes_decoder = [[1,2], [2,6], [3,12]],
)

opened by robinhad 0

Model improvement checklist
[x] Add Ukrainian accentor - https://github.com/egorsmkv/ukrainian-accentor

[ ] Fine-tune from existing checkpoint (e.g. VITS Ljspeech)

[ ] Try to increase fft_size, hop_length to match sample_rate accordingly

[ ] Include more dataset samples into model
opened by robinhad 0

Releases(v4.0.0)

v4.0.0(Dec 10, 2022)

This is a release of Ukrainian TTS model and checkpoint. License for this model is GNU GPL v3 License. This release has a stress support using + sign before vowels. Model was trained for 160 000 steps by @robinhad .
Source code(tar.gz)
Source code(zip)
config.yaml(7.88 KB)
model.pth(368.17 MB)
v3.0.0(Sep 14, 2022)
This is a release of Ukrainian TTS model and checkpoint. License for this model is GNU GPL v3 License. This release has a stress support using + sign before vowels. Model was trained for 280 000 steps by @robinhad . Kudos to @egorsmkv for providing dataset for this model. Kudos to @proger for providing alignment scripts. Kudos to @dchaplinsky for Dmytro voice.

Example:

Test sentence:

К+ам'ян+ець-Под+ільський - м+істо в Хмельн+ицькій +області Укра+їни, ц+ентр Кам'ян+ець-Под+ільської міськ+ої об'+єднаної територі+альної гром+ади +і Кам'ян+ець-Под+ільського рай+ону.

Mykyta (male):

https://user-images.githubusercontent.com/5759207/190852232-34956a1d-77a9-42b9-b96d-39d0091e3e34.mp4

Olena (female):

https://user-images.githubusercontent.com/5759207/190852238-366782c1-9472-45fc-8fea-31346242f927.mp4

Dmytro (male):

https://user-images.githubusercontent.com/5759207/190852251-db105567-52ba-47b5-8ec6-5053c3baac8c.mp4

Olha (female):

https://user-images.githubusercontent.com/5759207/190852259-c6746172-05c4-4918-8286-a459c654eef1.mp4

Lada (female):

https://user-images.githubusercontent.com/5759207/190852270-7aed2db9-dc08-4a9f-8775-07b745657ca1.mp4
Source code(tar.gz)
Source code(zip)
config.json(12.07 KB)
model-inference.pth(329.95 MB)
model.pth(989.97 MB)
speakers.pth(495 bytes)
v3.0.0-alpha(Sep 8, 2022)

Mykyta, Lada and Olena License: GPL v3 licence.
Source code(tar.gz)
Source code(zip)
config.json(9.94 KB)
model-inference.pth(329.95 MB)
model.pth(989.96 MB)
speakers.pth(431 bytes)
v2.0.0(Jul 10, 2022)
This is a release of Ukrainian TTS model and checkpoint using voice (7 hours) from Mykyta dataset. License for this model is GNU GPL v3 License. This release has a stress support using + sign before vowels. Model was trained for 140 000 steps by @robinhad . Kudos to @egorsmkv for providing Mykyta and Olena dataset.

Example:

Test sentence:

К+ам'ян+ець-Под+ільський - м+істо в Хмельн+ицькій +області Укра+їни, ц+ентр Кам'ян+ець-Под+ільської міськ+ої об'+єднаної територі+альної гром+ади +і Кам'ян+ець-Под+ільського рай+ону.

Mykyta (male):

https://user-images.githubusercontent.com/5759207/178158485-29a5d496-7eeb-4938-8ea7-c345bc9fed57.mp4

Olena (female):

https://user-images.githubusercontent.com/5759207/178158492-8504080e-2f13-43f1-83f0-489b1f9cd66b.mp4
Source code(tar.gz)
Source code(zip)
config.json(9.97 KB)
model-inference.pth(329.95 MB)
model.pth(989.72 MB)
optimized.pth(329.95 MB)
speakers.pth(431 bytes)
v2.0.0-beta(May 8, 2022)

This is a beta release of Ukrainian TTS model and checkpoint using voice (7 hours) from Mykyta dataset. License for this model is GNU GPL v3 License. This release has a stress support using + sign before vowels. Model was trained for 150 000 steps by @robinhad . Kudos to @egorsmkv for providing Mykyta dataset.

Example:

https://user-images.githubusercontent.com/5759207/167305810-2b023da7-0657-44ac-961f-5abf1aa6ea7d.mp4

:
Source code(tar.gz)
Source code(zip)
config.json(8.85 KB)
LICENSE(34.32 KB)
model-inference.pth(317.15 MB)
model.pth(951.32 MB)
tts_output.wav(1.11 MB)
v1.0.0(Jan 14, 2022)

This is release of Ukrainian TTS model and checkpoint using voice (10 hours) from M-AILABS Ukrainian dataset. Model was trained using 200 000 steps. Example:

https://user-images.githubusercontent.com/5759207/149566245-40656002-3999-48a8-b671-e0f74c3d6e2f.mp4
Source code(tar.gz)
Source code(zip)
config.json(6.44 KB)
example.mp4(32.07 KB)
v0.0.1(Oct 14, 2021)

This is release of Ukrainian TTS model and checkpoint using voice (10 hours) from M-AILABS Ukrainian dataset. Model was trained using 14145 steps. Example:

https://user-images.githubusercontent.com/5759207/140622395-9e734c95-159c-4d72-9f56-e8d1f1ac66c2.mp4
Source code(tar.gz)
Source code(zip)
config.json(4.96 KB)
test.mp4(32.73 KB)
vocoder_config.json(4.39 KB)

Owner

Yurii Paniv

AI and stuff

GitHub Repository https://huggingface.co/spaces/robinhad/ukrainian-tts

Text to speech for Vietnamese, ez to use, ez to update

Chào mọi người, đây là dự án mở nhằm giúp việc đọc được trở nên dễ dàng hơn. Rất cảm ơn đội ngũ Zalo đã cung cấp hạ tầng để mình có thể tạo ra app này

32 Jul 29, 2022

A Practitioner's Guide to Natural Language Processing

Learn how to process, classify, cluster, summarize, understand syntax, semantics and sentiment of text data with the power of Python! This repository contains code and datasets used in my book, Text

1.5k Jan 03, 2023

Deal or No Deal? End-to-End Learning for Negotiation Dialogues

Introduction This is a PyTorch implementation of the following research papers: (1) Hierarchical Text Generation and Planning for Strategic Dialogue (

1.4k Dec 29, 2022

Search with BERT vectors in Solr and Elasticsearch

123 Dec 29, 2022

Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.

🤗 Contributing to OpenSpeech 🤗 OpenSpeech provides reference implementations of various ASR modeling papers and three languages recipe to perform ta

513 Jan 03, 2023

pytorch-kaldi is a project for developing state-of-the-art DNN/RNN hybrid speech recognition systems. The DNN part is managed by pytorch, while feature extraction, label computation, and decoding are performed with the kaldi toolkit.

The PyTorch-Kaldi Speech Recognition Toolkit PyTorch-Kaldi is an open-source repository for developing state-of-the-art DNN/HMM speech recognition sys

2.3k Dec 27, 2022

SAVI2I: Continuous and Diverse Image-to-Image Translation via Signed Attribute Vectors

SAVI2I: Continuous and Diverse Image-to-Image Translation via Signed Attribute Vectors [Paper] [Project Website] Pytorch implementation for SAVI2I. We

44 Dec 30, 2022

Ray-based parallel data preprocessing for NLP and ML.

Wrangl Ray-based parallel data preprocessing for NLP and ML. pip install wrangl # for latest pip install git+https://github.com/vzhong/wrangl See exa

33 Dec 27, 2022

Python-zhuyin - An open source Python library that provides a unified interface for converting between Chinese pinyin and Zhuyin (bopomofo)

2 Dec 29, 2022

Official code for Spoken ObjectNet: A Bias-Controlled Spoken Caption Dataset

Official code for our Interspeech 2021 - Spoken ObjectNet: A Bias-Controlled Spoken Caption Dataset [1]*. Visually-grounded spoken language datasets c

3 Jan 26, 2022

Installation, test and evaluation of Scribosermo speech-to-text engine

Scribosermo STT Setup Scribosermo is a LGPL licensed, open-source speech recognition engine to "Train fast Speech-to-Text networks in different langua

3 Jun 20, 2022

Turkish Stop Words Türkçe Dolgu Sözcükleri

trstop Turkish Stop Words Türkçe Dolgu Sözcükleri In this repository I put Turkish stop words that is contained in the first 10 thousand words with th

103 Nov 12, 2022

This library is testing the ethics of language models by using natural adversarial texts.

prompt2slip This library is testing the ethics of language models by using natural adversarial texts. This tool allows for short and simple code and v

9 Dec 28, 2021

Blackstone is a spaCy model and library for processing long-form, unstructured legal text

Blackstone Blackstone is a spaCy model and library for processing long-form, unstructured legal text. Blackstone is an experimental research project f

579 Jan 08, 2023

تولید اسم های رندوم فینگیلیش

karafs کرفس تولید اسم های رندوم فینگیلیش installation ➜ pip install karafs usage دو زبانه ➜ karafs -n 10 توت فرنگی بی ناموس toot farangi-ye bi_namoos

36 Nov 24, 2022

Official implementations for various pre-training models of ERNIE-family, covering topics of Language Understanding & Generation, Multimodal Understanding & Generation, and beyond.

English|简体中文 ERNIE是百度开创性提出的基于知识增强的持续学习语义理解框架，该框架将大数据预训练与多源丰富知识相结合，通过持续学习技术，不断吸收海量文本数据中词汇、结构、语义等方面的知识，实现模型效果不断进化。ERNIE在累积 40 余个典型 NLP 任务取得 SOTA 效果，并在 G

5.4k Jan 03, 2023

Ukrainian TTS (text-to-speech) using Coqui TTS

Related tags

Overview

Ukrainian TTS 📢 🤖

Support

Example

How to use :

How to train:

Attribution

Comments

Error with file: speakers.pth

doc: fix examples in README

Problem

Fix

One vowel words in the end of the sentence aren't stressed

Error import StressOption

Vits improvements

Model improvement checklist

Releases(v4.0.0)

v4.0.0(Dec 10, 2022)

v3.0.0(Sep 14, 2022)

Example:

v3.0.0-alpha(Sep 8, 2022)

v2.0.0(Jul 10, 2022)

Example:

v2.0.0-beta(May 8, 2022)

Example:

v1.0.0(Jan 14, 2022)

v0.0.1(Oct 14, 2021)

Owner

Yurii Paniv

Text to speech for Vietnamese, ez to use, ez to update

A Practitioner's Guide to Natural Language Processing

Deal or No Deal? End-to-End Learning for Negotiation Dialogues

Search with BERT vectors in Solr and Elasticsearch

Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.

pytorch-kaldi is a project for developing state-of-the-art DNN/RNN hybrid speech recognition systems. The DNN part is managed by pytorch, while feature extraction, label computation, and decoding are performed with the kaldi toolkit.

SAVI2I: Continuous and Diverse Image-to-Image Translation via Signed Attribute Vectors

Ray-based parallel data preprocessing for NLP and ML.

Python-zhuyin - An open source Python library that provides a unified interface for converting between Chinese pinyin and Zhuyin (bopomofo)

Official code for Spoken ObjectNet: A Bias-Controlled Spoken Caption Dataset

Installation, test and evaluation of Scribosermo speech-to-text engine

Turkish Stop Words Türkçe Dolgu Sözcükleri

This library is testing the ethics of language models by using natural adversarial texts.

Blackstone is a spaCy model and library for processing long-form, unstructured legal text

تولید اسم های رندوم فینگیلیش

Official implementations for various pre-training models of ERNIE-family, covering topics of Language Understanding & Generation, Multimodal Understanding & Generation, and beyond.

A Japanese tokenizer based on recurrent neural networks

Unofficial Parallel WaveGAN (+ MelGAN & Multi-band MelGAN & HiFi-GAN & StyleMelGAN) with Pytorch

Tevatron is a simple and efficient toolkit for training and running dense retrievers with deep language models.

Topic Modelling for Humans