FG-transformer-TTS Fine-grained style control in transformer-based text-to-speech synthesis

Last update: Dec 30, 2022

Related tags

Overview

LST-TTS

Official implementation for the paper Fine-grained style control in transformer-based text-to-speech synthesis. Submitted to ICASSP 2022. Audio samples/demo for our system can be accessed here

Setting up submodules

git submodule update --init --recursive

Get the waveglow vocoder checkpoint from here (This is from the NVIDIA official WaveGlow repo).

Setup environment

See docker/Dockerfile for the packages need to be installed.

Dataset preprocessing

LJSpeech

python preprocess_LJSpeech.py --datadir LJSpeechDir --outputdir OutputDir

VCTK

Get the leading and trailing scilence marks from this repo, and put vctk-silences.0.92.txt in your VCTK dataset directory.

python preprocess_VCTK.py --datadir VCTKDir --outputdir Output_Train_Dir

python preprocess_VCTK.py --datadir VCTKDir --outputdir Output_Test_Dir --make_test_set

--make_test_set: specify this flag to process the speakers in the test set, otherwise only process training speakers.

Training

LJSpeech

python train_TTS.py --precision 16 \
                    --datadir FeatureDir \
                    --vocoder_ckpt_path WaveGlowCKPT_PATH \
                    --sampledir SampleDir \
                    --batch_size 128 \
                    --check_val_every_n_epoch 50 \
                    --use_guided_attn \
                    --training_step 250000 \
                    --n_guided_steps 250000 \
                    --saving_path Output_CKPT_DIR \
                    --datatype LJSpeech \
                    [--distributed]

--distributed: enable DDP multi-GPU training
--batch_size: batch size per GPU, scale down if you train with multi-GPU and want to keep the same batch size
--check_val_every_n_epoch: sample and validate every n epoch
--datadir: output directory of the preprocess scripts

VCTK

python train_TTS.py --precision 16 \
                    --datadir FeatureDir \
                    --vocoder_ckpt_path WaveGlowCKPT_PATH \
                    --sampledir SampleDir \
                    --batch_size 64 \
                    --check_val_every_n_epoch 50 \
                    --use_guided_attn \
                    --training_step 150000 \
                    --n_guided_steps 150000 \
                    --etts_checkpoint LJSpeech_Model_CKPT \
                    --saving_path Output_CKPT_DIR \
                    --datatype VCTK \
                    [--distributed]

--etts_checkpoint: the checkpoint path of pretrained model (on LJ Speech)

Synthesis

We provide examples for synthesis of the system in synthesis.py, you can adjust this script to your own usage. Example to run synthesis.py:

python synthesis.py --etts_checkpoint VCTK_Model_CKPT \
                    --sampledir SampleDir \
                    --datatype VCTK \
                    --vocoder_ckpt_path WaveGlowCKPT_PATH

FG-transformer-TTS Fine-grained style control in transformer-based text-to-speech synthesis

Related tags

Overview

LST-TTS

Setting up submodules

Setup environment

Dataset preprocessing

LJSpeech

VCTK

Training

LJSpeech

VCTK

Synthesis

Owner

Li-Wei Chen

Computer-Vision-Paper-Reviews - Computer Vision Paper Reviews with Key Summary along Papers & Codes

Official Implementation for Fast Training of Neural Lumigraph Representations using Meta Learning.

Nest - A flexible tool for building and sharing deep learning modules

Randomized Correspondence Algorithm for Structural Image Editing

Wordplay, an artificial Intelligence based crossword puzzle solver.

Code for Paper "Evidential Softmax for Sparse MultimodalDistributions in Deep Generative Models"

Deep Anomaly Detection with Outlier Exposure (ICLR 2019)

3D Generative Adversarial Network

CLIP (Contrastive Language–Image Pre-training) for Italian

AdvStyle - Official PyTorch Implementation

Predict stock movement with Machine Learning and Deep Learning algorithms

A Demo server serving Bert through ONNX with GPU written in Rust with <3

In this project we use both Resnet and Self-attention layer for cat, dog and flower classification.

Exploration of some patients clinical variables.

LSTM built using Keras Python package to predict time series steps and sequences. Includes sin wave and stock market data

This is the source code of the solver used to compete in the International Timetabling Competition 2019.

Official Code for ICML 2021 paper "Revisiting Point Cloud Shape Classification with a Simple and Effective Baseline"

Source code for our paper "Molecular Mechanics-Driven Graph Neural Network with Multiplex Graph for Molecular Structures"

Largest list of models for Core ML (for iOS 11+)

PSGAN running with ncnn⚡妆容迁移/仿妆⚡Imitation Makeup/Makeup Transfer⚡