glow-speak is a fast, local, neural text to speech system that uses eSpeak-ng as a text/phoneme front-end.

Last update: Dec 25, 2022

Related tags

Overview

Glow-Speak

glow-speak is a fast, local, neural text to speech system that uses eSpeak-ng as a text/phoneme front-end.

Installation

git clone https://github.com/rhasspy/glow-speak.git
cd glow-speak/

python3 -m venv .venv
source .venv/bin/activate
pip3 install --upgrade pip
pip3 install --upgrade setuptools wheel
pip3 install -f 'https://synesthesiam.github.io/prebuilt-apps/' -r requirements.txt

python3 setup.py develop
glow-speak --version

Voices

The following languages/voices are supported:

German
- de_thorsten
Chinese
- cmn_jing_li
Greek
- el_rapunzelina
English
- en-us_ljspeech
- en-us_mary_ann
Spanish
- es_tux
Finnish
- fi_harri_tapani_ylilammi
French
- fr_siwis
Hungarian
- hu_diana_majlinger
Italian
- it_riccardo_fasol
Korean
- ko_kss
Dutch
- nl_rdh
Russian
- ru_nikolaev
Swedish
- sv_talesyntese
Swahili
- sw_biblia_takatifu
Vietnamese
- vi_vais1000

Usage

Download Voices

glow-speak-download de_thorsten

Command-Line Synthesis

glow-speak -v en-us_mary_ann 'This is a test.' --output-file test.wav

HTTP Server

glow-speak-http-server --debug

Visit http://localhost:5002

Socket Server

Start the server:

glow-speak-socket-server --voice en-us_mary_ann --socket /tmp/glow-speak.sock

From a separate terminal:

echo 'This is a test.' | bin/glow-speak-socket-client --socket /tmp/glow-speak.sock | xargs aplay

Lines from client to server are synthesized, and the path to the WAV file is returned (usually in /tmp).

You might also like...

End-to-End Speech Processing Toolkit

ESPnet: end-to-end speech processing toolkit system/pytorch ver. 1.0.1 1.1.0 1.2.0 1.3.1 1.4.0 1.5.1 1.6.0 1.7.1 1.8.1 ubuntu18/python3.8/pip ubuntu18

5.9k Jan 3, 2023

Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.

OpenSpeech provides reference implementations of various ASR modeling papers and three languages recipe to perform tasks on automatic speech recogniti

26 Dec 14, 2022

Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.

OpenSpeech provides reference implementations of various ASR modeling papers and three languages recipe to perform tasks on automatic speech recogniti

86 Jun 11, 2021

Athena is an open-source implementation of end-to-end speech processing engine.

Athena is an open-source implementation of end-to-end speech processing engine. Our vision is to empower both industrial application and academic research on end-to-end models for speech processing. To make speech processing available to everyone, we're also releasing example implementation and recipe on some opensource dataset for various tasks (Automatic Speech Recognition, Speech Synthesis, Voice Conversion, Speaker Recognition, etc).

34 Sep 8, 2022

Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.

🤗 Contributing to OpenSpeech 🤗 OpenSpeech provides reference implementations of various ASR modeling papers and three languages recipe to perform ta

513 Jan 3, 2023

SHAS: Approaching optimal Segmentation for End-to-End Speech Translation

SHAS: Approaching optimal Segmentation for End-to-End Speech Translation In this repo you can find the code of the Supervised Hybrid Audio Segmentatio

21 Dec 20, 2022

An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition

CRNN paper：An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition 1. create your ow

3 Apr 2, 2022

Official PyTorch code for ClipBERT, an efficient framework for end-to-end learning on image-text and video-text tasks

Official PyTorch code for ClipBERT, an efficient framework for end-to-end learning on image-text and video-text tasks. It takes raw videos/images + text as inputs, and outputs task predictions. ClipBERT is designed based on 2D CNNs and transformers, and uses a sparse sampling strategy to enable efficient end-to-end video-and-language learning.

612 Jan 4, 2023

Neural Lexicon Reader: Reduce Pronunciation Errors in End-to-end TTS by Leveraging External Textual Knowledge

Neural Lexicon Reader: Reduce Pronunciation Errors in End-to-end TTS by Leveraging External Textual Knowledge This is an implementation of the paper,

19 Oct 14, 2022

Comments

AssertionError on web interface (only) - and Raspberry Pi Bullseye test

Hi Micheal,

great work again! :smiley:

I just saw this repository and thought I'd give it a try on my freshly installed Raspberry Pi 4 with 32bit Raspberry Pi OS Bullseye (Debian 11). Installation almost finished without errors! :partying_face: ... I just had to fix one thing: sudo apt-get install libatlas-base-dev After 15min I was already generating audio :grin: :+1:

When I tested en mary_ann and thorsten_de via the web interface I got this error as soon as my test sentence ended with a question mark:

DEBUG:glow-speak:ɪ_z ð_ɪ_s ɐ_n_ˈʌ_ð_ɚ t_ˈɛ_s_t? .
ERROR:glow_speak.http_server:
Traceback (most recent call last):
  File "/home/pi/glow-speak/.venv/lib/python3.9/site-packages/quart/app.py", line 1490, in full_dispatch_request
    result = await self.dispatch_request(request_context)
  File "/home/pi/glow-speak/.venv/lib/python3.9/site-packages/quart/app.py", line 1536, in dispatch_request
    return await self.ensure_async(handler)(**request_.view_args)
  File "/home/pi/glow-speak/glow_speak/http_server.py", line 484, in app_say
    wav_bytes = await text_to_wav(text, voice, **tts_args)
  File "/home/pi/glow-speak/glow_speak/http_server.py", line 323, in text_to_wav
    text_ids = text_to_ids(
  File "/home/pi/glow-speak/glow_speak/__init__.py", line 110, in text_to_ids
    text_ids = phonemes2ids(
  File "/home/pi/glow-speak/.venv/lib/python3.9/site-packages/phonemes2ids/__init__.py", line 190, in phonemes2ids
    maybe_extend_ids(sub_phoneme, word_ids, append_list=False)
  File "/home/pi/glow-speak/.venv/lib/python3.9/site-packages/phonemes2ids/__init__.py", line 108, in maybe_extend_ids
    maybe_ids = missing_func(phoneme)
  File "/home/pi/glow-speak/glow_speak/__init__.py", line 59, in guess_ids
    typing.List[Phoneme], guess_phonemes(phoneme, self.to_phonemes)
  File "/home/pi/glow-speak/.venv/lib/python3.9/site-packages/gruut_ipa/accent.py", line 159, in guess_phonemes
    assert dist_split is not None
AssertionError

Maybe some encoding error when reading the web input?

Speed seems pretty good, comparable to Larynx I'd say :+1: and I noticed the pronunciations have been improved for German :clap: :sunglasses:

opened by fquirin 0

Releases(v1.0)

v1.0(Oct 20, 2021)

Source code(tar.gz)
Source code(zip)
cmn_jing_li.tar.gz(101.49 MB)
de_thorsten.tar.gz(101.59 MB)
el_rapunzelina.tar.gz(101.34 MB)
en-us_ljspeech.tar.gz(101.66 MB)
en-us_mary_ann.tar.gz(101.69 MB)
es_tux.tar.gz(101.61 MB)
fi_harri_tapani_ylilammi.tar.gz(101.46 MB)
fr_siwis.tar.gz(101.59 MB)
hu_diana_majlinger.tar.gz(101.47 MB)
it_riccardo_fasol.tar.gz(101.70 MB)
ko_kss.tar.gz(101.58 MB)
nl_rdh.tar.gz(101.60 MB)
ru_nikolaev.tar.gz(101.64 MB)
sv_talesyntese.tar.gz(101.42 MB)
sw_biblia_takatifu.tar.gz(101.71 MB)
vi_vais1000.tar.gz(101.28 MB)

Owner

Rhasspy

Offline voice assistant

GitHub Repository

Research Code for NeurIPS 2020 Spotlight paper "Large-Scale Adversarial Training for Vision-and-Language Representation Learning": UNITER adversarial training part

VILLA: Vision-and-Language Adversarial Training This is the official repository of VILLA (NeurIPS 2020 Spotlight). This repository currently supports

109 Dec 31, 2022

TaCL: Improve BERT Pre-training with Token-aware Contrastive Learning

26 Oct 17, 2022

Two-stage text summarization with BERT and BART

Two-Stage Text Summarization Description We experiment with a 2-stage summarization model on CNN/DailyMail dataset that combines the ability to filter

6 Oct 22, 2022

TPlinker for NER 中文/英文命名实体识别

本项目是参考 TPLinker 中HandshakingTagging思想，将TPLinker由原来的关系抽取(RE)模型修改为命名实体识别(NER)模型。

113 Dec 28, 2022

Stack based programming language that compiles to x86_64 assembly or can alternatively be interpreted in Python

lang lang is a simple stack based programming language written in Python. It can

1 May 30, 2022

A complete NLP guideline for enthusiasts

NLP-NINJA A complete guide for Natural Language Processing in Python Table of Contents S.No. Topic Level Meaning 1 Tokenization 🤍 Beginner 2 Stemming

22 Dec 27, 2022

Knowledge Oriented Programming Language

KoPL: 面向知识的推理问答编程语言安装 | 快速开始 | 文档 KoPL全称 Knowledge oriented Programing Language, 是一个为复杂推理问答而设计的编程语言。我们可以将自然语言问题表示为由基本函数组合而成的KoPL程序，程序运行的结果就是问题的答案。目前，

62 Dec 12, 2022

Phrase-BERT: Improved Phrase Embeddings from BERT with an Application to Corpus Exploration

Phrase-BERT: Improved Phrase Embeddings from BERT with an Application to Corpus Exploration This is the official repository for the EMNLP 2021 long pa

70 Dec 11, 2022

Simple text to phones converter for multiple languages

Phonemizer -- foʊnmaɪzɚ The phonemizer allows simple phonemization of words and texts in many languages. Provides both the phonemize command-line tool

762 Dec 29, 2022

Official code for "Parser-Free Virtual Try-on via Distilling Appearance Flows", CVPR 2021

Parser-Free Virtual Try-on via Distilling Appearance Flows, CVPR 2021 Official code for CVPR 2021 paper 'Parser-Free Virtual Try-on via Distilling App

395 Jan 03, 2023

【原神】自动演奏风物之诗琴的程序

疯物之诗琴读取midi并自动演奏原神风物之诗琴。可以自定义配置文件自动调整音符来适配风物之诗琴。（原神1.4直播那天就开始做了！到现在才能放出来。。）如何使用在Release页面中下载打包好的程序和midi压缩包并解压。双击运行“疯物之诗琴.exe”。在原神中打开风物之诗琴，软件内输入

435 Jan 04, 2023

[Preprint] Escaping the Big Data Paradigm with Compact Transformers, 2021

Compact Transformers Preprint Link: Escaping the Big Data Paradigm with Compact Transformers By Ali Hassani[1]*, Steven Walton[1]*, Nikhil Shah[1], Ab

367 Dec 31, 2022

MHtyper is an end-to-end pipeline for recognized the Forensic microhaplotypes in Nanopore sequencing data.

MHtyper is an end-to-end pipeline for recognized the Forensic microhaplotypes in Nanopore sequencing data. It is implemented using Python.

6 Jun 27, 2022

HiFi DeepVariant + WhatsHap workflowHiFi DeepVariant + WhatsHap workflow

HiFi DeepVariant + WhatsHap workflow Workflow steps align HiFi reads to reference with pbmm2 call small variants with DeepVariant, using two-pass meth

2 May 14, 2022

Implementation of Natural Language Code Search in the project CodeBERT: A Pre-Trained Model for Programming and Natural Languages.

CodeBERT-Implementation In this repo we have replicated the paper CodeBERT: A Pre-Trained Model for Programming and Natural Languages. We are interest

4 Jul 01, 2022

Code for the paper "VisualBERT: A Simple and Performant Baseline for Vision and Language"

This repository contains code for the following two papers: VisualBERT: A Simple and Performant Baseline for Vision and Language (arxiv) with a short

464 Jan 04, 2023

Official code of our work, Unified Pre-training for Program Understanding and Generation [NAACL 2021].

PLBART Code pre-release of our work, Unified Pre-training for Program Understanding and Generation accepted at NAACL 2021. Note. A detailed documentat

138 Dec 30, 2022

keras implement of transformers for humans

4.8k Jan 03, 2023

A Multilingual Latent Dirichlet Allocation (LDA) Pipeline with Stop Words Removal, n-gram features, and Inverse Stemming, in Python.

Multilingual Latent Dirichlet Allocation (LDA) Pipeline This project is for text clustering using the Latent Dirichlet Allocation (LDA) algorithm. It

74 Oct 07, 2022

Experiments in converting wikidata to ftm

FollowTheMoney / Wikidata mappings This repo will contain tools for converting Wikidata entities into FtM schema. Prefixes: https://www.mediawiki.org/

2 Nov 12, 2021

glow-speak is a fast, local, neural text to speech system that uses eSpeak-ng as a text/phoneme front-end.

Related tags

Overview

Glow-Speak

Installation

Voices

Usage

Download Voices

Command-Line Synthesis

HTTP Server

Socket Server

You might also like...

End-to-End Speech Processing Toolkit

Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.

Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.

Athena is an open-source implementation of end-to-end speech processing engine.

Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.

SHAS: Approaching optimal Segmentation for End-to-End Speech Translation

An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition

Official PyTorch code for ClipBERT, an efficient framework for end-to-end learning on image-text and video-text tasks

Neural Lexicon Reader: Reduce Pronunciation Errors in End-to-end TTS by Leveraging External Textual Knowledge

Comments

AssertionError on web interface (only) - and Raspberry Pi Bullseye test

Releases(v1.0)

v1.0(Oct 20, 2021)

Owner

Rhasspy

Research Code for NeurIPS 2020 Spotlight paper "Large-Scale Adversarial Training for Vision-and-Language Representation Learning": UNITER adversarial training part

TaCL: Improve BERT Pre-training with Token-aware Contrastive Learning

Two-stage text summarization with BERT and BART

TPlinker for NER 中文/英文命名实体识别

Stack based programming language that compiles to x86_64 assembly or can alternatively be interpreted in Python

A complete NLP guideline for enthusiasts

Knowledge Oriented Programming Language

Phrase-BERT: Improved Phrase Embeddings from BERT with an Application to Corpus Exploration

Simple text to phones converter for multiple languages

Official code for "Parser-Free Virtual Try-on via Distilling Appearance Flows", CVPR 2021

【原神】自动演奏风物之诗琴的程序

[Preprint] Escaping the Big Data Paradigm with Compact Transformers, 2021

MHtyper is an end-to-end pipeline for recognized the Forensic microhaplotypes in Nanopore sequencing data.

HiFi DeepVariant + WhatsHap workflowHiFi DeepVariant + WhatsHap workflow

Implementation of Natural Language Code Search in the project CodeBERT: A Pre-Trained Model for Programming and Natural Languages.

Code for the paper "VisualBERT: A Simple and Performant Baseline for Vision and Language"

Official code of our work, Unified Pre-training for Program Understanding and Generation [NAACL 2021].

keras implement of transformers for humans

A Multilingual Latent Dirichlet Allocation (LDA) Pipeline with Stop Words Removal, n-gram features, and Inverse Stemming, in Python.

Experiments in converting wikidata to ftm