無料で使える中品質なテキスト読み上げソフトウェア、VOICEVOXのコア

Last update: Aug 29, 2022

Related tags

Audio voicevox_core

Overview

VOICEVOX CORE

VOICEVOX の音声合成コア。

Releases にビルド済みのコアライブラリ（.so/.dll）があります。

依存関係

CUDA 11.1 と CUDNN のインストールと LibTorch のダウンロードが必要です。

API

core.h をご参照ください。

サンプルの実行

まず Releases からコアライブラリが入った zip をダウンロードしておきます。

Python 3

ソースコードから実行

cd example/python

# example/python のディレクトリにコアライブラリが入った zip ファイルを展開

# Windowsの場合、DLLからLIBファイルの作成
./makelib.bat core

# 環境構築
pip install -r requirements.txt
python setup.py install  # Linuxの場合は先頭に `LIBRARY_PATH="$LIBRARY_PATH:."` が必要

# # うまく行かないときは毎回以下を実行すると良いかも
# python setup.py clean
# rm -r build *.cpp

# 実行（Windowsの場合）
PATH="$PATH:$HOME/libtorch/lib/" python run.py \
    --text "これは本当に実行できているんですか" \
    --speaker_id 1

# 実行（Windows以外の場合）
LD_LIBRARY_PATH="$LD_LIBRARY_PATH:/libtorch/lib/" python run.py \
    --text "これは本当に実行できているんですか" \
    --speaker_id 1

# 引数の紹介
# --text 読み上げるテキスト
# --speaker_id 話者ID
# --use_gpu GPUを使う
# --f0_speaker_id 音高の話者ID（デフォルト値はspeaker_id）
# --f0_correct 音高の補正値（デフォルト値は0。+-0.3くらいで結果が大きく変わります）

「ImportError: DLL load failed: 指定されたモジュールが見つかりません。」というエラーが出た場合は libtorch のパスが間違っているかもしれません。

Docker から

# イメージのビルド
docker build -t voicevox_core example/python

# コンテナの起動(音声を保存しておくボリュームを作成)
docker run -it -v ~/voicevox:/root/voice voicevox_core bash

# テスト音声 `おはようございます-1.wav` を生成
python run.py --text おはようございます --speaker_id 1
mv *.wav ~/voice
exit

# 音声の再生
aplay ~/voice/おはようございます-1.wav

その他の言語

サンプルコードを実装された際はぜひお知らせください。こちらに追記させて頂きます。

事例紹介

VOICEVOX ENGINE SHARP @yamachu ･･･ VOICEVOX ENGINE の C# 実装
Node VOICEVOX Engine @y-chan ･･･ VOICEVOX ENGINE の Node.js/C++ 実装

ライセンス

サンプルコードおよび core.h は MIT LICENSE です。

Releases にあるビルド済みのコアライブラリは別ライセンスなのでご注意ください。

無料で使える中品質なテキスト読み上げソフトウェア、VOICEVOXのコア

Related tags

Overview

VOICEVOX CORE

依存関係

API

サンプルの実行

Python 3

ソースコードから実行

Docker から

その他の言語

事例紹介

ライセンス

You might also like...

Releases(0.13.0-check-directml-rust.0)

0.13.0-check-directml-rust.0(Aug 18, 2022)

check-2022-04-11(Apr 10, 2022)

Owner

Hiroshiba

Python audio and music signal processing library

Full LAKH MIDI dataset converted to MuseNet MIDI output format (9 instruments + drums)

Nayeli: cool telegram groups vc music project

MUSIC-AVQA, CVPR2022 (ORAL)

Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding

Analysis of voices based on the Mel-frequency band

Some utils for auto speech recognition

controls volume using hand gestures

A python program to cut longer MP3 files (i.e. recordings of several songs) into the individual tracks.

Spotifyd - An open source Spotify client running as a UNIX daemon.

AudioDVP:Photorealistic Audio-driven Video Portraits

Terminal-based music player written in Python for the best music in the world 🎵 🎧 💻

GiantMIDI-Piano is a classical piano MIDI dataset contains 10,854 MIDI files of 2,786 composers

This is a realtime voice translator program which gets input from user at any language and converts it to the desired language that the user asks

Any-to-any voice conversion using synthetic specific-speaker speeches as intermedium features

A music player designed for a University Project.

無料で使える中品質なテキスト読み上げソフトウェア、VOICEVOXのコア

Bot duniya Music Player

Python implementation of the Short Term Objective Intelligibility measure

Stream Music 🎵 𝘼 𝙗𝙤𝙩 𝙩𝙝𝙖𝙩 𝙘𝙖𝙣 𝙥𝙡𝙖𝙮 𝙢𝙪𝙨𝙞𝙘 𝙤𝙣 𝙏𝙚𝙡𝙚𝙜𝙧𝙖𝙢 𝙂𝙧𝙤𝙪𝙥 𝙖𝙣𝙙 𝘾𝙝𝙖𝙣𝙣𝙚𝙡 𝙑𝙤𝙞𝙘𝙚 𝘾𝙝𝙖𝙩𝙨 𝘼𝙫𝙖𝙞𝙡?