speechbrain
/

tts-tacotron2-ljspeech

speech-synthesis

Model card Files Files and versions Community

Mirco commited on May 28, 2022

Commit

f7400ef

•

1 Parent(s): 56bef66

Update README.md

Files changed (1) hide show

README.md +72 -1

README.md CHANGED Viewed

@@ -19,8 +19,79 @@ metrics:
 <iframe src="https://ghbtns.com/github-btn.html?user=speechbrain&repo=speechbrain&type=star&count=true&size=large&v=2" frameborder="0" scrolling="0" width="170" height="30" title="GitHub"></iframe>
 <br/><br/>
-# Work-in-Progress
 ### Limitations
 The SpeechBrain team does not provide any warranty on the performance achieved by this model when used on other datasets.

 <iframe src="https://ghbtns.com/github-btn.html?user=speechbrain&repo=speechbrain&type=star&count=true&size=large&v=2" frameborder="0" scrolling="0" width="170" height="30" title="GitHub"></iframe>
 <br/><br/>
+# Text-to-Speech (TTS) with Tacotron2 trained on LJSpeech
+This repository provides all the necessary tools for Text-to-Speech (TTS)  with SpeechBrain using a [Tacotron2](https://arxiv.org/abs/1712.05884) pretrained on [LJSpeech](https://keithito.com/LJ-Speech-Dataset/).
+The pre-trained model takes in input a short text and produces a spectrogram in output. One can get the final waveform by applying a vocoder (e.g., HiFIGAN) on top of the generated spectrogram.
+## Install SpeechBrain
+First of all, currently you need to install SpeechBrain from the source:
+1. Clone SpeechBrain:
+```bash
+git clone https://github.com/speechbrain/speechbrain/
+```
+2. Install it:
+```
+cd speechbrain
+pip install -r requirements.txt
+pip install -e .
+```
+Please notice that we encourage you to read our tutorials and learn more about
+[SpeechBrain](https://speechbrain.github.io).
+### Perform Text-to-Speech (TTS)
+```
+from speechbrain.pretrained import Tacotron2
+tacotron2 = Tacotron2.from_hparams(source="speechbrain/TTS_Tacotron2", savedir="tmpdir")
+mel_output, mel_length, alignment = tacotron2.encode_text("Mary had a little lamb")
+```
+If you want to generate multiple sentences in one-shot, you can do in this way:
+```
+from speechbrain.pretrained import Tacotron2
+tacotron2 = Tacotron2.from_hparams(source="speechbrain/TTS_Tacotron2", savedir="tmpdir")
+items = [
+       "A quick brown fox jumped over the lazy dog",
+       "How much wood would a woodchuck chuck?",
+       "Never odd or even"
+     ]
+mel_outputs, mel_lengths, alignments = tacotron2.encode_batch(items)
+```
+### Inference on GPU
+To perform inference on the GPU, add  `run_opts={"device":"cuda"}`  when calling the `from_hparams` method.
+### Training
+The model was trained with SpeechBrain.
+To train it from scratch follow these steps:
+1. Clone SpeechBrain:
+```bash
+git clone https://github.com/speechbrain/speechbrain/
+```
+2. Install it:
+```bash
+cd speechbrain
+pip install -r requirements.txt
+pip install -e .
+```
+3. Run Training:
+```bash
+cd https://github.com/speechbrain/speechbrain/tree/develop/recipes/LJSpeech/TTS/tacotron2
+python train.py --device=cuda:0 --max_grad_norm=1.0 --data_folder=/your_folder/LJSpeech-1.1 hparams/train.yaml
+```
+You can find our training results (models, logs, etc) [here](https://drive.google.com/drive/folders/1PKju-_Nal3DQqd-n0PsaHK-bVIOlbf26?usp=sharing).
 ### Limitations
 The SpeechBrain team does not provide any warranty on the performance achieved by this model when used on other datasets.