Category: Speech Processing
A better way to create the full-context labels for HTS training data
One of my previous post describes my first attempt to generate training data for HTS system from recordings and transcripts: How to create full-context labels for your HTS system (update: not really…
ZureTTS from the eNTERFACE'14 workshop
Last month (from June 9 to July 4), I have a chance to go to Spain to work on a project name ZureTTS (which means Your Text-to-speech in Basque language) during the eNTERFACE'14 workshop. This was…
SailAlign and the error "ReadString: String too long"
If you have used SailAlign (or HTK) to do forced alignment on a large corpus, you may already encounter the error: ReadString: String too long. This error is actually thrown out from HTK, and a quick…
How to create full-context labels for your HTS system (update: not really worked)
Update : I later found out that the method described below did not work as expected. Tricking Festival by simply providing it with a custom monophone transcript will generates invalid .utt files.…
How to configure HTS for in-training synthesis with state-level alignment labels
Purpose Utilizing state-level alignment labels allows us to copy the prosody from one speaker and use it on another speaker’s acoustic model. This can be used to improve the synthesized results by…
How to configure HTS demo with STRAIGHT features for 16kHz training data
I have been using HTS for a while for my research on speech synthesis. Recently, I have had some problems when I tried to configure the HTS demo with STRAIGHT features to use 16k data instead of 48k.…