Try natural speech, Nepali and English, voice cloning, and expressive delivery on the RTX 5090. Audio plays as it arrives.
Checking server…Upload a clean 5–15 second recording of a voice you have permission to use. Add its exact transcript for better fidelity. Reference uploads are deleted after the request.
Ready.
First audio is measured in this browser, including network time. Real-time factor = generation time ÷ audio duration; lower is faster. VITS has been benchmarked and is stopped.
Evaluation lab • Boson Higgs TTS 3 • Test lab