66m open source onnx native super minimum delay
tts
SupertonicTTS: Towards Highly Efficient and Streamlined...
We introduce SupertonicTTS, a novel text-to-speech (TTS) system designed for efficient and streamlined speech synthesis. SupertonicTTS comprises three components: a speech autoencoder for...
https://arxiv.org/abs/2503.23108

Self-Purifying Flow Matching
Training Flow Matching Models with Reliable Labels via Self-Purification
Training datasets are inherently imperfect, often containing mislabeled samples due to human annotation errors, limitations of tagging models, and other sources of noise. Such label contamination...
https://arxiv.org/abs/2509.19091

RobustSpeechFlow
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via...
While flow-matching text-to-speech (TTS) achieves strong zero-shot speaker similarity and naturalness, it remains susceptible to content fidelity issues, particularly skip and repeat errors from...
https://arxiv.org/abs/2605.22083


Seonglae Cho