Open Source
CSM (Conversational Speech Model) by Sesame AI Labs is an open-source speech generation model that produces remarkably natural, human-sounding audio. It models conversational prosody, emotion, and rhythm in a way that outperforms most existing TTS systems on naturalness. CSM gained widespread attention for generating speech indistinguishable from real human voices. The model weights and inference code are available on GitHub and Hugging Face.