Check the catalog row before selecting a model. Voice names, output formats, and language support are provider-specific.
Families
Speech models
Speech generation and transcription providers.
Speech routes use the same paid inference pipeline as text and media routes. Text-to-speech can return binary audio instead of JSON.
Provider families with speech or transcription adapter paths include: