Skip to main content
Speech routes use the same paid inference pipeline as text and media routes. Text-to-speech can return binary audio instead of JSON. Provider families with speech or transcription adapter paths include: Check the catalog row before selecting a model. Voice names, output formats, and language support are provider-specific.