| Alibaba | Text (Qwen), speech (CosyVoice), video (Wan), vision, embeddings |
| Roboflow | Object detection, segmentation, OCR, image embeddings |
| Cloudflare | Edge inference, LLMs, vision, speech via Workers AI |
| Gemini | Multimodal text, vision, realtime, TTS, video |
| Azure | Hosted OpenAI models with data residency |
| OpenAI | GPT, Sora, embeddings, realtime |
| Fireworks | Open-source LLMs and rerankers |
| Cartesia | Low-latency text-to-speech |
| ElevenLabs | Voice generation, music, realtime transcription |
| Deepgram | Speech recognition (Nova-3) |
| ASICloud | Specialized inference |