AI / LLM — ranked by measured latency
Model inference endpoints where latency is the product. Every figure below comes from real client-side traffic through the APIdown SDK, not from vendor status pages. Sorted fastest p95 first.
| Anthropic | operational | — | — | — | 100% | 0 | A |
|---|---|---|---|---|---|---|---|
| Cohere | operational | — | — | — | 100% | 0 | A |
| Google Gemini | operational | — | — | — | 100% | 0 | A |
| Groq | operational | — | — | — | 100% | 0 | A |
| OpenAI | operational | — | — | — | 100% | 0 | A |
| Replicate | operational | — | — | — | 100% | 0 | A |
Dashes mean we had no signals for that API in the window — not that it was down. See the full reliability leaderboard, subscribe to the AI / LLM feed, or build a stack watchlist.