Skip to content

AI / LLM — ranked by measured latency

Model inference endpoints where latency is the product. Every figure below comes from real client-side traffic through the APIdown SDK, not from vendor status pages. Sorted fastest p95 first.

AI / LLM APIs ranked by latency, uptime, and reliability grade
Anthropicoperational100%0A
Cohereoperational100%0A
Google Geminioperational100%0A
Groqoperational100%0A
OpenAIoperational100%0A
Replicateoperational100%0A

Dashes mean we had no signals for that API in the window — not that it was down. See the full reliability leaderboard, subscribe to the AI / LLM feed, or build a stack watchlist.