Call this before every inference request

Decide if the call is worth it, where to send it, and how to optimise latency and cost across 15 providers.

15 providers 28 endpoints 24/7 monitoring 5 continents 99.7% uptime
First call — call this before every inference request
GET /v1/should-call
No parameters required • Returns: should_call, recommended_provider, expected_latency_ms, expected_cost, confidence_score
Value: typically reduces latency 20–60% and cost 10–40%
🛋 The Lazy Button
for when you want to know everything in a hurry

What is InferenceLatency.com?

InferenceLatency.com is a real-time AI inference monitoring platform that continuously tests 15 enabled LLM API providers — OpenAI, Anthropic, Groq, OpenRouter, Google Gemini, DeepSeek, Cohere, Together AI, Fireworks AI, Cerebras, SambaNova, xAI (Grok), Mistral AI, Perplexity, Nvidia NIM — measuring time to first token (TTFT), throughput, cost, and reliability. Google Gemini is currently rate-limited, so individual tests may report it as unavailable. Results are available via free JSON APIs with no authentication required.

The primary endpoint is GET /v1/should-call — a pre-inference decision engine that AI agents can call before every LLM request to get a recommended provider, expected latency, expected cost, and a confidence score. Calling this endpoint before inference typically reduces latency by 20–60% and cost by 10–40% compared to hardcoded provider selection.

Measurements use a standardised 1-token prompt sent simultaneously to all providers via their official APIs. Timing is millisecond-precise from request start to first token received (TTFT). Results are stored in a rolling 48-hour database to compute statistically reliable P50, P95, and P99 latency percentiles — the metrics that matter most for SLA planning and production reliability.

Google Gemini ⚠
🏆 Check out our partners at inferencewars.com for the latest inference war reports and provider leaderboards