Speed, measured and published
We do not claim to be fast: we measure it every hour, against our own proxy and against the provider directly, and publish both numbers. The gap between them is what going through us costs.
Last measured: Oct 5, 2026, 5:36 PM
Last 7 days
| Model | First token (p50 / p95) | Tokens per second (p50) | What we add |
|---|---|---|---|
| auto | 264 / 2442 ms through our API 210 / 435 ms straight to the provider | 45.1 | 54 ms |
| deepseek-v4 | 601 / 1460 ms through our API 830 / 1441 ms straight to the provider | 6.4 | -229 ms |
| gemma-4-26b | 183 / 250 ms through our API 163 / 214 ms straight to the provider | 0.0 | 20 ms |
| glm-5.2 | 390 / 5453 ms through our API 306 / 887 ms straight to the provider | 0.0 | 84 ms |
| gpt-oss-120b | 257 / 760 ms through our API 193 / 299 ms straight to the provider | 40.5 | 64 ms |
| gpt-oss-20b | 302 / 354 ms through our API 214 / 329 ms straight to the provider | 9.2 | 88 ms |
| llama-3.3-70b | 262 / 1365 ms through our API 231 / 320 ms straight to the provider | 58.3 | 31 ms |
| mistral-medium | 306 / 986 ms through our API 215 / 687 ms straight to the provider | 68.3 | 91 ms |
| mistral-small | 207 / 631 ms through our API 145 / 353 ms straight to the provider | 134.9 | 62 ms |
| qwen3-235b | 273 / 357 ms through our API 229 / 280 ms straight to the provider | 81.1 | 44 ms |
| qwen3.5-397b | 450 / 546 ms through our API 399 / 583 ms straight to the provider | 0.0 | 51 ms |
| qwen3.5-9b | 290 / 398 ms through our API 240 / 320 ms straight to the provider | 0.0 | 50 ms |
| qwen3.6-27b | 293 / 359 ms through our API 273 / 1122 ms straight to the provider | 0.0 | 20 ms |
| qwen3.6-35b | 309 / 671 ms through our API 237 / 657 ms straight to the provider | 0.0 | 72 ms |
| qwen3.8-27b | 313 / 448 ms through our API 273 / 9012 ms straight to the provider | 7.6 | 40 ms |
Last 30 days
| Model | First token (p50 / p95) | Tokens per second (p50) | What we add |
|---|---|---|---|
| auto | 274 / 1755 ms through our API 224 / 1067 ms straight to the provider | 32.0 | 54 ms |
| deepseek-v4 | 448 / 5510 ms through our API 441 / 1736 ms straight to the provider | 6.7 | -229 ms |
| gemma-4-26b | 195 / 284 ms through our API 158 / 549 ms straight to the provider | 0.0 | 20 ms |
| glm-5.2 | 510 / 4975 ms through our API 332 / 1974 ms straight to the provider | 0.0 | 84 ms |
| gpt-oss-120b | 260 / 1235 ms through our API 206 / 388 ms straight to the provider | 31.5 | 64 ms |
| gpt-oss-20b | 288 / 451 ms through our API 216 / 326 ms straight to the provider | 10.4 | 88 ms |
| llama-3.3-70b | 258 / 1436 ms through our API 217 / 456 ms straight to the provider | 81.3 | 31 ms |
| mistral-medium | 261 / 1694 ms through our API 203 / 828 ms straight to the provider | 65.5 | 91 ms |
| mistral-small | 206 / 427 ms through our API 144 / 334 ms straight to the provider | 136.8 | 62 ms |
| qwen3-235b | 244 / 752 ms through our API 202 / 558 ms straight to the provider | 89.0 | 44 ms |
| qwen3.5-397b | 434 / 639 ms through our API 380 / 609 ms straight to the provider | 0.0 | 51 ms |
| qwen3.5-9b | 270 / 2026 ms through our API 236 / 931 ms straight to the provider | 0.0 | 50 ms |
| qwen3.6-27b | 293 / 359 ms through our API 273 / 1122 ms straight to the provider | 0.0 | 20 ms |
| qwen3.6-35b | 295 / 667 ms through our API 247 / 859 ms straight to the provider | 0.0 | 72 ms |
| qwen3.8-27b | 309 / 417 ms through our API 268 / 3156 ms straight to the provider | 7.8 | 40 ms |
How auto decides
104 requests named auto in the last 30 days. This is where they went and why, over every account at once — no single customer is visible in these numbers.
- Sent to the small model
- 4.8%
- Small-model answers passing checks
- —
- Default-model answers passing checks
- —
- Shadow agreement with the default model
- —
- Shadowed calls
- 0
Pass rates come from the automatic checks on aliases that point at auto: 0 checked answers in the window. A dash means too few to say.
The shadow figures come from keys whose owners switched shadow sampling on: one call in a hundred is also sent to the small model and the two answers compared, at the owner's expense and within their budget. Auto trusts a task to the small model on the shadow alone when at least 100 shadowed calls agree 95 % of the time, and keeps the default model when 100 or more agree under 80 %, whatever the checks say.
| Reason | What it means | Calls |
|---|---|---|
| router_off | The router is switched off: everything goes to the default model. | 0 |
| profile | A task profile: your alias's task type has a model that proved itself on it. | 0 |
| tools | The request declared tools, which the small model handles less reliably. | 0 |
| structured | JSON output was required. | 0 |
| long_answer | A long answer was asked for (max_tokens above the small model's comfort). | 0 |
| long_prompt | The prompt itself was long. | 24 |
| code | The prompt looks like code. | 0 |
| short | None of the above: a short, plain request, which is what the small model is for. | 5 |
Every answer that came through auto carries the reason in its X-Auto-Reason header.
Method
Every hour we send 3 requests per model, with the same prompt of about 800 input tokens and at most 600 output tokens, at temperature 0.7 so no answer can come from the cache. Time to first token is measured at the first chunk carrying text, not at the first byte. The same requests are repeated straight against the provider, from the same region.
The probe's source is in the repository: core/app/bench/probe.py