Speed, measured and published

We do not claim to be fast: we measure it every hour, against our own proxy and against the provider directly, and publish both numbers. The gap between them is what going through us costs.

Last measured: Oct 5, 2026, 5:36 PM

Last 7 days

ModelFirst token (p50 / p95)Tokens per second (p50)What we add
auto
264 / 2442 ms through our API
210 / 435 ms straight to the provider
45.154 ms
deepseek-v4
601 / 1460 ms through our API
830 / 1441 ms straight to the provider
6.4-229 ms
gemma-4-26b
183 / 250 ms through our API
163 / 214 ms straight to the provider
0.020 ms
glm-5.2
390 / 5453 ms through our API
306 / 887 ms straight to the provider
0.084 ms
gpt-oss-120b
257 / 760 ms through our API
193 / 299 ms straight to the provider
40.564 ms
gpt-oss-20b
302 / 354 ms through our API
214 / 329 ms straight to the provider
9.288 ms
llama-3.3-70b
262 / 1365 ms through our API
231 / 320 ms straight to the provider
58.331 ms
mistral-medium
306 / 986 ms through our API
215 / 687 ms straight to the provider
68.391 ms
mistral-small
207 / 631 ms through our API
145 / 353 ms straight to the provider
134.962 ms
qwen3-235b
273 / 357 ms through our API
229 / 280 ms straight to the provider
81.144 ms
qwen3.5-397b
450 / 546 ms through our API
399 / 583 ms straight to the provider
0.051 ms
qwen3.5-9b
290 / 398 ms through our API
240 / 320 ms straight to the provider
0.050 ms
qwen3.6-27b
293 / 359 ms through our API
273 / 1122 ms straight to the provider
0.020 ms
qwen3.6-35b
309 / 671 ms through our API
237 / 657 ms straight to the provider
0.072 ms
qwen3.8-27b
313 / 448 ms through our API
273 / 9012 ms straight to the provider
7.640 ms

Last 30 days

ModelFirst token (p50 / p95)Tokens per second (p50)What we add
auto
274 / 1755 ms through our API
224 / 1067 ms straight to the provider
32.054 ms
deepseek-v4
448 / 5510 ms through our API
441 / 1736 ms straight to the provider
6.7-229 ms
gemma-4-26b
195 / 284 ms through our API
158 / 549 ms straight to the provider
0.020 ms
glm-5.2
510 / 4975 ms through our API
332 / 1974 ms straight to the provider
0.084 ms
gpt-oss-120b
260 / 1235 ms through our API
206 / 388 ms straight to the provider
31.564 ms
gpt-oss-20b
288 / 451 ms through our API
216 / 326 ms straight to the provider
10.488 ms
llama-3.3-70b
258 / 1436 ms through our API
217 / 456 ms straight to the provider
81.331 ms
mistral-medium
261 / 1694 ms through our API
203 / 828 ms straight to the provider
65.591 ms
mistral-small
206 / 427 ms through our API
144 / 334 ms straight to the provider
136.862 ms
qwen3-235b
244 / 752 ms through our API
202 / 558 ms straight to the provider
89.044 ms
qwen3.5-397b
434 / 639 ms through our API
380 / 609 ms straight to the provider
0.051 ms
qwen3.5-9b
270 / 2026 ms through our API
236 / 931 ms straight to the provider
0.050 ms
qwen3.6-27b
293 / 359 ms through our API
273 / 1122 ms straight to the provider
0.020 ms
qwen3.6-35b
295 / 667 ms through our API
247 / 859 ms straight to the provider
0.072 ms
qwen3.8-27b
309 / 417 ms through our API
268 / 3156 ms straight to the provider
7.840 ms

How auto decides

104 requests named auto in the last 30 days. This is where they went and why, over every account at once — no single customer is visible in these numbers.

Sent to the small model
4.8%
Small-model answers passing checks
—
Default-model answers passing checks
—
Shadow agreement with the default model
—
Shadowed calls
0

Pass rates come from the automatic checks on aliases that point at auto: 0 checked answers in the window. A dash means too few to say.

The shadow figures come from keys whose owners switched shadow sampling on: one call in a hundred is also sent to the small model and the two answers compared, at the owner's expense and within their budget. Auto trusts a task to the small model on the shadow alone when at least 100 shadowed calls agree 95 % of the time, and keeps the default model when 100 or more agree under 80 %, whatever the checks say.

Why auto routed each request, with the number of calls per reason
ReasonWhat it meansCalls
router_offThe router is switched off: everything goes to the default model.0
profileA task profile: your alias's task type has a model that proved itself on it.0
toolsThe request declared tools, which the small model handles less reliably.0
structuredJSON output was required.0
long_answerA long answer was asked for (max_tokens above the small model's comfort).0
long_promptThe prompt itself was long.24
codeThe prompt looks like code.0
shortNone of the above: a short, plain request, which is what the small model is for.5

Every answer that came through auto carries the reason in its X-Auto-Reason header.

Method

Every hour we send 3 requests per model, with the same prompt of about 800 input tokens and at most 600 output tokens, at temperature 0.7 so no answer can come from the cache. Time to first token is measured at the first chunk carrying text, not at the first byte. The same requests are repeated straight against the provider, from the same region.

The probe's source is in the repository: core/app/bench/probe.py