Limits
The limits exist so that a bug in an automation does not turn into an invoice.
| Limit | Note |
|---|---|
| 60 requests per minute | Per key. Configurable — write to us if you need more. |
| 100,000 tokens per minute | Per key. |
| 1 MB per request (25 MB for audio) | You get a 413 if you exceed it. |
| 20 active keys per project | Revoke the ones you are not using. |
| 5 comparisons a day without an account | 20 once you confirm your email; no reasonable limit with an account. |
| Reply length | If you send no max_tokens we apply 2,048. The per-request maximum is 8,192: above that it is trimmed to that value and the request is served anyway, with no error. |
| Spend caps you set yourself | Every key can carry a monthly cap, a daily cap and a ceiling on what one call may spend on its answer. All three are optional and all three are set in the panel or over the API; a key created by an agent must carry at least the monthly one. |
| Workspace and project budgets | Two more optional ceilings above the keys: everything the workspace spends in a month, and what one project — one client, for an agency — spends. Each answers with its own 402 code; the Budgets page lists all of them. |
Retries
We never retry a generation that already started: you would pay twice for the same work. We do switch provider if the first one never answers.
When a per-call ceiling shortens an answer
A ceiling on one call is applied as a length bound at the door, priced against the model you are billed for. When it is what shortened a reply, the response says so — otherwise a truncated answer looks like the model stopping on its own:
X-Call-Budget-Capped: true
X-Max-Tokens-Applied: 6000If you ask for a model we do not have
We never return an error for a model name: we serve it with “auto” and bill you its price, not the price of the model you named. So it is not a surprise, the response says so in its headers. If you are benchmarking us against another provider, read them: it is the difference between measuring the same model and measuring two different ones.
X-Requested-Model: gpt-4o
X-Resolved-Model: gpt-oss-120b
X-Model-Redirected: true
Warning: 199 - "gpt-4o is not in the catalogue; served gpt-oss-120b."A name of your own
The redirect rule above is about catalogue names. A name you created in the panel, `enroutia/keywords-prod`, is an alias: it resolves to whichever model you pointed it at, and `X-Resolved-Model` still names that catalogue model. If that model is ever retired, the alias falls back to “auto” the same way and never answers 404.
How auto decides
The auto alias reads the request before choosing. A task profile wins first: if your alias has a task type and the router has learnt a model for it, that model serves it (X-Auto-Reason: profile: followed by the task type). Otherwise the request goes to the default model when it declares tools, requires JSON, asks for a long answer, carries a long prompt or looks like code — reasons tools, structured, long_answer, long_prompt and code — and to the small model when none of that applies: short. With the router off, every call says router_off. The reason travels back in X-Auto-Reason on every answer served through auto, and the running tally is published on the speed page.
X-Requested-Model: auto
X-Resolved-Model: mistral-small
X-Auto-Reason: shortHedging
Every primary deployment that has a failover carries a read timeout on its stream, hedge_ttft_ms. If the provider has sent nothing for that long — before the first token, or between two chunks of an answer already under way — the request leaves it and goes to the failover provider for that alias, the same one that serves it when the first provider is down, and the answer says so with X-Served-By: fallback. It is a timeout between bytes, not a clock on the first token: a stalled generation can be abandoned and repeated on the failover, which is why the value is never rendered under 5 seconds. Whether the proxy also measures the first token on its own is being verified (V-133); until then this is what the setting does. It is set per instance, and 0 turns it off. A primary with no failover has no hedge: a timeout there would turn a slow answer into no answer.
X-Served-By: fallbackSemantic cache
The exact cache serves an identical request from memory for up to an hour. The semantic cache, switched on per key from Keys, does the same for a request that is not identical but means the same: the last user message is embedded with the embed alias and compared against what the same alias answered in the last hour for the same key, under the same system prompt, instructions and earlier turns; at or above the similarity threshold (0.97 by default) the stored answer is served without a call. Two keys never share an answer, and the same question under a different system prompt is a different question. The embedding is billed to the key at the embed price.
It only applies with temperature=0, without tools, and on a key that has neither Confidential Mode nor redaction on: a cached answer is the right answer only where the same question is meant to get the same reply. Entries live in memory, per key and per conversation context, for an hour, with the same promise as the exact cache — nothing on disk.
See also: Budgets · Aliases · Credits and balance · Tags and cost per automation · Privacy and aggregated metrics