EnroutiaEU

Privacy and aggregated metrics

Prompts and responses are never stored. What follows is what that leaves us with, what the automatic checks keep, and the two things you can choose to share.

What stays private

Your prompt and the model's response pass through the gateway in memory and are gone when the call ends. The automatic checks run in that same memory and keep only booleans and numbers: whether the JSON was valid, whether the required fields were present, whether the language was the expected one, whether the length was within range, whether the provider errored, the latency, the cost, and whether the call was retried. No text, no fragments, no hashes of the text.

Option one: content audit

If you want to read what your automations sent and received — to debug a workflow, or to judge quality by eye — turn on the content audit for your workspace and choose for how many days the content is kept. It is off by default, and turning it off deletes what was kept.

Option two: aggregated metrics

The panel has a switch that reads: “Contribute aggregated metrics to improve Enroutia's recommendations”. With it on, the outcome of your calls joins a pool that tells every workspace which model does best at which task. This is exactly what is shared:

  • the task type
  • the language
  • the model and its version
  • a token range — never the count, never the content
  • the cost
  • the latency
  • the validation result
  • whether the run was marked useful, needs edit, or incorrect
  • the day, aggregated

And this is what is never shared:

  • the prompt
  • the response
  • your name, or your customer's
  • private URLs
  • end users
  • personal data
  • unique prompt hashes

The cohort minimum

Nothing is shown to anyone until at least 10 workspaces and 100 comparable runs have contributed for a task and language. Below that the panel says “Not enough aggregated data yet” and gives no counts at all, so a small pool cannot be read back to the one workspace in it.

Revocation

Turn the switch off and the rows you contributed are deleted, not merely stopped. The benchmarks are recomputed without them the next time they run.

Known limits

The checks are heuristic. Language detection is a heuristic and can misjudge short or mixed-language text. There is no semantic judgement of quality without the content audit: the gateway can tell you the answer was valid JSON, in Spanish, of the right length — not that it was right. Human review is not part of the product.

PII redaction at the edge

A key can ask for personal identifiers to be replaced by placeholders before the prompt leaves the gateway. It is off by default and switched on per key from Keys. With it on, the text of every message — and every text block, and the input of a Responses call — is scanned for:

  • email addresses → [EMAIL_1], [EMAIL_2]…
  • phone numbers → [PHONE_n]: international ones with a country code; Spanish ones when written with phone spacing (612 345 678, 612 34 56 78) or, bare, when a phone word precedes them (tel, teléfono, móvil, whatsapp, +34, llámame, contacto). Never after pedido, factura, expediente, ref, nº or importe, and never before € or EUR — an order number and an amount are not phones
  • DNI, NIE and NIF numbers → [ID_n]
  • IBANs → [IBAN_n]
  • card numbers that pass the Luhn check → [CARD_n]
  • IPv4 addresses → [IP_n], unless a version word precedes them (versión 10.0.0.1 is software)
  • personal names, only when one or more titles precede them (Sr., Sra., Dña., D., Dr., Prof. Dr.) → [NAME_n]; a title before a post (Sr. Director General) or inside an abbreviation (R.D.) is not a name

The table that maps each placeholder back to its original lives in the memory of that one request and nowhere else: not in the usage record, not in a log, not in the callbacks that bill the call. When the provider answers, the placeholders in the answer — in the text and in the arguments of any tool call — are put back and the table is discarded. The answer carries X-PII-Redacted with the number of replacements, so a workflow can tell a redacted call from a plain one. Two limits to know: placeholders are numbered per request, so in a conversation whose history already carries [EMAIL_1] from an earlier answer the new turn's [EMAIL_1] may be a different address — keep your own map if you keep redacted history; and a bare identifier with no context at all (a nine-digit number alone) is left in place, because a false negative is cheaper than a prompt full of holes. Card numbers after IMEI, serie or S/N are device serials, not cards.

Streamed answers are restored too, as they stream, on all three dialects. A model often writes a placeholder in pieces — [EMA, then IL_1] — so the gateway holds back the few characters that could still become one until they are whole, and sends the original in their place; nothing longer than the longest placeholder of that request is ever held, and only on a key with redaction on. What cannot be restored is a placeholder the model did not write back as it was given: one it rewrote ([Email 1]) or one cut off by max_tokens arrives as the model wrote it. X-PII-Redacted is on streamed answers too.

X-PII-Redacted: 3

"content": "Write to [EMAIL_1] and call [PHONE_1] about invoice [IBAN_1]."

A redacted call is never cached — the placeholders vary between requests — and never shadow-sampled. The gateway counts redactions per kind as a metric with no content in it; nothing about what was redacted is kept.

Shadow sampling

Off by default, and switched on per key from Keys by an owner or admin. With it on, one call in a hundred on that key is also sent to a cheaper model from the catalogue — a second provider, chosen by us — with the same prompt. The two answers are compared in memory and only the score is stored: which model, which candidate, the task type, whether they agreed, and what the shadow call cost. Not the prompt, not either answer.

The shadow call is a real call: you pay it, at the candidate's price, out of your balance. Each key carries a monthly shadow budget, 50 cents by default and settable from 1 to 500 cents; when the month's shadow calls reach it, sampling pauses on that key until the month turns, the panel says so on the key, and you get a shadow_budget_reached webhook if you subscribed. A shadow call is never itself sampled.

Two settings turn it off on their own, whatever the switch says: Confidential Mode, because the mode exists to keep the prompt from going anywhere but the one provider that answers it, and PII redaction, because a prompt full of placeholders answered twice compares nothing. Turn the shadow off and nothing more is sent; the scores already kept stay, since they contain no content.

See also: Aliases · Credits and balance · Limits · Teams and roles