EnroutiaEU

Audits: replay your real tasks on cheaper models

An audit takes the prompts your automations actually send, with the answers your current model gave, replays them on the catalogue's models and says — per task — which cheaper model would have agreed with production and how often. The report names the candidates as Model A, B, C; unlocking the names is the thing that is sold, and it is done with a person. Everything up to that point is now yours to drive: from the panel or over HTTP.

The collector

The collector is an n8n workflow we publish: it reads your executions, groups them by task — the workflow and the node that called a model — keeps a sample of each, and posts the lot to POST /v1/audits with an account token carrying the audits permission. A continuous mode posts a batch a day against a fixed audit id. Anything can play collector: the payload is documented below.

Sending us content is the one thing nothing else you do implies, so the intake asks for one switch first: Keep call content for auditing, in the panel's privacy settings. Off, the collector's POST is a 403 content_audit_disabled. Samples are encrypted at rest and deleted with the audit. What is stored, and for how long

POST /v1/audits

One payload: a source, and a list of tasks with their current model, monthly volume, prompt template and samples — each sample the messages that were sent, the answer production gave and the token counts. The receipt says how many tasks are auditable: a task without a classifiable shape (free-form chat, say) is kept but marked out of scope.

curl -X POST https://api.enroutia.com/v1/audits \
  -H "Authorization: Bearer pat_live_…" \
  -H "Content-Type: application/json" \
  -d '{
    "audit_id": "1f0c8e3a-…",
    "source": "n8n",
    "tasks": [
      {
        "task_id": "Lead scoring · Classify",
        "current_model": "gpt-4o-mini",
        "monthly_calls": 12000,
        "prompt_template": "Classify this lead: {{ $json.text }}",
        "samples": [
          {
            "messages": [{"role": "user", "content": "Classify this lead: …"}],
            "production_output": "{\"tier\": \"warm\"}",
            "tokens_in": 412, "tokens_out": 18
          }
        ]
      }
    ]
  }'

HTTP/1.1 201 Created
{"audit_id": "1f0c8e3a-…", "status": "received", "tasks_received": 1,
 "tasks_auditable": 1, "tasks_out_of_scope": 0, "samples_received": 1}

POST /v1/audits/{id}/batches takes the same body against a standing id: the first batch creates the audit and every later one appends, deduplicating samples on content. GET /v1/audits/{id} reads it; DELETE removes it, samples and results included, immediately.

Reading your audits

The list carries counts and the last job; the detail carries the tasks with their type, current model, monthly calls and sample count. Neither ever returns stored content.

GET https://api.enroutia.com/v1/audits
{"audits": [{"id": "1f0c8e3a-…", "status": "approved", "tasks": 3, "samples": 84,
             "created_at": "2026-09-20T08:14:00+00:00",
             "last_job": {"id": "b2…", "kind": "run", "status": "done",
                          "created_at": "…", "finished_at": "…", "error": null}}]}

GET https://api.enroutia.com/v1/audits/{id}
{"audit_id": "1f0c8e3a-…", "source": "n8n", "status": "approved", "created_at": "…",
 "tasks": [{"task_id": "Lead scoring · Classify", "task_type": "lead_classification",
            "current_model": "gpt-4o-mini", "auditable": true,
            "monthly_calls": 12000, "window_days": 30, "samples": 40}],
 "collections": {"count": 2, "first_at": "…", "last_at": "…"}}

Approve

An audit arrives as received. Approving it says the samples are yours to replay; until then nothing runs. It is one POST, and the panel has a button for it.

POST https://api.enroutia.com/v1/audits/{id}/approve
→ 200 {"status": "approved"}      · 409 unless the audit is in "received"

Run the replay

Running queues a job; the scheduler picks it up within a minute, replays every auditable sample on every candidate model, judges the answers against production, and writes the verdicts. The 202 carries the job; poll the jobs route or subscribe to the webhook. Three 409s can come back:

POST https://api.enroutia.com/v1/audits/{id}/run
→ 202 {"job": {"id": "b2…", "kind": "run", "status": "queued"}}

GET https://api.enroutia.com/v1/audits/{id}/jobs
{"jobs": [{"id": "b2…", "kind": "run", "status": "running",
           "created_at": "…", "started_at": "…", "finished_at": null, "error": null}]}
  • audit_needs_admin — past the self-serve limits below. Write to us and a person runs it.
  • already_running — a job on this audit is queued or running. Wait for it.
  • bad_state — the audit is not approved, or has nothing auditable.

Self-serve limits

The replay runs on Enroutia's credit, not yours: you pay nothing to find out. The fence that makes that possible:

  • 3 self-serve runs per audit. A fourth answers 409 audit_needs_admin.
  • 500 cells per run — a cell is one sample on one candidate model. Bigger audits are run by a person, on request.
  • Nothing is debited from your balance for a replay. The verdicts and the anonymised report are free.
  • Unlocking the model names is not on the API: it is a purchase, agreed with a person, and applied at our side.

Verdicts and the report

Per task and per candidate: how many samples were compared, how many the candidate agreed with production on, and how many cells failed (the provider errored or the answer was unusable). Counters only. The report is one HTML file, anonymised, with the same numbers and the disagreements explained in shape, never in content.

GET https://api.enroutia.com/v1/audits/{id}/verdicts
{"tasks": [
  {"task_id": "Lead scoring · Classify", "task_type": "lead_classification",
   "current_model": "gpt-4o-mini",
   "models": [
     {"alias": "Model A", "samples_compared": 40, "agreements": 38, "failed_cells": 0},
     {"alias": "Model B", "samples_compared": 40, "agreements": 33, "failed_cells": 2}
   ]}
]}

GET https://api.enroutia.com/v1/audits/{id}/report?state=anon      → text/html, self-contained

state=anon is the only state HTTP serves. The report opens in the browser with your session, or over the API with the token: it is self-contained, so it can be forwarded as it is.

Webhooks

Subscribe an account webhook to audit_run_completed and audit_run_failed. Both carry the audit and job ids; the completed one adds the cells that ran and, when known, what the run cost us in millicents — so the number you would have been billed for is on the record even though you were not. Signed like every other event: see the brand reports page for the signature.

X-Webhook-Event: audit_run_completed

{
  "event": "audit_run_completed",
  "created_at": "2026-09-20T09:02:11+00:00",
  "project_id": "…",
  "data": {"audit_id": "1f0c8e3a-…", "job_id": "b2…", "cells": 120, "cost_millicents": 8400}
}

X-Webhook-Event: audit_run_failed
{"event": "audit_run_failed", "data": {"audit_id": "1f0c8e3a-…", "job_id": "b2…", "cells": 0}}

In the panel

Audits, after Tests in the rail: the list, the detail with its tasks and anonymised verdicts, and the three buttons — Approve, Run replay, Open report. The run button asks first and says what it will replay and on whose credit.

See also: Privacy and aggregated metrics · Brand reports · MCP · Limits