EnroutiaEU

Responses API and the Anthropic SDK

The gateway speaks three dialects on one base URL and one key: chat completions, OpenAI's Responses API and Anthropic's Messages API. All three reach the same models — any catalogue alias, auto, or an alias of yours — and answer with the same headers. This page is the two newer ones.

POST /v1/responses

The Responses API takes input instead of messages, and max_output_tokens instead of max_tokens. Send a string for a one-turn call, or a list of items with a role and a content for a conversation. stream: true streams the same server-sent events the OpenAI SDK expects.

curl https://api.enroutia.com/v1/responses \
  -H "Authorization: Bearer YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "input": "Summarise this in one sentence: …",
    "max_output_tokens": 200
  }'

A list works the way messages does: a system item first if you want one, then the turns.

"input": [
  {"role": "system", "content": "Answer in Spanish."},
  {"role": "user", "content": "¿Qué es una API?"}
]

The answer is one output item of type message whose content carries the text; usage counts input_tokens and output_tokens, which is what your credit is debited on — the same AI Run arithmetic as chat completions.

HTTP/1.1 200 OK
X-Resolved-Model: mistral-small
X-Request-Id: req_…
X-Auto-Reason: short

{
  "id": "resp_…",
  "object": "response",
  "model": "auto",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [{"type": "output_text", "text": "Una API es …"}]
    }
  ],
  "usage": {"input_tokens": 21, "output_tokens": 38, "total_tokens": 59}
}

With the OpenAI Python SDK

Version 1.66 or later has client.responses. Change base_url and api_key and nothing else; output_text is the SDK's shortcut to the text.

from openai import OpenAI

client = OpenAI(base_url="https://api.enroutia.com/v1", api_key="YOUR_KEY")

response = client.responses.create(
    model="auto",
    input="Summarise this in one sentence: …",
    max_output_tokens=200,
)
print(response.output_text)

# Streaming: the same call with stream=True, one event per delta.
with client.responses.stream(model="auto", input="Di hola en una frase.") as stream:
    for event in stream:
        if event.type == "response.output_text.delta":
            print(event.delta, end="")

POST /v1/messages

Anthropic's Messages API. max_tokens is required by the format, so a request without it is a 400 from the API's own validation; system is a top-level field, not a message; and the key may travel either as x-api-key or as Authorization: Bearer — the gateway accepts both.

curl https://api.enroutia.com/v1/messages \
  -H "x-api-key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "max_tokens": 256,
    "system": "Answer in Spanish.",
    "messages": [{"role": "user", "content": "¿Qué es una API?"}]
  }'

The answer is a message whose content is a list of text blocks, with the usage the Anthropic SDK reads.

{
  "id": "msg_…",
  "type": "message",
  "role": "assistant",
  "model": "auto",
  "content": [{"type": "text", "text": "Una API es …"}],
  "stop_reason": "end_turn",
  "usage": {"input_tokens": 21, "output_tokens": 38}
}

With the Anthropic Python SDK

base_url is the host without /v1: the SDK adds /v1/messages itself. Put your Enroutia key in api_key and the SDK sends it as x-api-key.

from anthropic import Anthropic

client = Anthropic(base_url="https://api.enroutia.com", api_key="YOUR_KEY")

message = client.messages.create(
    model="auto",            # or any catalogue alias, or an alias of yours
    max_tokens=256,          # required by the Messages API
    system="Answer in Spanish.",
    messages=[{"role": "user", "content": "¿Qué es una API?"}],
)
print(message.content[0].text)

# Streaming
with client.messages.stream(
    model="auto", max_tokens=256, messages=[{"role": "user", "content": "Hola"}]
) as stream:
    for text in stream.text_stream:
        print(text, end="")

If a proxy of yours strips x-api-key, pass default_headers with Authorization: Bearer YOUR_KEY instead; the gateway treats the two the same.

What comes back with every answer

Whichever dialect you spoke, the response carries the model that actually answered, the request id to quote if something is wrong, and — through auto — the reason it picked that model. The Tags page explains the reasons.

X-Requested-Model: auto
X-Resolved-Model: mistral-small
X-Request-Id: req_…
X-Auto-Reason: short

Streamed or not, the answer names the model you asked for — on every event, including response.model inside the Responses lifecycle events and message.model in Anthropic's message_start — and never the provider's own model id. X-Resolved-Model is the catalogue model that answered.

Not supported, yet

Both endpoints are translated to chat completions on the way in. What that translation cannot carry is refused rather than silently dropped, so a workflow finds out on the first call and not in production:

  • previous_response_id — there is no server-side conversation state; send the whole conversation as input.
  • Built-in tools — web search, file search, computer use. Your own function tools work as they do on chat completions.
  • Files and file ids — attach text inline instead.
  • The guaranteed-JSON check and retry — on chat completions a malformed JSON answer is retried once for you; on these two endpoints it is not, for now.
  • The answer cache — identical calls are not deduplicated on these endpoints, for now.

See also: Start in five minutes · Guide per tool · Aliases · Error reference