Skip to content

Messages (Anthropic)

POST /v1/messages

Accepts requests in the Anthropic Messages API format. Requests are translated to the internal OpenAI-compatible format, processed, and the response is translated back to the Anthropic format before returning.

This lets you use the Anthropic SDK with ai& without changing your code.

from anthropic import Anthropic
client = Anthropic(
api_key="sk-your-api-key",
base_url="https://api.aiand.com/v1",
)
message = client.messages.create(
model="openai/gpt-oss-120b",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello!"}],
)
print(message.content[0].text)
  1. Your request arrives in Anthropic Messages format
  2. ai& translates it to OpenAI format and runs inference
  3. The response is translated back to Anthropic format

Streaming is supported. Authentication uses the same Authorization: Bearer header as all other endpoints.

Standard Anthropic Messages API parameters are accepted. Parameters with no equivalent in the underlying model are silently ignored, with the exceptions below.

  • tools — client tools (type custom, or no type) are translated to function tools. Anthropic’s server-side tools (web_search_*, web_fetch_*, code_execution_*, tool_search_*, …) run inside Anthropic’s API and cannot be served here. Anthropic-defined client tools (bash_*, text_editor_*, computer_*, …) carry no input_schema for a non-Anthropic model to follow; send them as custom tools instead. A request that includes any Anthropic-defined tool type is rejected with 400 invalid_request_error naming the tool.
  • tool_choice — auto, any, none and tool are mapped; disable_parallel_tool_use: true becomes parallel_tool_calls: false.
  • tool_result images — on vision-capable models, images inside a tool result are handed to the model in a user message that follows the turn’s tool results, each labelled with its tool_use_id; on text-only models the tool text carries an [Image omitted…] placeholder.

Errors on this endpoint use Anthropic’s envelope, {"type":"error","error":{"type":"…","message":"…"},"request_id":"…"}, with Anthropic’s type names (invalid_request_error, authentication_error, permission_error, not_found_error, rate_limit_error, api_error, …). A request that exceeds the model’s context window is a 400 whose message begins prompt is too long and carries the capability_rejected: prompt_too_long token, which is what Claude Code’s automatic compaction looks for.

A streaming response that stays idle for 15 seconds receives a keep-alive: an SSE comment line before message_start, a ping event after it. Clients that count bytes to detect a stalled stream (Claude Code aborts after 300 seconds of silence) keep the connection through long reasoning spans.

The thinking block does not set the reasoning level — neither budget_tokens nor type maps onto our effort scale, so depth is left to the model’s own default. To control the level, send reasoning_effort with one of the model’s published reasoning_efforts, or Anthropic’s own output_config: { "effort": "…" }, which Claude Code sends for /effort; an explicit reasoning_effort wins when both are present. Either is validated exactly as on /v1/chat/completions: an unsupported value is a 400 naming the field you sent and the values the model takes, and the applied level comes back in X-Reasoning-Effort.

What thinking does control is whether that reasoning is returned to you. Send thinking: { "type": "enabled" } — or "adaptive", which is accepted identically — and the reasoning comes back as a thinking content block ahead of the text block, streaming as thinking_delta events. Omit it and you get the answer alone.

The thinking blocks we return carry no signature — nothing upstream signs them, so one replayed to Anthropic’s own API is rejected. Echoing one back to ai& is safe: a thinking block on a request is dropped before it reaches the model.

Some reasoning models put their whole answer inside the reasoning and leave the message content empty. Those turns come back as an ordinary text block whether or not you asked for thinking, so a successful response always carries at least one content block.

Image content blocks are supported on vision-capable models. All three Anthropic source types work — base64, url, and file (referencing a file_id from POST /v1/files). See Anthropic’s content blocks docs for the wire format. Requests against models without the vision capability are rejected with 400 invalid_request_error before inference.