A8 reference
The wire contract for A8: one operation, four interfaces to call it through, the fields that carry meaning, and what comes back. The concepts behind them are on the A8 overview and abstention.
Base URL and authentication
| Base URL | https://api.u22a8.ai/eval/v1 |
|---|---|
| List models | GET /models |
| Authentication | Authorization: Bearer <key>, or x-api-key: <key>. Inference needs the models:infer permission. A request carrying expected (teaching) needs models:train. |
| Rate limit | No fixed published limit. Sustained or abusive traffic is throttled at the edge and receives 429. |
Table 1. The base URL, and the credential a request carries.
Interfaces
A request has three parts: the criterion (what is being judged), the subject (the thing under judgment), and the shape of the verdict. Each interface carries the same three parts in its own fields, and all four return the same verdict for the same request.
| Interface | Route | Criterion | Subject | Verdict shape |
|---|---|---|---|---|
| OpenAI Chat Completions | POST /chat/completions | system messages | the last user message | response_format |
| OpenAI Responses | POST /responses | instructions | input | text.format |
| Anthropic Messages | POST /messages | system | the last user message | output_config.format |
| Native | POST /verdicts | criterion | subject | schema |
Table 2. The four interfaces, relative to the base URL. An OpenAI SDK takes the base URL as it is. An Anthropic SDK takes https://api.u22a8.ai/eval and adds /v1/messages itself. POST to the base URL is the Chat Completions route.
Use a compatible interface when a client or framework already speaks it. Use the native route when you write the request yourself: the parts are fields of their own, and the verdict comes back as JSON instead of a string.
Request
| Part | Role |
|---|---|
model | Required. a8 for the live model, a pinned horizon, or a8-1-0@reference for the reference model with none of your corrections in it. See horizons. |
| Criterion | Required, unless every field of the verdict shape describes itself. In Chat Completions and Responses, developer messages count as system, and several are joined. |
| Subject | Required. Text. A message may be a string or a list of text parts. A part that is not text is rejected. |
| Verdict shape | Optional. A JSON Schema object. Without it the verdict is a single label or number. |
min_accuracy | Optional. The accuracy to answer at: an integer 0 to 99, where 0 asks for no promise. See abstention. Defaults to 90. |
expected | Optional. The verdict you wanted, which fine-tunes the model. See fine-tuning. |
Table 3. The parts of a request. min_accuracy and expected are top-level fields of the body in every interface. The OpenAI and Anthropic SDKs send them through extra_body.
In the three compatible interfaces, every other field of that API is accepted and ignored: temperature, max_tokens, seed and the rest. A body written for another provider does not break. Three things are rejected with 400: a field that misspells min_accuracy or expected, n above 1, and the json_object response format, which names no fields to answer in. The native route rejects any field it does not define.
Describing the verdict shape
Each property of the schema is one field of the verdict. Its type says what kind of judgment it holds.
| Property | Verdict |
|---|---|
boolean, or an enum | A label: one of the declared values. |
number or integer | A score between minimum and maximum, which default to 0 and 100. |
| Any other | Not judged. A8 writes no text, so a string comes back empty, or null where the schema allows it. |
Table 4. How a schema property is read. A $ref into $defs and a nullable anyOf are followed, so a schema generated from a Pydantic or Zod type reads the same as one written by hand.
When a property carries a description, that description is its own criterion and is judged independently. One request can carry several criteria at once. The request’s criterion applies only to properties that don’t describe themselves: here, queue.
A forced tool call is read the same way. Where a request names one function it requires (tool_choice), the function’s parameters are the verdict shape, and the verdict comes back as that tool call.
Giving the subject a context
Where the judgment is relational (grounded in a source, answering a question, matching a reference), the subject is an object: output, plus any of reference, context and question. Each is judged in relation to the others. The same words sent as one formatted string are one opaque subject, and the relation is lost.
context takes a string or a list of them. reference and question take a string.
The native route takes the object as it is. In a compatible interface the subject is a message, so the object travels inside it as a JSON string, with output and at least one of the others. Anything else is read as plain text.
The native route
POST /verdicts takes the parts by name and returns the verdict as JSON.
verdict is an object matching the schema, a label or a number where the request carried no schema, and null on an abstention.
Response
A compatible interface answers with that API’s own object. The verdict is where the API puts the model’s output: the message content in Chat Completions, the output_text part in Responses, the text block in Anthropic Messages. This is the Chat Completions form.
| Field | Meaning |
|---|---|
model | The exact horizon that answered, more specific than what was asked for. Send it back verbatim to reproduce this verdict. See horizons. |
usage | The tokens the model read. Nothing writes the verdict, so an interface with an output count reports 0 in it. |
system_fingerprint | A change detector covering everything that determined the verdict. See reproducibility. |
min_accuracy | Present on every answer: the level the verdict is promised at, which is the level asked for or the nearest earned one above it. 0 means no promise was requested. See abstention. |
abstention | Present when the model declined to answer: reason for a single verdict, fields when a schema’s properties declined separately. It names the check the evidence failed. Its wording and its numbers are not part of the contract. |
interval / intervals | Present on a graded verdict answered under a promise: the [low, high] range the promised level guarantees, in the schema’s declared units. Absent when min_accuracy is 0. |
Table 5. Response fields, under the same names in every interface. The last three are extensions that standard clients ignore.
An abstention
The request succeeds with 200. Each API has a place for an answer the model declined to give, and A8 uses it, so a typed client reads an abstention without a parsing error.
| Interface | An abstention is |
|---|---|
| Chat Completions | message.refusal, with content set to null |
| Responses | a refusal content part |
| Anthropic Messages | stop_reason: "refusal", with stop_details and no content |
| Native | verdict: null |
Table 6. Where an abstention appears. The abstention field is present in all four.
Where a schema carries several criteria and only some decline, the response is an ordinary answer. The declined fields are null, and abstention.fields names the check each one failed.
Streaming
With stream: true, the three compatible interfaces answer in that API’s event stream. The verdict arrives whole, in one delta. The native route does not stream.
Examples
Status codes
| Status | Condition |
|---|---|
400 | The body cannot be read as an eval: no subject, no criterion, a min_accuracy outside 0 to 99, or a misspelled min_accuracy or expected. A misspelled parameter is rejected rather than silently replaced by the default. Also returned when model names a version that is not published. |
401 | No credential, or the key is invalid or revoked. |
403 | The key lacks the required permission: models:train for a body carrying expected, models:infer otherwise. |
404 | The pinned horizon names nothing this organization has fitted. See horizons. |
429 | Rate limited. Back off and retry. |
503 | Temporary: the pinned horizon is still fine-tuning (Retry-After: 30), an upstream dependency is briefly unavailable (Retry-After: 5), or the key verifier is unreachable (no Retry-After; use your own backoff). The OpenAI and Anthropic SDKs retry this automatically. |
Table 7. Status codes. Every response carries an x-request-id header worth quoting in a support request.
An error body follows the interface it answers. type is the class that API’s clients branch on. code names the exact condition, and param the field to fix. On the Anthropic route the same content sits in Anthropic’s envelope, {"type": "error", "error": {...}, "request_id": ...}.