A8 reference

The wire contract for A8: one operation, four interfaces to call it through, the fields that carry meaning, and what comes back. The concepts behind them are on the A8 overview and abstention.

Base URL and authentication

Base URLhttps://api.u22a8.ai/eval/v1
List modelsGET /models
AuthenticationAuthorization: Bearer <key>, or x-api-key: <key>. Inference needs the models:infer permission. A request carrying expected (teaching) needs models:train.
Rate limitNo fixed published limit. Sustained or abusive traffic is throttled at the edge and receives 429.

Table 1. The base URL, and the credential a request carries.

Interfaces

A request has three parts: the criterion (what is being judged), the subject (the thing under judgment), and the shape of the verdict. Each interface carries the same three parts in its own fields, and all four return the same verdict for the same request.

InterfaceRouteCriterionSubjectVerdict shape
OpenAI Chat CompletionsPOST /chat/completionssystem messagesthe last user messageresponse_format
OpenAI ResponsesPOST /responsesinstructionsinputtext.format
Anthropic MessagesPOST /messagessystemthe last user messageoutput_config.format
NativePOST /verdictscriterionsubjectschema

Table 2. The four interfaces, relative to the base URL. An OpenAI SDK takes the base URL as it is. An Anthropic SDK takes https://api.u22a8.ai/eval and adds /v1/messages itself. POST to the base URL is the Chat Completions route.

Use a compatible interface when a client or framework already speaks it. Use the native route when you write the request yourself: the parts are fields of their own, and the verdict comes back as JSON instead of a string.

Compatible in shapeA8 reads the criterion and the subject from the fields in Table 2. A prompt that mixes them in one message, as the built-in graders of some eval frameworks do, is read as one subject with no criterion.

Request

PartRole
modelRequired. a8 for the live model, a pinned horizon, or a8-1-0@reference for the reference model with none of your corrections in it. See horizons.
CriterionRequired, unless every field of the verdict shape describes itself. In Chat Completions and Responses, developer messages count as system, and several are joined.
SubjectRequired. Text. A message may be a string or a list of text parts. A part that is not text is rejected.
Verdict shapeOptional. A JSON Schema object. Without it the verdict is a single label or number.
min_accuracyOptional. The accuracy to answer at: an integer 0 to 99, where 0 asks for no promise. See abstention. Defaults to 90.
expectedOptional. The verdict you wanted, which fine-tunes the model. See fine-tuning.

Table 3. The parts of a request. min_accuracy and expected are top-level fields of the body in every interface. The OpenAI and Anthropic SDKs send them through extra_body.

In the three compatible interfaces, every other field of that API is accepted and ignored: temperature, max_tokens, seed and the rest. A body written for another provider does not break. Three things are rejected with 400: a field that misspells min_accuracy or expected, n above 1, and the json_object response format, which names no fields to answer in. The native route rejects any field it does not define.

Describing the verdict shape

Each property of the schema is one field of the verdict. Its type says what kind of judgment it holds.

PropertyVerdict
boolean, or an enumA label: one of the declared values.
number or integerA score between minimum and maximum, which default to 0 and 100.
Any otherNot judged. A8 writes no text, so a string comes back empty, or null where the schema allows it.

Table 4. How a schema property is read. A $ref into $defs and a nullable anyOf are followed, so a schema generated from a Pydantic or Zod type reads the same as one written by hand.

When a property carries a description, that description is its own criterion and is judged independently. One request can carry several criteria at once. The request’s criterion applies only to properties that don’t describe themselves: here, queue.

{ "type": "object", "properties": { "queue": {"enum": ["billing", "shipping", "account"]}, "urgent": {"type": "boolean", "description": "Does the customer need an answer today?"}, "tone": {"type": "integer", "minimum": 1, "maximum": 5, "description": "How polite is the message?"} } }

A forced tool call is read the same way. Where a request names one function it requires (tool_choice), the function’s parameters are the verdict shape, and the verdict comes back as that tool call.

Giving the subject a context

Where the judgment is relational (grounded in a source, answering a question, matching a reference), the subject is an object: output, plus any of reference, context and question. Each is judged in relation to the others. The same words sent as one formatted string are one opaque subject, and the relation is lost.

context takes a string or a list of them. reference and question take a string.

The native route takes the object as it is. In a compatible interface the subject is a message, so the object travels inside it as a JSON string, with output and at least one of the others. Anything else is read as plain text.

{ "role": "user", "content": "{\"output\": \"The tower was finished in 1889.\", \"context\": [\"Construction ended in March 1889.\"]}" }

The native route

POST /verdicts takes the parts by name and returns the verdict as JSON.

{ "model": "a8", "criterion": "Is the answer grounded in the context?", "subject": { "output": "The tower was finished in 1889.", "context": ["Construction ended in March 1889."] }, "schema": { "type": "object", "properties": {"grounded": {"type": "boolean"}} }, "min_accuracy": 95 }
{ "id": "eval-9f2c1d84ab30e5f7", "object": "verdict", "model": "a8-1-0@20260715t140322z", "verdict": {"grounded": true}, "min_accuracy": 95, "usage": {"input_tokens": 812}, "system_fingerprint": "fit-1a2b3c4d5e6f7890.9e2a1b3c" }

verdict is an object matching the schema, a label or a number where the request carried no schema, and null on an abstention.

Response

A compatible interface answers with that API’s own object. The verdict is where the API puts the model’s output: the message content in Chat Completions, the output_text part in Responses, the text block in Anthropic Messages. This is the Chat Completions form.

{ "id": "eval-9f2c1d84ab30e5f7", "object": "chat.completion", "model": "a8-1-0@20260715t140322z", "choices": [{ "index": 0, "message": {"role": "assistant", "content": "{\"grounded\": true}", "refusal": null}, "finish_reason": "stop" }], "usage": {"prompt_tokens": 812, "completion_tokens": 0, "total_tokens": 812}, "system_fingerprint": "fit-1a2b3c4d5e6f7890.9e2a1b3c", "min_accuracy": 95 }
FieldMeaning
modelThe exact horizon that answered, more specific than what was asked for. Send it back verbatim to reproduce this verdict. See horizons.
usageThe tokens the model read. Nothing writes the verdict, so an interface with an output count reports 0 in it.
system_fingerprintA change detector covering everything that determined the verdict. See reproducibility.
min_accuracyPresent on every answer: the level the verdict is promised at, which is the level asked for or the nearest earned one above it. 0 means no promise was requested. See abstention.
abstentionPresent when the model declined to answer: reason for a single verdict, fields when a schema’s properties declined separately. It names the check the evidence failed. Its wording and its numbers are not part of the contract.
interval / intervalsPresent on a graded verdict answered under a promise: the [low, high] range the promised level guarantees, in the schema’s declared units. Absent when min_accuracy is 0.

Table 5. Response fields, under the same names in every interface. The last three are extensions that standard clients ignore.

An abstention

The request succeeds with 200. Each API has a place for an answer the model declined to give, and A8 uses it, so a typed client reads an abstention without a parsing error.

InterfaceAn abstention is
Chat Completionsmessage.refusal, with content set to null
Responsesa refusal content part
Anthropic Messagesstop_reason: "refusal", with stop_details and no content
Nativeverdict: null

Table 6. Where an abstention appears. The abstention field is present in all four.

Where a schema carries several criteria and only some decline, the response is an ordinary answer. The declined fields are null, and abstention.fields names the check each one failed.

Streaming

With stream: true, the three compatible interfaces answer in that API’s event stream. The verdict arrives whole, in one delta. The native route does not stream.

Examples

# curl, the native route curl -s https://api.u22a8.ai/eval/v1/verdicts \ -H "Authorization: Bearer $U22A8_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "a8", "criterion": "Does this reply meet our support quality standard?", "subject": "Sorry for the delay. Your parcel left the depot this morning." }'
# python, the OpenAI SDK from openai import OpenAI from pydantic import BaseModel client = OpenAI(base_url="https://api.u22a8.ai/eval/v1", api_key="<your key>") class Verdict(BaseModel): meets_standard: bool resp = client.chat.completions.parse( model="a8", messages=[ {"role": "system", "content": "Does this reply meet our support quality standard?"}, {"role": "user", "content": "Sorry for the delay. Your parcel left the depot this morning."}, ], response_format=Verdict, extra_body={"min_accuracy": 95}, ) message = resp.choices[0].message message.parsed # Verdict(meets_standard=True), or None on an abstention message.refusal # None, or the check the evidence failed
# python, the Anthropic SDK from anthropic import Anthropic client = Anthropic(base_url="https://api.u22a8.ai/eval", api_key="<your key>") message = client.messages.parse( model="a8", max_tokens=1024, system="Does this reply meet our support quality standard?", messages=[ {"role": "user", "content": "Sorry for the delay. Your parcel left the depot this morning."}, ], output_format=Verdict, extra_body={"min_accuracy": 95}, ) message.parsed_output # Verdict(meets_standard=True) message.stop_reason # "end_turn", or "refusal" on an abstention

Status codes

StatusCondition
400The body cannot be read as an eval: no subject, no criterion, a min_accuracy outside 0 to 99, or a misspelled min_accuracy or expected. A misspelled parameter is rejected rather than silently replaced by the default. Also returned when model names a version that is not published.
401No credential, or the key is invalid or revoked.
403The key lacks the required permission: models:train for a body carrying expected, models:infer otherwise.
404The pinned horizon names nothing this organization has fitted. See horizons.
429Rate limited. Back off and retry.
503Temporary: the pinned horizon is still fine-tuning (Retry-After: 30), an upstream dependency is briefly unavailable (Retry-After: 5), or the key verifier is unreachable (no Retry-After; use your own backoff). The OpenAI and Anthropic SDKs retry this automatically.

Table 7. Status codes. Every response carries an x-request-id header worth quoting in a support request.

An error body follows the interface it answers. type is the class that API’s clients branch on. code names the exact condition, and param the field to fix. On the Anthropic route the same content sits in Anthropic’s envelope, {"type": "error", "error": {...}, "request_id": ...}.

{ "error": { "message": "unknown parameter: 'min_acuracy' (did you mean 'min_accuracy'?)", "type": "invalid_request_error", "param": "min_acuracy", "code": "malformed_request", "request_id": "req_5b456a869547474cc316c4fc48017e19" } }