Documentation

A8 (Touchstone) is a general eval model: state any criterion and it returns a verdict. It answers the OpenAI and Anthropic APIs, so the SDK you already use calls it.

The verdict is scored, so the same input always returns the same verdict.

§1Quickstart

A first verdict from A8, in three steps. The full request and response contract is on the A8 reference.

§1.1Get a key

Issue one under the console → API keys. It needs the models:infer permission; teaching with expected needs models:train. See API keys.

§1.2Point a client at the base URL

Any client that takes a custom base URL works: the OpenAI and Anthropic SDKs, LangChain, LlamaIndex, the Vercel AI SDK, Pydantic AI and LiteLLM among them. With the OpenAI SDK:

# pip install openai from openai import OpenAI client = OpenAI( base_url="https://api.u22a8.ai/eval/v1", api_key="<your key>", )

§1.3State a criterion, send a subject

The system message is the criterion: what is being judged. The last user message is the subject: the thing under judgment.

resp = client.chat.completions.create( model="a8", messages=[ {"role": "system", "content": "Is the reply free of hedging?"}, {"role": "user", "content": "It may possibly work, but I could be wrong."}, ], ) print(resp.choices[0].message.content) # the verdict, or None on an abstention print(resp.model) # the horizon that answered

A8 answers where its evidence clears the accuracy the request sets, 90 by default, and abstains where it does not. On a criterion it has not learned, send the verdict you wanted as expected and it learns it.

The same request through the Anthropic SDK, the Responses API or the native route is in the reference.

§2Where to go next

A8 TouchstoneWhat A8 is, the shape of a call, and where to go next.Read first →