Create claims
Records claims for the caller's org: one to correct a number, many to import a leaderboard.
POST /v1/benchmark/claims
| Address | https://api.hanzo.ai/v1/benchmark/claims |
| Method | POST |
| Operation | post_benchmark_claims |
| Auth | Authorization: Bearer $HANZO_API_KEY |
Records claims for the caller's org: one to correct a number, many to import a leaderboard. Any signed-in caller may write; a caller with no verified principal is refused 401.
The org and the author are the verified caller's. Nothing in the body names either, and a body that tries is not read. Claims are private to that org unless visibility is "public". Writing as org "admin" curates the leaderboard, so it takes a SuperAdmin; anyone else acting there is refused 403, and the refusal is audited.
Every row must carry a Source, because a claim without its citation is a number nobody can check, and a benchmark id from /catalog, because an unknown id would sit in the store invisible to every read. A row names its benchmark, model, provider and protocol in at most 128 characters each, cites a source of at most 2048 bytes, and scores a percentage from 0 to 100; a row outside that is rejected by number.
A request carries at most 500 rows in at most 1 MiB (413 past either), and an org writes at most 2000 rows per UTC day (429 past that, and nothing from the request is written).
The trail takes the call's intent, naming every row, BEFORE the first row lands; a trail that cannot take it answers 503 and nothing is written. A second record then names which rows were stored and which failed.
Writes are append-only, so this never destroys the value it replaces. A vendor restating a score leaves both rows on disk, which is how the restating itself becomes visible.
Request
8 fields, body application/json (required).
| Field | In | Type | Required | Description |
|---|---|---|---|---|
data | body | benchmark.publishedClaim[] | — | Data is the claims to record. |
data[].benchmark | body | string | — | Benchmark is the canonical test id the claim is about, from /catalog. |
data[].model | body | string | — | Model is the system the score is claimed for. |
data[].protocol | body | string | — | Protocol records HOW it was scored — provider-reported, agentic, third-party-leaderboard — because a provider card and a third party running its own harness are different kinds of number and must not be blended. |
data[].provider | body | string | — | Provider is who the claim belongs to — the lab or leaderboard whose number this is. |
data[].score | body | number (double) | — | Score is the reported aggregate, as a percentage. |
data[].source | body | string | — | Source is the citation the row was read from. |
visibility | body | string | — | Visibility is who may read every row in Data: "private", the default, keeps them to the caller's org; "public" shows them to anyone. |
Response
| Status | Body | Meaning |
|---|---|---|
200 | benchmark.putClaimsOut | ok |
default | problem-details | refused |
200 body — 4 fields.
| Field | In | Type | Always | Description |
|---|---|---|---|---|
org | body | string | — | Org is the org the rows were recorded under: the caller's own. |
recorded | body | integer (int64) | — | Recorded is how many rows were written. |
rejected | body | string[] | — | Rejected names the rows that were not, and why. |
visibility | body | string | — | Visibility is who may read them. |
Failure carries the platform error shape — see Errors.
Examples
hanzo benchmark claims createimport { Configuration, BenchmarkApi } from 'hanzoai';
const api = new BenchmarkApi(new Configuration({ accessToken: process.env.HANZO_API_KEY }));
const { data } = await api.postBenchmarkClaims({ data: [{"benchmark":"<benchmark>","model":"<model>","protocol":"<protocol>","provider":"<provider>"}], visibility: "<visibility>" });from hanzoai.cloud import ApiClient, Configuration
from hanzoai.cloud.api import BenchmarkApi
client = ApiClient(Configuration(access_token=os.environ["HANZO_API_KEY"]))
result = BenchmarkApi(client).post_benchmark_claims(data=[{"benchmark":"<benchmark>","model":"<model>","protocol":"<protocol>","provider":"<provider>"}], visibility="<visibility>")cfg := hanzoai.NewConfiguration()
cfg.AddDefaultHeader("Authorization", "Bearer "+os.Getenv("HANZO_API_KEY"))
client := hanzoai.NewAPIClient(cfg)
resp, _, err := client.BenchmarkAPI.PostBenchmarkClaims(context.Background()).Execute()
if err != nil {
return err
}use hanzo_client::apis::{configuration::Configuration, benchmark_api};
let mut cfg = Configuration::new();
cfg.bearer_access_token = std::env::var("HANZO_API_KEY").ok();
let result = benchmark_api::post_benchmark_claims(&cfg, Default::default()).await?;import ai.hanzo.cloud.ApiClient;
import ai.hanzo.cloud.api.BenchmarkApi;
ApiClient client = new ApiClient();
client.setBearerToken(System.getenv("HANZO_API_KEY"));
var result = new BenchmarkApi(client).postBenchmarkClaims();curl -X POST https://api.hanzo.ai/v1/benchmark/claims \
-H "Authorization: Bearer $HANZO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"data": [
{
"benchmark": "<benchmark>",
"model": "<model>",
"protocol": "<protocol>",
"provider": "<provider>"
}
],
"visibility": "<visibility>"
}'MCP reaches benchmark through the benchmark tool, which names its 9 operations with its own verbs — this one among them, under a name only MCP declares. describe explains any of them:
curl -X POST https://api.hanzo.ai/v1/mcp \
-H "Content-Type: application/json" \
-d '{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "describe",
"arguments": {
"op": "get_benchmark_catalog"
}
}
}'