Hanzo

Benchmark

Package benchmark is one honest score for any model, on the tests everyone quotes.

Package benchmark is one honest score for any model, on the tests everyone quotes.

Base URLhttps://api.hanzo.ai
Operations6
AuthAuthorization: Bearer $HANZO_API_KEY

benchmark

GET /v1/benchmark/catalog

The canonical public benchmarks this arena runs

Lists the top-14 set every major provider reports — the id, title, axis, item count and upstream source of each — with native marking the ones the standardized harness runs today; the rest are registered and adapter-pending. These ids are the vocabulary the rest of the surface takes: a run names them, and the leaderboard and compare read them from ?benchmark=. The catalog is deployment-wide and identical for every caller — there is no tenant in it.

GET /v1/benchmark/compare

The only sound head-to-head: two models on the items they BOTH answered

Scores model ?a= against model ?b= on one benchmark, paired over the items both arms actually completed. It answers the common-item count, each arm's correct count, the rescues each way (items one got right and the other did not), the net, and a two-sided exact McNemar p over the discordant pairs.

Pairing is what makes it valid. Reading two leaderboard rows against each other compares one model's coverage with another's, so an arm that only ran the easy subset looks better than it is; this endpoint refuses that by construction — items only one arm attempted are dropped before anything is counted. A p of 1 with zero discordant pairs means the arms never disagreed, not that they are identical. Both a and b are required (400 without them); the benchmark defaults to GPQA-Diamond.

GET /v1/benchmark/leaderboard

Per-model scores for one benchmark: what we measured beside what the vendor claims

Answers one row per model for the benchmark named by ?benchmark= (GPQA-Diamond when omitted), carrying measured — the accuracy our own harness got — beside published, the provider's own claim, and gap, the claim minus the measurement. The gap is the point of the arena; provider-reported claims have run materially hot against one standardized harness.

The two planes are NEVER blended, and that is the rule to read the rows by: a model we have measured but no vendor has claimed for shows published null, a model with only a claim shows measured null, and gap exists only where both do. Each row also carries n, the number of items actually attempted — coverage differs between models, so two measured values at different n are not comparable and the compare endpoint is what settles that properly. Rows are ordered by measured accuracy, unmeasured last. Scores are deployment-wide evidence, not per-tenant.

GET /v1/benchmark/presets

The router blends available to compose from

Lists preset router blends — a named set of model arms, the rank they escalate through and the panel width that bounds fan-out — each served by the model layer as enso-<name>. Today it answers exactly one row, the reference blend: a worked example written in models we name, published as an example of the FORM. It is deliberately not the composition of a Hanzo-served tier — the tier name exists to abstract that — so fork it and swap arms by what the leaderboard measures on your own tasks rather than reading it as a disclosure.

POST /v1/benchmark/presets

Compose a router blend from the arms that win your tasks

Validates a blend — name, its arms, the rank they escalate through and the panel fan-out width — and answers 202 with the preset and the enso-<name> it would be served as. It VALIDATES AND ECHOES: the definition is not persisted yet, so a preset accepted here is not one the model layer will resolve. Treat the response as a check on the blend, not a promise to serve it.

Defaults fill the shape rather than refusing it: an omitted rank becomes the arms in declared order and a panel below 1 becomes 1. The one real invariant is that rank may only name arms the blend declares — the same rule the model catalog enforces — and a rank naming anything else is a 422 listing exactly which entries were undeclared. A blend with no name or no arms is a 400.

POST /v1/benchmark/runs

Queue a benchmark run against a catalog model or your own endpoint

Admits a request to run one or more catalog benchmarks against model — a catalog model id — or against endpoint, an endpoint of your own on the chat-completions wire, and answers 202 with what was queued. It ADMITS AND QUEUES ONLY: nothing is executed on this call and no scores come back with it. Results land in the leaderboard as the worker completes them.

Cost is bounded by the store rather than by a quota: attempts are append-only and keyed by (benchmark, item, model), so an (item, model) pair already attempted is skipped instead of re-spent, and re-queuing the same run is close to free. Validation is up front and total — a request with neither model nor endpoint is a 400, one with no benchmarks is a 400, and any benchmark id outside the catalog is a 422 naming exactly which ids were unknown, so a typo never silently queues a partial run.


All Hanzo APIs · Interactive reference

How is this guide?

On this page