Parts of this page show illustrative data: invented, and marked where it stands, until the lab has its own.

SampleInvented by Claude to show the page. Nothing on it was measured, judged or reviewed.Illustrative data

Which way of choosing a model keeps quality at the lowest cost?

No finding: this is a sample experiment, and its figures measure nothing. A real experiment states its finding here, in one sentence, with the number in it.

Revision 2 · Sep 24 · The judge's agreement recomputed; the conclusion unchanged. What changed ↓

State
Finished
Writing
Draft
Started
21d ago
Finished
6d ago

Question

Hypothesis

A cheap way of choosing a model per question can send most questions to a small model and keep at least 95% of the reference model's judged quality, at no more than half its cost.

Baseline
Always answer with Model E
Accept if
quality ≥ 95% of the reference and cost ≤ 50%
Registered
Sep 4, before the first run

Setup

Dataset
sample-qa v1 · 4,000 invented questions, 2,800 to fit the policies and 1,200 held out
Models
Five, named A to E; Model E is the reference
Policies
Rules; a cascade, Model A first and Model E when unsure; an LLM router, where Model B decides; kNN over embeddings of past questions
Judge
Pairwise against the reference answer, 3 samples per question
Cost
Measured tokens at the list prices of the day, the policies' own calls included

Evidence

What the figure shows

Each point is a model that always answers, or a policy that chooses one per question: its judged quality against what 1,000 answers cost. The shaded corner passes both bars, and the line joins the cheapest point at each level of quality.

Judged quality vs cost per 1,000 answers Quality as % of the reference model. Cost on a log scale, cheaper is left. Bars show 95% intervals.
Illustrative data
80%85%90%95%100% $1$2$5$10$20 Passes both bars 50% of reference cost Threshold 95% Model A Model B Model C Model D Model E · reference Rules Cascade LLM router 96.1% · $3.72 kNN router Cost per 1,000 answers, USD 80859095100 $1$2$5$10$20
○ single model◇ policy● chosen
Source
run r-0412
Dataset
sample-qa@v1 · 7f3c21e
Judge and harness
judge-prompt v5 · harness 1.8.2
Invented for the design
Sep 25 · rev 2 Sep 24
PolicyQuality, % of refCost / 1kvs referencep95 latencyRouter's own cost
Always Model E (reference) 100.0 $12.00 — 6.1 s —
kNN router, threshold 95% (chosen) 96.1 ± 0.9 $3.72 −69% 3.1 s $0.02
LLM router, Model B decides 95.5 ± 0.8 $6.96 −42% 4.9 s $3.36
Cascade, Model A then Model E 95.0 ± 0.7 $4.70 −61% 5.8 s $0.31
Rules 91.0 ± 1.1 $5.10 −58% 3.4 s $0.00
Always Model A 84.0 ± 1.4 $2.25 −81% 2.2 s —

Failures

Illustrative data
  1. F1

    A contract clause to translate went to Model A and scored 2 of 5.

    Few translations were in the fitting set, so the question's nearest neighbours were emails.

  2. F2

    A SQL query with a window function went to Model B. It ran and returned wrong totals.

    A query that runs is not a query that is right: the judge now runs SQL against a fixture.

  3. F3

    The judge preferred Model D's answers more often than the human labels did.

    The judge and Model D come from one family; Model D is marked judge-sensitive in the table.

Nothing excluded. 14 timeouts were retried and kept; 3 refusals count as failed answers.

Reproduce

Illustrative data
git checkout <the run's commit>
lab run sample --config configs/sample.yaml --dataset sample-qa@v1

Raw results, CSVJudge prompt v5Human labels

A sample has no files: its links are off.

Decision

A sample decides nothing

A real experiment ends here with the decision that followed: what changed in the lab because of its result, and the next question it leads to.

Revisions

rev 2
Sep 24
150 human labels added; the judge's agreement recomputed, κ 0.64 → 0.71.
rev 1
Sep 19
The first version.