Parts of this page show illustrative data: invented, and marked where it stands, until the lab has its own.
Which way of choosing a model keeps quality at the lowest cost?
No finding: this is a sample experiment, and its figures measure nothing. A real experiment states its finding here, in one sentence, with the number in it.
Revision 2 · Sep 24 · The judge's agreement recomputed; the conclusion unchanged. What changed ↓
- State
- Finished
- Writing
- Draft
- Started
- 21d ago
- Finished
- 6d ago
Hypothesis
A cheap way of choosing a model per question can send most questions to a small model and keep at least 95% of the reference model's judged quality, at no more than half its cost.
- Baseline
- Always answer with Model E
- Accept if
- quality ≥ 95% of the reference and cost ≤ 50%
- Registered
- Sep 4, before the first run
Setup
- Dataset
- sample-qa v1 · 4,000 invented questions, 2,800 to fit the policies and 1,200 held out
- Models
- Five, named A to E; Model E is the reference
- Policies
- Rules; a cascade, Model A first and Model E when unsure; an LLM router, where Model B decides; kNN over embeddings of past questions
- Judge
- Pairwise against the reference answer, 3 samples per question
- Cost
- Measured tokens at the list prices of the day, the policies' own calls included
What the figure shows
Each point is a model that always answers, or a policy that chooses one per question: its judged quality against what 1,000 answers cost. The shaded corner passes both bars, and the line joins the cheapest point at each level of quality.
- Source
- run r-0412
- Dataset
- sample-qa@v1 · 7f3c21e
- Judge and harness
- judge-prompt v5 · harness 1.8.2
- Invented for the design
- Sep 25 · rev 2 Sep 24
| Policy | Quality, % of ref | Cost / 1k | vs reference | p95 latency | Router's own cost |
|---|---|---|---|---|---|
| Always Model E (reference) | 100.0 | $12.00 | — | 6.1 s | — |
| kNN router, threshold 95% (chosen) | 96.1 ± 0.9 | $3.72 | −69% | 3.1 s | $0.02 |
| LLM router, Model B decides | 95.5 ± 0.8 | $6.96 | −42% | 4.9 s | $3.36 |
| Cascade, Model A then Model E | 95.0 ± 0.7 | $4.70 | −61% | 5.8 s | $0.31 |
| Rules | 91.0 ± 1.1 | $5.10 | −58% | 3.4 s | $0.00 |
| Always Model A | 84.0 ± 1.4 | $2.25 | −81% | 2.2 s | — |
Failures
Illustrative data- F1
A contract clause to translate went to Model A and scored 2 of 5.
Few translations were in the fitting set, so the question's nearest neighbours were emails.
- F2
A SQL query with a window function went to Model B. It ran and returned wrong totals.
A query that runs is not a query that is right: the judge now runs SQL against a fixture.
- F3
The judge preferred Model D's answers more often than the human labels did.
The judge and Model D come from one family; Model D is marked judge-sensitive in the table.
Reproduce
Illustrative datagit checkout <the run's commit>
lab run sample --config configs/sample.yaml --dataset sample-qa@v1
Raw results, CSVJudge prompt v5Human labels
A sample decides nothing
A real experiment ends here with the decision that followed: what changed in the lab because of its result, and the next question it leads to.
Revisions
- rev 2
- Sep 24
- 150 human labels added; the judge's agreement recomputed, κ 0.64 → 0.71.
- rev 1
- Sep 19
- The first version.