Keep confidence-width estimation, power for paired candidate comparisons, stratified slice coverage, and the full candidate/judge/tool/adjudication budget as separate design decisions.
Decision this answers: How many items and paired observations are needed, which slices remain underpowered, and what will the complete evaluation cost before it starts?
Scenario
Decision summary
Precision sample
385exact-derived
Items for one binomial proportion at the selected confidence and margin.
Use the official surface to confirm native prices, counting rules, limits, and product eligibility.
This offline workbench
NIST supplies the statistical planning foundation. This page adds paired LLM-candidate comparisons, visible slice floors, repetitions, model and judge calls, tools, human adjudication, and fixed program cost.
Sources and methodology
Every result-affecting reference is visible here without JavaScript and is retained in the JSON export.
[1] Sample sizes required
Authority
NIST
Native identifier
NIST/SEMATECH e-Handbook 7.2.4.2
Unit
observations
Evidence
exact-derived
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
The binomial planning equation is evaluated and rounded up; 0.5 is the conservative no-pilot case.