Portfolio 19 · evidence before leaderboard

LLM evaluation design and budget planner

Keep confidence-width estimation, power for paired candidate comparisons, stratified slice coverage, and the full candidate/judge/tool/adjudication budget as separate design decisions.

Decision this answers: How many items and paired observations are needed, which slices remain underpowered, and what will the complete evaluation cost before it starts?

Scenario

Numeric scenario inputs

1.96 corresponds to a two-sided 95% normal approximation; this is not power.

exact-derived[1]

Use 0.5 when no defensible pilot estimate exists; it is conservative for this formula.

model estimate[1]

Desired half-width for one proportion, distinct from minimum detectable effect.

model estimate[2]

Every item is run against every candidate.

measured[3]

Repeated stochastic runs per item and candidate.

measured[4]

Model plus application execution for one candidate run.

measured[4]

Automated judging and external tool meters per run.

measured[4]

The complete default scenario is server rendered. Interactions stay in this browser.

Decision summary

Precision sample

385exact-derived

Items for one binomial proportion at the selected confidence and margin.

[1]

Paired comparison sample

221exact-derived

Within-item difference variance and minimum detectable effect; not an unpaired formula.

[2]

Candidate runs

2310exact-derived

Items × candidates × repetitions.

[4]

Complete evaluation budget

$251.67exact-derived

Candidate, judge, tool, adjudication, and fixed program cost.

[4][3]
Direct-labeled result profile

The tables below are the accessible source of truth; bar lengths never carry identity alone.

Stratified coverage plan

Stratified coverage plan
CandidatePopulation weightAllocated itemsMinimum visible floorDecision statusEvidence
Routine requests70%270100Precision-drivenexact-derived[1]
Safety-sensitive20%77100Raise to floorlower bound[3]
Rare edge cases10%38100Raise to floorlower bound[3]

Budget reconciliation

Budget reconciliation
CandidateUnitVolumeDefault subtotalEvidenceEvidence
Candidate executionUSD/run2310$41.58Measured metermeasured[4]
Automated judge + toolsUSD/run2310$20.79Measured metermeasured[4]
Human adjudication + fixedUSD/item + USD46 items$189.30Planning estimatemodel estimate[3]

Official surface versus this site

This offline workbench

NIST supplies the statistical planning foundation. This page adds paired LLM-candidate comparisons, visible slice floors, repetitions, model and judge calls, tools, human adjudication, and fixed program cost.

Sources and methodology

Every result-affecting reference is visible here without JavaScript and is retained in the JSON export.

[1] Sample sizes required
Authority
NIST
Native identifier
NIST/SEMATECH e-Handbook 7.2.4.2
Unit
observations
Evidence
exact-derived
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
The binomial planning equation is evaluated and rounded up; 0.5 is the conservative no-pilot case.

Open the authoritative source

[2] Confidence intervals for proportions
Authority
NIST
Native identifier
NIST/SEMATECH e-Handbook 7.2.4.1
Unit
proportion interval
Evidence
exact-derived
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
Confidence-width terminology is kept distinct from power and minimum detectable effect.

Open the authoritative source

[3] Artificial Intelligence Risk Management Framework
Authority
NIST
Native identifier
NIST AI 100-1
Unit
evaluation governance
Evidence
exact-derived
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
Risk-relevant slices and human adjudication are explicit planning lines, not statistical claims from the framework.

Open the authoritative source

[4] Evaluation best practices
Authority
OpenAI
Native identifier
evaluation best practices
Unit
evaluation runs
Evidence
measured
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
Candidate, judge, tool, and repetition meters are supplied locally; no evaluation data is uploaded.

Open the authoritative source

Limits and non-claims

  • The normal approximation is a planning tool; analysis may need an exact or score interval for small or extreme samples.
  • A minimum slice floor improves visibility but changes total sample size and must remain explicit.
  • Judge agreement and human adjudication quality are evaluation outputs, not assumed constants.