A two-point gap on 40 examples is almost never a real improvement — it sits inside the margin of error you'd get from random sampling alone. At that sample size, confidence intervals on accuracy routinely span 10+ percentage points, so the "win" is statistically indistinguishable from noise until you add more examples or repeat runs.

Quick answer: A score difference only means something if it's larger than the confidence interval around it. Detecting a genuine 2-point difference on a binary pass/fail eval usually takes several hundred to several thousand examples, not 40. Below that, run each variant more than once and compare with a paired test instead of trusting a single number.

Why a Two-Point Gap on 40 Examples Doesn't Mean What You Think

Any score computed from a small sample carries sampling error — the same 40 examples, redrawn from the same underlying task distribution, would likely score several points differently even with an identical model. On a 40-item binary eval scoring near 80% accuracy, the standard margin of error is roughly ±12-13 percentage points at conventional 95% confidence, which swallows almost any two-point gap.

The math behind this is simple. For a proportion like accuracy, the standard error is sqrt(p × (1-p) / n). At p = 0.8 and n = 40, that's sqrt(0.16 / 40) ≈ 0.063, or about 6.3 percentage points of standard error — and a 95% confidence interval is roughly ±1.96 standard errors wide.

Small sample sizes make this worse in a predictable way: because the margin of error shrinks with the square root of n, cutting it in half requires roughly four times as many examples, not two. That's why a 40-example eval and a 400-example eval aren't "10x more precise" — they're only about 3x more precise.

Examples (n)Approx. margin of error at ~80% accuracy (95% CI)What a 2-point gap tells you
20±17-18 pointsMeaningless — pure noise
40±12-13 pointsStill noise
100±7-8 pointsProbably still noise
200±5-6 pointsBorderline
500±3-4 pointsStarting to be detectable
1,000±2-3 pointsA 2-point gap becomes plausible

These figures assume ~80% baseline accuracy; your exact margin shifts with your true score, but the shape holds — steep gains from the first few hundred examples, then diminishing returns. The normal approximation also breaks down near 0% or 100% accuracy, exactly where many production evals sit.

Statistician Edwin Wilson proposed a fix back in 1927, now known as the Wilson score interval — the version most modern statistics libraries recommend over the naive formula at small or extreme sample sizes.

If you're still assembling your evaluation set, this is the math behind why our guide to building your first golden dataset of 50 examples treats that number as a floor, not a finish line — a set that size can catch obvious regressions, but it can't adjudicate a close race between two decent variants.

What a Confidence Interval Actually Tells You

A confidence interval is a range around your measured score that would contain the true score most of the time if you repeated the measurement — it describes the reliability of the method, not a guarantee about this one number. Two variants whose confidence intervals overlap substantially cannot be confidently ranked against each other, no matter which one scored higher on this particular run.

This reframes the entire question. Instead of asking "did Model B beat Model A," the right question is "do Model A's and Model B's plausible-score ranges overlap." If they do, you don't have a winner — you have two candidates that need more data or a better comparison design before anyone can call it.

Three common ways to compute that interval, in order of how often teams reach for them:

MethodBest forStrengthWatch-out
Normal (Wald) approximationQuick back-of-envelope estimates, large nTrivial to compute by handDistorted near 0%/100% or with small n
Wilson score intervalBinary pass/fail evals, any nStays accurate at small samples and extreme scoresSlightly less familiar formula
Bootstrap resamplingContinuous or judge-scored metrics (1-5 scales, weighted rubrics)No distributional assumptions; works for any metricNeeds the raw per-example scores, not just the aggregate

Bootstrap resampling — repeatedly resampling your existing examples with replacement and recomputing the score thousands of times to build an empirical distribution — was formalized by statistician Bradley Efron in 1979 and has since become the default tool in NLP research for comparing systems when scores don't come from a clean binomial process, such as pass@1 composites or rubric-graded outputs. It's also the practical answer when your eval report only stores a single aggregate score: to compute any of these intervals, you need the per-example pass/fail or score data, not just the headline number.

This is where judge quality starts to matter as much as sample size. If your scoring itself is inconsistent — the same output judged differently on different passes — your interval is wide for a second reason layered on top of sampling error. Our piece on when to trust an LLM-as-judge setup covers how to check that the judge itself isn't the noisiest part of the pipeline before you even get to the sample-size question.

The Rule of Thumb: How Many Examples to Detect a Given Difference

As a working rule of thumb, examples needed per variant ≈ 16 × p × (1-p) / d², where p is your baseline accuracy and d is the difference you want to reliably detect, both as decimals. This approximation targets the conventional 80% statistical power at a 0.05 significance threshold that Jacob Cohen's foundational 1988 work on power analysis established as the default bar in behavioral and social-science research, and it's since been adopted broadly across applied ML evaluation.

Plugging in a ~80% baseline accuracy, here's what that formula implies for different gap sizes:

Effect size you want to detectApprox. examples needed per variant (80% power)
2 points~6,000+
5 points~1,000
10 points~250
15 points~115
20 points~65

Two things move these numbers. First, baseline accuracy matters: scores near 50% have more variance and need more examples, while scores near the extremes need fewer. Second, a paired comparison — running the same examples through both variants rather than independent samples — reduces the effective sample size needed, often substantially, because it cancels out item-level difficulty differences instead of adding them as noise.

This isn't just theoretical. Researchers Dror, Baumer, Shlomov, and Reichart surveyed NLP publications in a widely cited 2018 paper and found that most skipped formal significance testing entirely. A follow-up 2020 study by Card, Henderson, and colleagues, "With Little Power Comes Great Responsibility," found many published NLP experiments were underpowered to reliably detect the effect sizes their authors claimed. Small evals with confident headlines are a well-documented pattern, not a hypothetical.

The broader discipline these numbers sit inside is covered in our complete guide to AI evals; if you're deciding whether a given comparison even belongs in an offline test set versus something you should be watching in production traffic instead, our piece on offline vs. online AI evals is the right next read — sample-size math applies differently once you're pulling from a live, growing stream of usage data rather than a fixed golden set.

The Anti-Pattern: Celebrating a One-Run Win

The most common mistake is running each variant exactly once, seeing Model B beat Model A by a few points, and shipping — without checking whether that gap would reproduce. LLM outputs are stochastic by design (sampling temperature, decoding randomness, judge variance), so a single run's score is itself a noisy draw, and simply re-running the same model can shift the score almost as much as a genuine improvement would.

You're likely in this anti-pattern if any of the following are true:

  • Each variant was run exactly once, with no repeat runs or seed variation
  • The eval report shows a single number per variant with no error bars or confidence range
  • The comparison used different example orderings, prompts, or subsets between variants
  • An LLM judge scored the outputs without any check on judge-to-judge or run-to-run agreement
  • The "win" was declared before anyone asked how big the gap would need to be to matter

Run-to-run variance comes from more sources than people expect: nonzero temperature, few-shot example ordering, and — if you're using an automated grader — judge inconsistency on borderline outputs. That last one compounds the sampling problem instead of sitting separately from it, which is why validating your grading approach (see the llm-as-judge piece above) has to happen before a significance calculation means anything.

There's also a human factor behind this anti-pattern that's worth naming honestly: the pressure to declare a win usually comes from somewhere outside the numbers — a roadmap deadline, a stakeholder update, a launch already scheduled. The same discipline that stops a team from mistaking a loud, specific feature request for the actual underlying need — the core move in jobs-to-be-done thinking — applies here too: separate the signal you actually have from the story you'd like to tell about it.

The Fix: Repeat Runs, Error Bars, and Paired Comparisons

The fix is procedural, not statistical wizardry: run each variant multiple times, report a mean with a confidence interval instead of one number, and compare variants on the same examples using a paired test so item-level difficulty cancels out rather than adding noise. None of this requires a statistics background beyond what's above — it requires discipline about what you report.

In practice, that looks like:

  1. Run each variant 3-5+ times with different seeds or sampling draws whenever generation isn't fully deterministic.
  2. Report score ± interval, never a bare number — "82% (95% CI: 76-88%)" instead of "82%."
  3. Use a paired design — same examples, same order, across variants — since it's more statistically powerful than independent samples for the same total number of examples run.
  4. Match the test to the metric type: for binary pass/fail comparisons on paired data, statistician Quinn McNemar's 1947 test for paired nominal outcomes is the standard tool; for continuous or judge-scored metrics, paired bootstrap resampling (the same Efron-style method from the confidence-interval section) is the more common modern choice.
  5. Decide the bar before you look at the score. Pre-registering what counts as a real win — not just "higher" but "higher by at least X, with a confidence interval that excludes zero" — is the same discipline behind well-designed numeric thresholds for evals: set the bar first, so the number can't retroactively justify itself.

Steps 1 and 3 do most of the work. Multiple runs expose whether a variant is reliably better or just had a lucky draw; pairing on the same examples means both variants are judged against identical item difficulty, so the comparison isn't diluted by "Model A's sample happened to include harder questions."

Where This Fits Into Your Eval Practice

None of this replaces judgment — it constrains it. An eval plan with numeric pass/fail thresholds is only trustworthy if those thresholds account for sample size and run-to-run variance, not just a single directional arrow pointing up.

This is exactly the gap Prodinja's eval-planning step is built to sit inside. It's designed to walk you through setting numeric thresholds for what counts as pass or fail before you run a single comparison, so the bar exists before the data does.

The statistics above are what should sit behind that number — a threshold sized to a difference your sample can actually detect, not one picked because it matched what a single run happened to show. Used this way, it's a structure for your own judgment, not a replacement for the repeat runs and paired comparisons — you still have to bring those.

Key Takeaways

  • A score difference is only real if it's bigger than the confidence interval around it — overlapping intervals mean you don't have a winner yet, regardless of which number is higher.
  • Small evals have wide margins of error: a 40-example binary eval can easily carry a ±12-13 point margin of error at 80% baseline accuracy, swallowing most small "wins."
  • Detecting a genuine 2-point difference typically takes several hundred to several thousand examples, while a 15-20 point gap is detectable with a few hundred — effect size and sample size trade off directly.
  • A single run is a noisy draw, not ground truth — sampling temperature, prompt ordering, and judge inconsistency all add variance that one run can't distinguish from a real improvement.
  • Paired comparisons beat independent samples — running both variants on the same examples cancels out item-difficulty noise instead of adding it, and lets you use tests like McNemar's designed for exactly this setup.
  • Decide your threshold before you see the score — a bar chosen after the fact isn't a threshold, it's a rationalization.
  • Report a range, not a point estimate — "82% (76-88% CI)" is more honest, and more useful, than "82%" on its own.

Frequently Asked Questions

How many examples do I need for an LLM eval result to be statistically significant?

It depends on the size of the difference you're trying to detect: roughly a few hundred examples for a 10-15 point gap, and several thousand for a 2-point gap, using the standard two-proportion power approximation at 80% power. There's no single universal number — smaller true differences always require larger samples to detect reliably.

Can I trust a single eval run to compare two models or prompts?

No — a single run is one noisy draw from a distribution of possible outcomes, not a definitive measurement. Because LLM generation involves sampling and judge scoring involves some inconsistency, the same model run twice can produce different scores; at least 3-5 repeat runs with a reported range are needed before treating a gap as real.

What's the difference between statistical significance and practical significance in evals?

Statistical significance means a gap is unlikely to be pure chance given your sample size; practical significance means the gap is large enough to matter for your product or users. A large, well-powered eval can detect a statistically real but practically trivial half-point difference — significance alone doesn't tell you whether to ship.

Should I use a paired test or an independent-samples test to compare two model variants?

Use a paired test — such as McNemar's test for binary outcomes or paired bootstrap resampling for continuous scores — whenever both variants are run on the identical set of examples. Pairing removes item-difficulty variance from the comparison, which typically makes real differences easier to detect with the same number of examples.

Why did my eval score change when I reran the exact same model?

Nonzero sampling temperature, prompt or few-shot ordering, and inconsistency in an LLM judge's grading all introduce run-to-run variance even with an unchanged model. This is expected, not a bug — it's the reason a single run's score should never be treated as the model's "true" score without a confidence interval or repeat runs around it.