How noisy are LLM evals in CI?
You change a prompt, open a pull request, and your eval suite goes from 48 passing to 46. Did your change break two evals, or would they have failed anyway? We measured how often passing evals fail on the next run in the CI of public repositories, and how big a break a suite can see through that noise.
Passing evals fail by chance about 1 time in 40
We took the eval results each repository uploads from GitHub Actions on its default branch, lined up consecutive runs, and counted the evals that passed in one run and failed in the next. Failures that never reached a model (a missing API key, a rate limit, a timeout) were left out.
| Repository | Format | Evals per run | Runs | Passing evals that failed on the next run |
|---|---|---|---|---|
| JRichlen/agent-plugins | promptfoo | about 12 | 33 | 2 of 389 (0.5%) |
| PrefectHQ/prefect-mcp-server | pytest, JUnit XML | 21 | 27 | 0 of 508 (0%) |
| abpai/skills | JUnit XML | 28 | 12 | 2 of 272 (0.7%) |
| launchdarkly/ai-tooling | promptfoo | 59 | 13 | 18 of 651 (2.8%) |
| evloghq/evlog | JUnit XML | about 17 | 53 | 46 of 870 (5.3%) |
Across the five, 68 of 2,690 passing evals failed on the next run: 2.5%. Commits landed on the default branch between runs, so a few of those failures may be real changes. These rates are an upper bound on chance alone.
A larger suite shows chance on its own. onyx-dot-app/agent-wiki runs about 2,268 evals every night. Over 56 nights, a median of 73 evals that passed one night failed the next, and the count never reached zero. On the 22 nights when the code hadn't changed at all, the median was 73.5. Over the whole period, 541 of the 2,268 evals both passed and failed at least once.
At 2.5 to 3%, a suite of 50 evals loses one or two between any two runs, and a suite of 300 loses eight or nine. A pull request that changes nothing will often show "regressions".
Two suites we left out, and why
Two more repositories were measured and left out of the table, because their numbers measured something other than the model.
- 0% that measured nothing. The only eval in TimMoyence/Innov-mind-museum's results passed on every run. The endpoint it called returned HTTP 401 "Token required", and the scorer accepted the error message as an answer.
- 45%. chmonitor/chmonitor's evals call its production endpoint, behind Cloudflare. Many of its outputs were Cloudflare challenge pages ("Just a moment...") and "Worker exceeded resource limits" errors, not answers.
This is common. On PrefectHQ/prefect-mcp-server#200, 14 of 20 evals failed on the first push. All 14 were the same 404: the judge model, claude-opus-4-1-20250805, had been retired. The pull request broke nothing, and the next push swapped the judge and all 20 passed. (Case study)
Before you read anything into a pass rate, check that each result, passing or failing, came from the model and not from an error page.
How big a break can your suite see?
Say each passing eval in a suite fails by chance with probability p on any run. Then a run with no real change loses about n × p evals. A pull request's failures stand out only when there are more of them than chance would produce in 19 pull requests out of 20. Here is how many evals a change has to break before that happens:
| Suite | Fail by chance per run | Beyond chance at | A break seen half the time | Seen 80% of the time | Chance of seeing a 2-eval break |
|---|---|---|---|---|---|
| 20 evals, 2% noise | 0.4 | 3 failures | 3 evals | 3 evals | 30% |
| 50 evals, 3% noise | 1.5 | 5 failures | 4 evals | 5 evals | 17% |
| 100 evals, 3% noise | 3 | 7 failures | 4 evals | 6 evals | 17% |
| 300 evals, 5% noise | 15 | 22 failures | 8 evals | 11 evals | 11% |
| 2,268 evals, 3.2% noise | 73 | 88 failures | 16 evals | 23 evals | 6% |
A bigger suite doesn't see small breaks better when you count failures across the whole suite. Its chance failures grow with it, and a few real ones get lost among them. The 2,268-eval suite above can lose 15 evals to a real regression, and more often than not it looks like any other night.
What it can see is a break in an eval that never fails. On HolmesGPT/holmesgpt#1452, a refactor stopped sending the agent its runbook catalog, and one eval in a suite of 8 failed: 162_get_runbooks, which had passed all 230 of its runs since early January. The repository's own eval check stayed green. A single failure of an eval like that isn't chance. (Case study)
What helps
Measure your own noise. Rerun your evals on the unchanged default branch, or keep the results of every default-branch run, and count how often passing evals fail. Every other decision depends on that number, and it differs a lot between suites, as the table shows.
Compare each eval with its own history. Don't compare a pass rate against a fixed threshold. Pair the pull request's results with the same evals on the base branch, and judge a failure by how often that eval failed before. A failure in an eval that never fails means more than three failures in evals that fail every week.
Keep errors apart from answers. A missing key, a rate limit, a retired model or an error page isn't a result. Count it as "couldn't run", not as a pass or a failure.
Gate on steady evals, and track the rest. Evals tied to a specific behavior, which pass on every run, can block a merge on a single failure. Broad quality sweeps are better watched as a trend than used as a gate.
Repeat the evals that matter. Running an eval 3 to 10 times and judging the share that pass cuts its noise. On HumanEval, Meta's researchers found that pairing and averaging cut the smallest difference that reaches significance from 12 points to 2 to 4. If you then test every eval on its own, correct for how many you tested, or some will look broken by chance.
Rerun before you believe one failure. An eval that fails by chance 3% of the time fails twice in a row about 1 time in 1,000. On launchdarkly/ai-tooling#189, the repository's own eval report showed configs-create down 25 points against a baseline recorded in May, and its gate failed on onboarding, 2 of 4. The test behind the drop had also failed on 7 of the 12 nights on main before the pull request, and rerunning the same commit passed the gate. (Case study)
How evalship does this
evalship reads the eval results your GitHub Actions already upload and compares every pull request with your default branch, using the method above. It measures each suite's noise over its last 30 runs on the default branch, marks failures that chance explains as "within chance", sets aside results that couldn't run, and comments on the pull request with what's left. How PR analysis works has the details. It's free for public repositories, and setup is one prompt you paste into your coding agent: get started.
Further reading
- Evan Miller, Adding Error Bars to Evals (Anthropic, 2024): standard errors, clustering, paired comparisons and power analysis for evals.
- Wang and others, Measuring all the noises of LLM Evals (Meta, 2025): why averaging repeated runs helps paired comparisons.
- OpenRouter, How to Gate Pull Requests on LLM Evals in CI: measuring the noise floor before setting a threshold.