What evalship checks

evalship checks your evals twice. Once at setup, the review asks whether your evals can catch a bad change at all. Then on every pull request, evalship compares the evals' results with the base branch's and asks whether the change broke anything, and whether the results can be trusted enough to say so.

Every number in the pull request checks is computed by deterministic code, and so are failure patterns, likely causes and fixed review findings wherever rules can settle them. The model writes eval notes, and names a pattern, a cause or a fix only when the rules can't.

The review, once at setup

Your coding agent reads every scorer, judge and prompt, and reports at most 4 findings, most valuable first. Each one cites a file, a line and an exact quote, and evalship checks every citation against your repository: a finding whose citation doesn't hold is hidden. See The audit.

Check What it catches
Eval bug A flaw in how an eval decides pass or fail: the answer leaks into the prompt, an average hides a failing hard rule, errors or skips count as passes, the scorer sees truncated or wrong input, or correct answers fail.
Untested instruction A rule in a prompt, such as "never promise refunds above $100", that no case checks.
Judge problem An LLM judge you can't trust: it grades its own model's output, is swayed by the answer's claims, or uses an undefined scale.
Untested prompt A prompt or LLM call that no suite exercises.
Evals never run in CI Evals exist, but CI never runs them, or not on pull requests.
Suspiciously easy suite Cases that can't realistically fail, such as checking only for non-empty output, or a threshold of 0.
Judge-only suite Every assertion is an LLM judge, where deterministic checks would catch the same failures cheaper.
Thin or skewed dataset Few cases, near-duplicates, no edge or adversarial cases, or one language while the prompt handles many.
Nondeterminism Temperature above 0 without repeats or tolerances, live network calls, or no seeds.
Cost and speed The most expensive suites, repeated model calls, no caching, or a large judge model where a small one would do.
Mixed or duplicated setup Two frameworks testing the same thing, or dead eval files.

Each finding comes with what to paste into your coding agent to fix it. The review keeps working after setup: a pull request that fixes a finding says so in its comment, and the report marks the finding fixed once it merges.

On every pull request

Did it break anything?

Check What it does
Regressions Finds the evals that passed on the base branch and fail now, and those whose score dropped by more than the tolerance (0.05 by default).
Beyond chance Counts a suite's failures as regressions only when it got worse beyond the noise measured over its last 30 runs on the default branch. The rest are listed as within chance, with a suggestion to rerun them.
Flaky evals Sets aside evals that flipped between passing and failing at least twice in their last 10 runs on the default branch, or whose scores swing widely. They never count as regressions or fail the gate.
Repeats Judges an eval that runs several times on all its runs, with Fisher's exact test: 3 of 3 to 1 of 3 is within chance, 10 of 10 to 6 of 10 is a regression.
Failure patterns When 2 or more new failures fail for the same visible reason, names the pattern and the line in the diff that likely caused it.
New evals that fail Calls out new evals that fail, first: they have no result on the default branch to compare with.

Can the results be trusted?

Check What it does
Evals that couldn't run Leaves out of the comparison evals that never reached the model: a missing API key (pull requests from forks and Dependabot get no secrets), a rate limit, a timeout, a provider or judge error, or an output that is an HTTP error page, even if it passed. When every eval in a suite fails with the same error, it blames the harness, not the answers.
Evals that didn't run Notes a suite that had results on the base branch and has none here. Evals a run skipped on purpose (a tag, a filter, a path-filtered job) are marked not run, not failed.
Scoring changed Warns when the pull request rewrites the line where a suite decides pass or fail, its scorer, or its cases' expected answers, so before and after may not compare.
Model swapped Warns when the pull request changes the model an eval runs or judges with.
Eval workflow changed Warns when the pull request deletes the workflow that runs the evals, removes its pull_request trigger, removes the results upload or its if: always(), or turns a step off with if: false.

Is everything that changed tested?

Check What it does
Untested changes Lists the LLM calls, prompt files and tool descriptions the pull request adds or changes that no eval exercises, with a prompt for your coding agent to add one. See Untested changes.
Not run on pull requests Notes a change that only suites run by hand, on a schedule or by label exercise.
Eval notes The model reads the diff and your review, and flags evals that can't catch the change, expected answers loosened in the same pull request as the prompt they grade, new suites CI never runs, and changes that repeat a review finding.

What did it cost?

Check What it does
Cost and duration Notes when the total eval cost moves by 20% or more, or the duration by 50% or more.

See The PR comment for how each check shows up, and Configuration for the tolerance, the flaky rule and the gate.