About evalship

evalship is quality control for your LLM evals. It lives in GitHub, next to the evals you already have.

Why it exists

Evals are meant to catch an LLM feature getting worse before your users do. Often they don't. The judge counts errors as passes. No case tests the refund limit in your system prompt. A pull request rewrites a prompt that nothing covers. One flaky run fails, and the team learns to ignore red.

A score of 0.92 on a dashboard shows none of that. evalship checks the evals themselves, then keeps watch on every pull request.

What it does

  • Reviews your evals at setup. Your coding agent audits them with the setup prompt, makes them run in GitHub Actions and opens a setup PR.
  • Comments on every pull request. One comment, an evalship check and a README badge, whichever eval frameworks you run.
  • Shows untested changes. Prompts and LLM calls in the diff that no eval covers.
  • Separates regressions from noise using each eval's history on your default branch.

What it doesn't do

evalship doesn't run your evals: your workflow does, with your API keys. It doesn't store datasets, trace production traffic, manage prompts or compare experiments. If you use Langfuse, Braintrust or LangSmith, keep them; each case can link to its trace.

Who's behind it

evalship is made by Garm Tech B.V., a software company in Amsterdam. Questions, ideas or a bug: [email protected].

Garm Tech B.V., Staalkade 6 O, 1011 JN Amsterdam, the Netherlands. Chamber of Commerce 99602075, VAT NL869056748B01.

Set up https://evalship.com/setup.md