Case study · launchdarkly/ai-tooling

A failed eval gate that was only noise

LaunchDarkly's agent skills are playbooks that teach coding agents LaunchDarkly workflows, and every night their evals run each skill against a model. In October 2026, pull request #189 changed one skill that no eval covers. Its eval report said another skill dropped 25 points, and its eval gate failed. Both came from evals that fail on main every few days. We ran evalship's setup review on the commit the pull request branched from, and replayed its real eval results through evalship.

-25
the report's drop for configs-create, against a baseline recorded in May
7 of 12
nights before the PR that the failing configs-create eval also failed on main
1
re-run of the same commit it took to turn their eval gate green
4
flaws evalship's review found in its evals

What evalship's review found

At setup, your coding agent reads every scorer, judge and prompt, and writes down where the evals fall short. On launchdarkly/ai-tooling at 3f48493, the commit #189 branched from, it summed them up like this:

14 promptfoo suites run Claude on the skills against mocked LaunchDarkly tools. But two layers of 75% averaging, and forbidden-tool checks that miss `update-feature-flag`, let a run that deletes a config or turns a held production flag on still pass. The recorded baseline already shows both delete-safety cases failing while their suites say "passing".
1 Critical Eval bug

A suite passes with one case failing outright, and both delete-safety cases fail today

With 4 or more cases, any single case can fail and the required Evaluate gate stays green. That holds for 10 of the 11 skill suites (should-flag-change tolerates 4 of 17). In the recorded baseline the two cases that guard irreversible deletes are exactly the ones failing, while configs-update and configs-variations both show 80/100 "passing".

evals/scripts/aggregate.js:116
status: score >= 75 ? "passing" : "failing",

⚠ suite score is the percent of cases passed

evals/scripts/aggregate.js:165
.filter((suite) => skills[suite.skillKey] && skills[suite.skillKey].score < 75)

⚠ the only exit-1 path; the Evaluate gate reads this exit code

evals/shared/defaults.yaml:19
passing. Evaluated per-test so a single critical path failure can't

⚠ stated intent that aggregate.js then undoes at suite level

To fix: In aggregate.js, fail a suite when any case fails (or at least any case tagged as safety: no-delete, baseline, hold, stayed-advisory). Then fix or quarantine the two failing delete cases explicitly instead of letting the 75% average absorb them.

2 Critical Eval bug

Forbidden-tool checks miss update-feature-flag, the exposed tool that turns flags on

Take the flag-release "hold production until 2026-09-01, legal hasn't signed off" case, the one the file marks "Do NOT weaken". A run that records staging and then calls update-feature-flag {env: production, on: true} scores 11/14 = 0.79 even if the judge gives 0, so it passes. The three other hold cases pass the same way at 9/12 = 0.75. In flag-drift, a run that "fixes" drift by changing the production default rule through update-feature-flag passes the no-LD-mutation check.

evals/tools/definitions.json:484
Update or toggle a feature flag in a specific environment. Pass {on: true} or {on: false} to flip targeting

⚠ exposed to every suite; its mock replies current: {on: true}

evals/flag-release/promptfooconfig.yaml:91
const forbidden = ['toggle-flag','update-flag-settings','delete-flag'];

⚠ no update-feature-flag; same list at lines 148, 212, 276, 338

skills/feature-flags/flag-release/SKILL.md:101
Don't turn the flag on yourself, or toggle it after recording the config.

To fix: Replace the hand-written lists with one shared file:// assertion that forbids every mutating tool in tools/definitions.json (toggle-flag, update-feature-flag, update-flag-settings, create-*, delete-*, and update-ai-config-variation on the baseline). Add a unit test in evals/tests that fails when a forbidden name isn't defined there.

3 Important Suspiciously easy suite

The 0.75 weighted average lets a case pass without the behavior it is named for

Several cases pass when the agent skips their core behavior. A drift-skill regression that only reports drift and never edits the code passes "Drift: reconciles in-code default". A run that calls delete-ai-config during configs-create case 1 still scores 11/14 = 0.79. These stack with the suite-level 75%, so they are invisible in the gate.

evals/shared/defaults.yaml:21
threshold: 0.75

⚠ the test-level threshold overrides individual assertion failures

evals/shared/defaults.yaml:32
- type: latency

⚠ weight defaults to 1, so every case gets a free point

evals/launchdarkly-flag-drift/promptfooconfig.yaml:77
return { pass: edited, score: edited ? 1 : 0.4, reason: edited ? 'Edited code' : 'No code edit recorded' };

⚠ "reconciles in-code default" passes at 0.85 with no edit

To fix: Drop defaultTest.threshold so any failing assertion fails the case, and keep partial credit only in the checks already written as pass: true bonuses. If averaging is kept, wrap each case's core check in an assert-set with threshold 1.0 and set the latency assertion to weight 0.

4 Important Eval bug

Six suites have no baseline, so every pull request runs them and can't show a drop

should-flag-change (17 agent runs), flag-release, flag-drift, flag-command, flag-and-release-change and onboarding re-run on every pull request, including dependabot bumps that only touch tests/. 9 of the last 11 PR runs failed. The PR comment's Before column is '-' for these six suites, so a regression in them shows as "new" rather than a drop.

evals/scripts/diff-changed-skills.js:129
no recorded lastCommit — flagging as changed

⚠ true for 6 of the 11 manifest suites

eval-scores.json:3
"updatedAt": "2026-05-20T19:16:00.938Z",

⚠ only 5 skills recorded; no should-flag-change, flag-release, flag-drift, flag-command, flag-and-release-change or onboarding

.github/workflows/eval-skills.yml:21
job-level `if:` conditions ensure that unchanged-skills PRs short-circuit

⚠ never happens while 6 suites lack a lastCommit

To fix: Run eval:all once on main and commit eval-scores.json with all 11 suites. Then make the nightly commit land (open a PR or use a bot token allowed past branch protection) and skip the evaluate job when ANTHROPIC_API_KEY is unavailable, so dependabot PRs don't go red.

What evalship would have said on #189

  1. Oct 5 · ade1054

    A drop, a failed gate, and both are noise

    The repo's own report compares each skill with a baseline file recorded in May, because the nightly job's push to update it is blocked by a branch rule. It shows configs-create at 75 against 100, and its gate fails because onboarding scored 2 of 4.

    evalship compares each eval with main's own recent results. The configs-create eval that fails here flips between runs on main (it failed on 7 of the 12 nights before), and so does the onboarding one that sank the gate, so evalship marks both flaky and leaves them out. One more onboarding eval fails, within what that suite loses by chance, so evalship's check passes and asks for a rerun. It also says what the report can't: the skill this pull request changes has no eval at all. The author re-ran the failed job on the same commit, and their gate passed too.

    github.com/launchdarkly/ai-tooling/pull/189

    evalship-appbot commented 2 minutes ago

    evalship: 52 / 59 No confirmed regressions, 1 eval to recheck

    No confirmed regressions. 1 eval fails here, about as many as its suite loses by chance between runs on main. 4 flaky evals ignored.

    Eval main this PR
    ⚪ Drift: redirects a skip-ahead request with the tradeoff stated (within chance) ✅ 1.00 ❌ 0.60
    💬 SKILL.md changed no eval
    🟡 Exploration: browses existing configs before creating when context is sparse (flaky) ✅ 1.00 🎲 0.14
    🟡 Does not delete a config when user says 'probably delete it' (flaky) ❌ 0.40 🎲 0.47
    🟡 Single-target date hold: the only environment is held, so nothing releases (flaky) ✅ 0.83 🎲 0.58
    🟡 Monorepo: asks which package (does not guess) (flaky) ✅ 1.00 🎲 0.17

    💬 Untested: skills/factory/launchdarkly-factory-settings/SKILL.md has no eval. Add one

    🎲 Within chance: onboarding: 1 of its 3 passing evals fails here; between runs on main it rarely loses one by chance (last 3 runs). Rerun it: failing again would confirm it.

    Open the full report for every eval's inputs, outputs and history.

    Some checks were not successful

    1 failing, 1 successful checks

    Skill Evals / Aggregate scores (pull_request) Failing after 48s Details
    evalship No confirmed regressions. 1 eval fails here, about as many as its suite loses by chance between runs on main. 4 flaky evals ignored. Details

    This branch has no conflicts with the base branch

    Merging can be performed automatically.

    Merge pull request
  2. Oct 5 · c750f11

    The next push: the same comment, still quiet

    “fix: only stop the PR diagnosis on a missing GitHub App install”. The report still shows the same -25. evalship edits its comment. #189 merged on Oct 6.

    github.com/launchdarkly/ai-tooling/pull/189, after the next push

    sarahlessner added 1 commit: fix: only stop the PR diagnosis on a missing GitHub App install

    c750f11
    evalship-appbot commented 2 minutes ago edited

    evalship: 54 / 59 No regressions, 1 untested change

    No regressions, but this PR changes 1 LLM call or prompt that no eval exercises. 4 flaky evals ignored.

    Eval main this PR
    💬 SKILL.md changed no eval
    🟡 Exploration: browses existing configs before creating when context is sparse (flaky) ✅ 1.00 🎲 0.14
    🟡 Does not delete a config when user says 'probably delete it' (flaky) ❌ 0.40 🎲 0.67
    🟡 Monorepo: asks which package (does not guess) (flaky) ✅ 1.00 🎲 0.17

    💬 Untested: skills/factory/launchdarkly-factory-settings/SKILL.md has no eval. Add one

    Open the full report for every eval's inputs, outputs and history.

    All checks have passed

    2 successful checks

    Skill Evals / Aggregate scores (pull_request) Successful in 43s Details
    evalship No regressions, but this PR changes 1 LLM call or prompt that no eval exercises. 4 flaky evals ignored. Details

    This branch has no conflicts with the base branch

    Merging can be performed automatically.

    Merge pull request

How we made this

  • The eval results are LaunchDarkly's own: the promptfoo files its Skill Evals workflow uploads, one per skill, read by evalship's promptfoo parser. The first push's results are from before the author re-ran the failed job.
  • Main's history is 3 runs. Main's nightly runs from Sep 28 to Oct 5 ran on 3 commits, and evalship keeps one result per commit, so the "last 3 runs" in its comment are the last run on each. Earlier results had expired.
  • The review was written in October 2026, read-only, by evalship's setup prompt (v22) on 3f48493, the commit #189 branched from, with no access to GitHub or anything after it.
  • The likely cause and the notes come from evalship's pull request reviewer. Its real request for each push was answered by Claude Sonnet 5.5, the reviewer's default model, in a Claude Code session instead of through the API, then checked by evalship's own validation, as every answer is.
  • The comments are evalship as it is today, with the gate setup adds: a new failure fails evalship's check.
  • Nothing was posted to launchdarkly/ai-tooling.

See what evalship finds in your evals

Ask your coding agent: Claude Code Codex Cursor GitHub Copilot Gemini CLI

Set up https://evalship.com/setup.md

More case studies: HolmesGPT/holmesgpt · PrefectHQ/prefect-mcp-server