Case study · PrefectHQ/prefect-mcp-server

14 failing evals that never got graded

Prefect's MCP server gives AI assistants read-only tools for inspecting Prefect workspaces. Its evals read like support cases: set up a broken workspace, ask the agent what a user would ask, and have a model check that the answer names the real cause. In August 2026, pull request #200 fixed two tools, and its eval report came back 14 failing. None of those evals had been graded: the model that grades them had been retired. We ran evalship's setup review on the commit the pull request branched from, and replayed its real eval results through evalship.

14 of 20
evals the repo's eval report listed as failing on the first push
0 of 14
of those were graded: the judge model returned 404
0
regressions evalship reported; it set the 14 aside as couldn't run, with the error
4
flaws evalship's review found in its evals

What evalship's review found

At setup, your coding agent reads every scorer, judge and prompt, and writes down where the evals fall short. On PrefectHQ/prefect-mcp-server at 2940a3e, the commit #200 branched from, it summed them up like this:

20 agent evals (pytest with Pydantic AI, 14 of them judged by Claude Opus) run the real MCP server against a throwaway Prefect API on pull requests. They monitor and never fail the build. Several cases don't set up what they claim to test: three concurrency cases blame a limit no run is using, the shell tool reports every successful CLI command as exit 124 (its timeout code), and one automation case puts the answer in the fixture.
1 Important Eval bug

Three concurrency evals blame a limit that no run is using

In the work-pool, work-queue and tag cases every run is forced to Late and nothing occupies the one slot. So the judge passes an agent that blames any configured limit of 1, and fails one that checks usage and correctly says the limit is free. For the tag case, the tool output itself says over_limit: false. No case has a limit that is set but free, so a PR that makes the agent, or the queue hint, blame idle limits stays green.

evals/late_runs/test_work_pool_concurrency.py:63
# Create multiple flow runs - first will consume the slot, others will be Late

⚠ lines 73-76 then force all three runs to Late, so nothing holds the slot

evals/late_runs/test_work_pool_concurrency.py:132
that its concurrency limit is exhausted/full, not just give generic

⚠ the judge requires 'exhausted' for a limit with no run in it

evals/late_runs/test_work_queue_concurrency.py:131
pool/queue name and that the queue's concurrency limit is exhausted,

⚠ this fixture also forces every run to Late (line 82)

To fix: Fill the slot the way test_deployment_concurrency.py:71-77 does, with a Running flow run in the pool or queue, and rethink the tag case, since its limit counts tasks a Late run never starts. Add one case with a free limit of 1, where blaming the limit fails.

2 Important Eval bug

The eval shell tool reports every successful prefect command as exit code 124

All four evals that change state through the Prefect CLI (trigger a run, cancel late runs, create a reactive or proactive automation) show the agent a timeout exit code for every command that worked. An agent that sensibly retries fails the trigger case's single-run assert, and one that ignores exit codes passes. Real clients run the CLI in a real shell, so these cases measure how the model copes with a broken tool, not how it uses the CLI.

evals/_tools/run_shell_command.py:49
exit_code=proc.returncode or 124, stdout=out.decode(), stderr=err.decode()

⚠ a successful command returns 0, which is falsy, so the agent sees 124

evals/_tools/run_shell_command.py:47
return ShellResult(exit_code=124, stdout="", stderr="Timed out after 20s")

⚠ 124 is the tool's own timeout code

evals/test_trigger_deployment_run.py:75
assert len(flow_runs) == 1

⚠ an agent that retries the apparently failed `prefect deployment run` creates a second run and fails here

To fix: Return proc.returncode unchanged (it is set once communicate() returns) and keep 124 for the timeout branch only. Add a unit test next to tests/test_eval_tool_spy.py that runs `prefect version` through the tool and expects exit code 0.

3 Important Suspiciously easy suite

The automation-not-firing case puts its answer in the automation's description

The agent can answer 'it needs 3 failures and only 1 happened' from the compact get_automations listing without reading the trigger. A PR that breaks the detail path (the id lookup at automations.py:22-29 or the trigger dump at line 74) still passes this case, which is the only one about debugging automations.

evals/automations/test_debug_not_firing.py:28
description="Automation that triggers after 3 consecutive failures",

⚠ states the answer; it is also inexact, since threshold=3 within=300 counts any 3 failures in 5 minutes

src/prefect_mcp_server/_prefect_client/automations.py:66
"description": automation.description,

⚠ included in the compact listing; the trigger config appears only in detail mode (line 74)

evals/automations/test_debug_not_firing.py:81
"Does the agent identify that the automation has a threshold of 3 but only 1 failure occurred, "

⚠ the description plus the one failure named in the user's question are enough to pass

To fix: Give the fixture a neutral description and assert that get_automations was called with an id filter for this automation, so the agent has to read trigger.threshold.

4 Warning Suspiciously easy suite

The log-correlation rate-limit case passes without checking rate limits

This is the only case that starts from log evidence, and the step it exists to check (reaching for review_rate_limits) is optional. A PR that rewrites the review_rate_limits description (server.py:518-541) so the agent stops using it on logs still passes: the agent reads the 429 lines, says it looks like rate limiting, and offers to check.

evals/rate_limits/test_cloud_correlate_logs.py:81
could offer to review the rate limits as the next step of investigation.

⚠ the judge accepts an answer that never calls review_rate_limits

evals/rate_limits/test_cloud_correlate_logs.py:26
logger.warning("Prefect API returned 429 Too Many Requests")

⚠ the flow's own logs already name the 429

evals/README.md:66
verifies agent can correlate 429 warnings in flow logs with rate limit data (Cloud)

⚠ no rate-limit eval asserts that review_rate_limits ran

To fix: Drop the 'Alternatively' sentence from the rubric, and add tool_call_spy.assert_tool_was_called("review_rate_limits") to this case.

What evalship would have said on #200

  1. Aug 6 · d2a9f10

    14 evals fail, and none of them ran

    The repo's eval report compares counts with the base branch: “14 ❌ +13”. Every one of the 14 failures is the same error, a 404 for claude-opus-4-1-20250805, the model that grades the agent's answers. The agent ran; nothing graded it.

    evalship reads each failure's message, sees an eval that never got a verdict, and says so in one line with the error, instead of counting 14 regressions. The 6 evals that were graded are compared with main as usual. Its reviewer adds what the eval report can't: none of the rate-limit evals calls the tool with the new workspace_id parameter, so the fix itself is only covered by unit tests.

    github.com/PrefectHQ/prefect-mcp-server/pull/200

    evalship-appbot commented 2 minutes ago

    evalship: 6 / 20 14 evals couldn't run

    No regressions in the 6 evals that ran. 14 evals couldn't run, so they aren't compared.

    🔌 Couldn't run: 14 evals in cloud_correlate_logs, cloud_direct, cloud_no_throttling, and 11 more failed before they were scored (pydantic_ai.exceptions.ModelHTTPError: status_code: 404, model_name: claude-opus-4-1-20250805, body: {'type': 'error'...), so they aren't compared with main@2940a3e.

    🐛 Eval gap: The new workspace_id parameter on review_rate_limits is only covered by unit tests; the rate-limits-and-cloud OAuth case (test_cloud_oauth_cross_workspace.py) never calls review_rate_limits or get_identity. Add a case there that asks for throttling in the second workspace and asserts the tool is called with that workspace_id (src/prefect_mcp_server/server.py:503).

    Open the full report for every eval's inputs, outputs and history.

    All checks have passed

    3 successful checks

    CI / evals (pull_request) Successful in 4m Details
    Evaluation Results 14 fail, 6 pass in 3m 57s Details
    evalship No regressions in the 6 evals that ran. 14 evals couldn't run, so they aren't compared. Details

    This branch has no conflicts with the base branch

    Merging can be performed automatically.

    Merge pull request
  2. Aug 7 · 0e86bd0

    The judge model is swapped

    “fix: replace retired claude-opus-4-1 eval judge model with claude-opus-5”. Every eval passes again. Changing the grader changes what every score means, so evalship's comment says the results may not compare with main's.

  3. Aug 7 · 8227588

    Green, with the grader change named

    “style: ruff format after eval model swap”. evalship edits the same comment. #200 merged on Aug 7.

    github.com/PrefectHQ/prefect-mcp-server/pull/200, after the last push

    zzstoatzz added 1 commit: style: ruff format after eval model swap

    8227588
    evalship-appbot commented 2 minutes ago edited

    evalship: 20 / 20 No regressions vs main

    No regressions, 1 improvement.

    Eval main this PR
    🟢 test_agent_correlates_logs_with_rate_limiting ❌ ✅

    ⚖️ Not comparable: core-diagnostics now uses anthropic:claude-opus-5 (evals/conftest.py), so its results won't compare with those on main@2940a3e.

    Open the full report for every eval's inputs, outputs and history.

    All checks have passed

    3 successful checks

    CI / evals (pull_request) Successful in 5m Details
    Evaluation Results All 20 tests pass in 4m 35s Details
    evalship No regressions, 1 improvement. Details

    This branch has no conflicts with the base branch

    Merging can be performed automatically.

    Merge pull request

How we made this

  • Replaying this pull request found three bugs in evalship, fixed before this page went up: a judge's verdict that mentioned "429 warnings" was read as a rate-limited run, a changed tool listed in the review was also reported as untested, and the check's title said "No regressions" without saying that most evals couldn't run.
  • The eval results are Prefect's own: the JUnit files its CI uploads on every pull request and push to main, read by evalship's JUnit parser. Main's history is the 8 runs from Jul 21 to Aug 4 whose results hadn't expired.
  • The review was written in October 2026, read-only, by evalship's setup prompt (v22) on 2940a3e, the commit #200 branched from, with no access to GitHub or anything after it.
  • The likely cause and the notes come from evalship's pull request reviewer. Its real request for each push was answered by Claude Sonnet 5.5, the reviewer's default model, in a Claude Code session instead of through the API, then checked by evalship's own validation, as every answer is.
  • The comments are evalship as it is today, with the gate setup adds: a new failure fails evalship's check.
  • Nothing was posted to PrefectHQ/prefect-mcp-server.

See what evalship finds in your evals

Ask your coding agent: Claude Code Codex Cursor GitHub Copilot Gemini CLI

Set up https://evalship.com/setup.md

More case studies: HolmesGPT/holmesgpt · launchdarkly/ai-tooling