A pull request that breaks a toolset is reported as a setup failure, not a failed eval
Example: a PR that renames a BashExecutorConfig field, breaks bash_instructions.jinja2, or breaks the Loki health check (toolset_grafana_loki.py:71-93) silently disables that toolset for users. In the eval, all 9 PR regression cases become '🚧 setup failure' under a '✅ Results of HolmesGPT evals' header with no failure warning, and the benchmark score is unchanged. Toolset config changes like the recent claude/standardize-toolset-config-fields branch are exactly this kind of PR.
bash:
⚠ bash, kubernetes/core, kubernetes/logs, helm/core, internet, robusta and others are enabled for every case, so their product prerequisite checks run in every eval
self.config = BashExecutorConfig(**config)
⚠ product code in bash's prerequisite; if it raises, holmes/core/tools.py:835-837 marks the toolset FAILED
raise ToolsetPrerequisiteError(
⚠ the harness turns any FAILED toolset that the config enables into this error
To fix: In property_manager.handle_test_error, count ToolsetPrerequisiteError as a regression instead of a setup failure, and have the PR comment show a failure warning whenever setup failures are above zero.