Langfuse, Braintrust and LangSmith

Keep your eval platform. It stays where you store datasets, read traces and compare experiments; evalship reviews the results in the pull request. Your eval script keeps reporting to the platform and also writes an evalship JSON file that CI uploads. Each case's url points at its trace, so every changed eval in the PR comment links straight to it, and the report links to it too.

Each recipe below adds about 15 lines to the script that runs your experiment. They take a case's lowest score, so one failing scorer fails the case: an average can hide a hard rule that broke. Scores should be between 0 and 1.

The workflow

Run the script on pull requests and on pushes to your default branch, then upload the file. Add your platform's keys and your model provider's key as secrets (gh secret set LANGFUSE_SECRET_KEY).

      - run: python evals/run_experiment.py
        env:
          LANGFUSE_PUBLIC_KEY: ${{ secrets.LANGFUSE_PUBLIC_KEY }}
          LANGFUSE_SECRET_KEY: ${{ secrets.LANGFUSE_SECRET_KEY }}
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
      - uses: actions/upload-artifact@v4
        if: always()
        with:
          name: evalship-results
          path: eval-results.json

The scripts don't exit non-zero when cases fail. The evalship check reports failures, and .evalship.yml can make it fail the build. See Connecting CI for the rest.

Langfuse

For the Python SDK v3 and later, with a dataset stored in Langfuse. Set LANGFUSE_BASE_URL too if you don't use Langfuse Cloud EU.

import json
from langfuse import get_client, Evaluation
from app.support_bot import answer

THRESHOLD = 0.8
langfuse = get_client()

def task(*, item, **kwargs):
    return answer(item.input)

def exact(*, output, expected_output, **kwargs):
    return Evaluation(name="exact", value=float(output == expected_output))

def trace_url(trace_id):
    try:
        return langfuse.get_trace_url(trace_id=trace_id)
    except Exception:  # a missing link shouldn't fail the run
        return None

dataset = langfuse.get_dataset("support-bot")
result = dataset.run_experiment(name="ci", task=task, evaluators=[exact])

cases, ran = [], set()
for r in result.item_results:
    ran.add(r.item.id)
    scores = [e.value for e in r.evaluations if isinstance(e.value, (int, float))]
    score = min(scores) if scores else None
    cases.append({"key": r.item.id, "score": score, "threshold": THRESHOLD,
                  "passed": score is not None and score >= THRESHOLD,
                  "input": r.item.input, "output": r.output, "url": trace_url(r.trace_id)})
# run_experiment leaves out items whose task raised, so add them back as failures
cases += [{"key": item.id, "passed": False, "input": item.input, "failure": "The task raised an error."}
          for item in dataset.items if item.id not in ran]

json.dump({"schema": "evalship/v1", "suite": "support-bot", "url": result.dataset_run_url, "cases": cases},
          open("eval-results.json", "w"), default=str)
  • The result is an object: use result.dataset_run_url, not result["dataset_run_url"] as some examples show.
  • With local data (langfuse.run_experiment(data=[...])), r.item is your own dict and dataset_run_url is None. Use an id from your data as the key.

Braintrust

Run the script with python. Under braintrust eval, Eval() only registers the evaluator and returns no results.

import json
from braintrust import Eval, init_dataset
from app.support_bot import answer

THRESHOLD = 0.8

def task(input, hooks):
    hooks.metadata["trace_url"] = hooks.span.link()  # this case's row in the experiment
    return answer(input)

def exact(input, output, expected):
    return 1.0 if output == expected else 0.0

result = Eval("support-bot", data=init_dataset("support-bot", "golden"), task=task, scores=[exact])

cases = []
for r in result.results:
    scores = [s for s in r.scores.values() if s is not None]
    score = min(scores) if scores else None
    cases.append({"key": (r.origin or {}).get("id") or r.metadata["id"], "score": score, "threshold": THRESHOLD,
                  "passed": r.error is None and score is not None and score >= THRESHOLD,
                  "failure": str(r.error) if r.error else None,
                  "input": r.input, "output": r.output, "url": r.metadata.get("trace_url")})

json.dump({"schema": "evalship/v1", "suite": "support-bot", "url": result.summary.experiment_url, "cases": cases},
          open("eval-results.json", "w"), default=str)
  • Cases from a Braintrust dataset carry the record id in r.origin. With local data, give each record "metadata": {"id": "..."}.
  • A case whose task raised has no scores and sets r.error, so it counts as failed.

LangSmith

Needs langsmith 0.7.25 or later for results.url.

import json
from langsmith import Client
from app.support_bot import answer

THRESHOLD = 0.8
client = Client()

def target(inputs: dict) -> dict:
    return {"answer": answer(inputs["question"])}

def exact(outputs: dict, reference_outputs: dict) -> bool:
    return outputs["answer"] == reference_outputs["answer"]

results = client.evaluate(target, data="support-bot", evaluators=[exact], experiment_prefix="ci")

cases = []
for row in results:
    run, example = row["run"], row["example"]
    scores = [float(e.score) for e in row["evaluation_results"]["results"] if e.score is not None]
    score = min(scores) if scores else None
    cases.append({"key": str(example.id), "score": score, "threshold": THRESHOLD,
                  "passed": run.error is None and score is not None and score >= THRESHOLD,
                  "failure": run.error, "input": example.inputs, "output": run.outputs,
                  "url": run.get_url()})

json.dump({"schema": "evalship/v1", "suite": "support-bot", "url": results.url, "cases": cases},
          open("eval-results.json", "w"), default=str)
  • Boolean evaluators return True or False as the score, hence float().
  • run.get_url() makes one request per case.

TypeScript

The same fields exist in camelCase. Langfuse: result.itemResults, r.traceId, result.datasetRunUrl and await langfuse.getTraceUrl(traceId). Braintrust: result.summary.experimentUrl. LangSmith's TypeScript results have no URL, so leave url out.

Keys

evalship matches cases across runs by suite and key. The recipes use dataset item and example ids, which stay the same as long as you edit the dataset rather than recreate it. Recreating it starts every case's history over. See Stable keys.