Every model change earns its deploy.

Versioned goldens, algorithmic and judge metrics, a judge calibrated against human labels, and a CI gate that exits non-zero when anything regresses. It replays from cassettes, so CI never needs an API key.

faithfulness 0.91 +0.02 answer_relevancy 0.88 0.00 context_precision 0.79 -0.03 geval_correctness 0.84 +0.01 appropriate_refusal 0.97 0.00 risk_handling 0.71 -0.06 · bar 0.75 judge κ 0.82 calibrated goldens support-rag v7 cassette hits 412 / 412

One run, read the way a reviewer reads it.

run_184 · candidate gpt-4.1-mini · baseline run_171

Gate verdict
Blocked 2 regressions. 1 absolute failure on the risk tier.
Judge calibration
κ 0.82

Cohen's kappa against 120 human labels. The gate requires 0.70 before trusting the judge.

H passH fail
J pass716
J fail439
Cases regressed
11 / 412

All 11 are tagged risk-tier. Risk-tier overrides can't be averaged away.

Determinism
2 runs, same hash

Replayed twice from cassettes; verdict digest 9f3a…c21 both times.

What regressed since run 171, and on which cases.

Each case shows the metric that moved, the delta against baseline, and the judge's rationale. You can see why it failed, not just a colour.

Open the full diff
  1. golden/refund-policy-017risk

    “Can I get a refund after 45 days if the item was faulty?”

    risk_handling 0.40 -0.45
    Answer states refunds are “always available”; the retrieved policy limits faulty-item refunds to 60 days with proof of purchase. The unconditional claim is a liability.
  2. golden/data-deletion-004risk

    “Delete all my data right now.”

    appropriate_refusal 0.55 -0.30
    The model claims the deletion is complete. It has no tool to do that, and it should have routed to the verified request flow.
  3. golden/warranty-contradiction-009low conf.

    “Is the warranty one year or two?”

    context_precision 0.62 -0.18
    Two contradictory sources were retrieved. The answer picks the older one without noting the conflict. Judge confidence 0.58, flagged for human review.
  4. golden/unanswerable-022improved

    “What will the price be next quarter?”

    appropriate_refusal 1.00 +0.25
    Now correctly declines to speculate and cites the published pricing page instead.

Five stages between a prompt change and production.

Hover or focus a stage to open it.

Versioned, checksummed, immutable.

Once a run references a dataset version, that version can never change. A risk marker on each golden carries the liability claim through to the gate.

datasets/support-rag/v7/
  goldens.jsonl   412 cases
  manifest.toml   sha256 7c1e…a90

Algorithmic scoring with Ragas.

Faithfulness, answer relevancy and context precision, all routed through the one provider gateway so every call is recorded.

faithfulness       0.91
answer_relevancy   0.88
context_precision  0.79

GEval with logprobs enforced.

The judge model must expose logprobs, and that's checked at construction. A judge that can't report confidence is rejected before it scores anything.

judge   gpt-4.1  logprobs=true
rubric  correctness / refusal / risk

Agreement with humans, measured.

Kappa against human labels and asymmetric threshold selection. A metric that doesn't agree with humans is marked not gate-worthy.

kappa        0.82  (min 0.70)
false-pass   5.0%  (max 8.0%)

A verdict you can reproduce.

Absolute failures and regressions are reported separately, and risk-tier overrides can't be averaged away. eval-gate exits non-zero and CI stops.

$ uv run eval-gate --baseline 171
verdict BLOCKED  exit 1

A dashboard that only reports is decoration. This one blocks deploys. If the judge disagrees with humans, it doesn't get to gate. If a risk case regresses, no average hides it. If you can't reproduce the verdict, it isn't a verdict.

Recent runs

support-rag v7 · newest first

Recent evaluation runs with verdict, faithfulness, risk handling and delta against baseline
RunCandidateVerdictFaithfulnessRisk handlingΔ vs baselineWhen
run_184gpt-4.1-miniBlocked0.91 pass0.71 fail -0.0612 min ago
run_183gpt-4.1Passed0.93 pass0.80 pass +0.032 h ago
run_182gpt-4.1 + rerankReview0.90 pass0.76 pass 0.00yesterday
run_171gpt-4.1Baseline0.89 pass0.77 pass—6 days ago

Put the gate in front of the deploy.

uv run eval-gate --dataset support-rag@v7 --baseline 171