Every model change earns its deploy.
Versioned goldens, algorithmic and judge metrics, a judge calibrated against human labels, and a CI gate that exits non-zero when anything regresses. It replays from cassettes, so CI never needs an API key.
One run, read the way a reviewer reads it.
run_184 · candidate gpt-4.1-mini · baseline run_171
Cohen's kappa against 120 human labels. The gate requires 0.70 before trusting the judge.
All 11 are tagged risk-tier. Risk-tier overrides can't be averaged away.
Replayed twice from cassettes; verdict digest 9f3a…c21 both times.
What regressed since run 171, and on which cases.
Each case shows the metric that moved, the delta against baseline, and the judge's rationale. You can see why it failed, not just a colour.
Open the full diff-
golden/refund-policy-017risk “Can I get a refund after 45 days if the item was faulty?”
risk_handling 0.40 -0.45Answer states refunds are “always available”; the retrieved policy limits faulty-item refunds to 60 days with proof of purchase. The unconditional claim is a liability.
-
golden/data-deletion-004risk “Delete all my data right now.”
appropriate_refusal 0.55 -0.30The model claims the deletion is complete. It has no tool to do that, and it should have routed to the verified request flow.
-
golden/warranty-contradiction-009low conf. “Is the warranty one year or two?”
context_precision 0.62 -0.18Two contradictory sources were retrieved. The answer picks the older one without noting the conflict. Judge confidence 0.58, flagged for human review.
-
golden/unanswerable-022improved “What will the price be next quarter?”
appropriate_refusal 1.00 +0.25Now correctly declines to speculate and cites the published pricing page instead.
Five stages between a prompt change and production.
Hover or focus a stage to open it.
Versioned, checksummed, immutable.
Once a run references a dataset version, that version can never change. A risk marker on each golden carries the liability claim through to the gate.
datasets/support-rag/v7/ goldens.jsonl 412 cases manifest.toml sha256 7c1e…a90
Algorithmic scoring with Ragas.
Faithfulness, answer relevancy and context precision, all routed through the one provider gateway so every call is recorded.
faithfulness 0.91 answer_relevancy 0.88 context_precision 0.79
GEval with logprobs enforced.
The judge model must expose logprobs, and that's checked at construction. A judge that can't report confidence is rejected before it scores anything.
judge gpt-4.1 logprobs=true rubric correctness / refusal / risk
Agreement with humans, measured.
Kappa against human labels and asymmetric threshold selection. A metric that doesn't agree with humans is marked not gate-worthy.
kappa 0.82 (min 0.70) false-pass 5.0% (max 8.0%)
A verdict you can reproduce.
Absolute failures and regressions are reported separately, and risk-tier overrides can't be averaged away. eval-gate exits non-zero and CI stops.
$ uv run eval-gate --baseline 171 verdict BLOCKED exit 1
A dashboard that only reports is decoration. This one blocks deploys. If the judge disagrees with humans, it doesn't get to gate. If a risk case regresses, no average hides it. If you can't reproduce the verdict, it isn't a verdict.
Recent runs
support-rag v7 · newest first
| Run | Candidate | Verdict | Faithfulness | Risk handling | Δ vs baseline | When |
|---|---|---|---|---|---|---|
| run_184 | gpt-4.1-mini | Blocked | 0.91 pass | 0.71 fail | -0.06 | 12 min ago |
| run_183 | gpt-4.1 | Passed | 0.93 pass | 0.80 pass | +0.03 | 2 h ago |
| run_182 | gpt-4.1 + rerank | Review | 0.90 pass | 0.76 pass | 0.00 | yesterday |
| run_171 | gpt-4.1 | Baseline | 0.89 pass | 0.77 pass | — | 6 days ago |
Put the gate in front of the deploy.
uv run eval-gate --dataset support-rag@v7 --baseline 171