//EVALS
Reliability evidence
A scorecard, published as-is. 60 cases test the real pipeline — real Anthropic calls, real Supabase writes — across all seven categories in the project spec, at the spec's own target counts. 18 were tuned directly against while building the rule engine (two real bugs were found and fixed that way); the other 42 were designed from the scoring spec and run once, held out from tuning. Both numbers are shown below, not blended into one.
100%
Accuracy across 60 cases
0
False-score rate — a number emitted on evidence that should have refused
0
False-refusal rate — a refusal on evidence that should have scored
1.0s
Mean latency per case · ~$0.60 total
//DEV SET VS. HELD-OUT SET
100%
Dev set (tuned against) · 18/18 cases
100%
Held-out set (not tuned against) · 42/42 cases
//BY CATEGORY
| Category | Cases | Passed | Accuracy |
|---|---|---|---|
| Sales-ready | 15 | 15 | 100% |
| Needs review | 10 | 10 | 100% |
| Nurture | 10 | 10 | 100% |
| Disqualified | 10 | 10 | 100% |
| Duplicate / merge review | 5 | 5 | 100% |
| Insufficient evidence | 5 | 5 | 100% |
| Adversarial (prompt injection) | 5 | 5 | 100% |
Last run: 8/7/2026, 4:18:21 PM · eval set v2-60cases