01 / The price of a pass
Similar scores.
Very different recorded costs.
21.5×
gap in recorded cost
per successful test
Claude Sonnet 4.6≥ 14.42¢
Spending on all 90 attempts divided by passes. Bars start at zero.
For my product, that was enough reason to keep testing lower-cost models.
These bills are incomplete, so both costs are minimums. The 21.5× gap describes recorded spending, not guaranteed savings. Scoring costs are separate.
Compare all six shortlisted models ↗02 / A separate DeepSeek experiment
I changed the setup.
The pass rate reached 90%.
Clearer instructions. Better tools. One check before answering.
60%→90%
Early run: 18 of 30 passed.
Final version: 81 of 90 passed across three runs.
$1.52to run the final 90 tests
+ $3.66 to score them
Pass Fail · 24 of 30 tasks passed every time.
The model matters. So does everything I put around it.
Several changes were tested together, using questions I had already developed against. This shows progress on those tasks—not proof of which change caused it, or how it handles new tasks.
See what helped—and what didn’t ↗