Writing
Measurement, attribution and the models that decide
-
18 September 2026
Jev vs Claude as an eval judge new
We ran TypeSafe's Jev against Haiku 4.5, Sonnet 5 and Opus 5 on 212 cases from our own production traffic. It tied on accuracy at 57x less cost and a 10x tighter latency tail.
-
10 September 2026
Is a cheaper model cheaper per resolved task?
Compare complete model costs, retries, fallbacks and resolved tasks, using a worked example that shows when a lower price per token still costs more per outcome.
-
10 September 2026
How to link AI retries to business outcomes
Keep retries, duplicate telemetry and business outcomes distinct, with a runnable Python ledger that attributes every attempt to the result it produced.
-
9 September 2026
How to measure AI cost per successful workflow
Include failed attempts, retries and delayed outcomes when measuring what an AI workflow costs to complete, not just what the successful calls billed.