← valuemaxx

Is a cheaper model cheaper per resolved task?

Compare the total model cost of two matched groups of tasks, including retries, fallbacks and unsuccessful work, divided by the resolved tasks in each group. Check resolution rate and completion time alongside that ratio. A lower price per call doesn't settle the decision.

The unit you're buying is a finished customer task. A model can produce a cheaper first attempt while sending more tasks through a fallback, or leaving more work unresolved. Those paths belong in the comparison.

Define “resolved” before changing the model

Choose a business outcome you can observe independently of the model's answer. For a support workflow, a hypothetical rule is resolution within three days of arrival, with no reopen in the following seven days. Allow every task ten days from arrival before comparing the groups. Use the same rule and window for both.

Keep tasks whose window is still open separate from mature results. After the window closes, an observed task that missed the rule is unsuccessful. A task whose outcome couldn't be collected is unknown; elapsed time doesn't repair missing telemetry.

Include the whole model path

Here's a hypothetical comparison, not a benchmark or customer result. Both groups contain 1,000 tasks with comparable difficulty and complete records. All outcome windows have closed. Costs cover model calls only, including calls on unsuccessful tasks; labor, tools and infrastructure are excluded.

Hypothetical model costs and mature resolutions
MeasurementCurrent model pathCheaper first model
First-attempt model cost$80$40
Additional retry and fallback model cost$20$50
Total model cost$100$90
Resolved tasks800600
Observed unsuccessful tasks200400
Resolution rate80%60%
Model cost per resolved task$0.125$0.150

The cheaper first model reduces total model spend by $10, but produces 200 fewer resolutions. Its model cost per resolution is 20% higher: $90 ÷ 600, compared with $100 ÷ 800. The extra $50 includes every retry and fallback charge; it doesn't count the initial $40 again.

This example demonstrates the arithmetic. It doesn't predict any particular model's performance. If a group has no resolutions, report the spend and “no resolved tasks”; a zero denominator isn't a zero cost per resolution.

Keep the comparison fair

Randomly assign eligible incoming tasks to the two paths when practical, keeping that assignment through retries and fallbacks. Balance important categories, such as language, task type and difficulty. A cheaper path receiving mostly easy tasks won't tell you how it handles your normal workload.

Hold prompts, tools, retrieval, retry limits and outcome scoring fixed while testing the first model. If several components must change together, label it a comparison of two workflow configurations. You can't attribute the result solely to the model.

Decide the trial size, review date and acceptable resolution-rate, latency and safety limits in advance. Start with bounded traffic. Stop for a breached guardrail; don't promote a path just because an early cost ratio looks good. Review uncertainty and results within the important task categories before expanding it.

Show what's missing

Report cost-record coverage and outcome coverage for each group. Don't silently drop tasks with missing usage or outcomes: that can make either path look better. Keep charges with unknown task IDs visible as unassigned spend. Hold the final cost comparison until those gaps are reconciled.

Label billed charges separately from price-based estimates. If human handling changes, the model-only ratio can't tell you total delivery cost. Measure labor separately before making that claim; missing labor data doesn't mean free labor.

Start with one workflow and one controlled comparison. The cost-per-successful-workflow guide explains the denominator. The retry-to-outcome implementation guide shows how to keep attempts and charges attached to the original task.