Define an accepted result
Write a rule or rubric for whether a task succeeded. An extraction can require valid fields and correct values; a generated change can require passing project checks.
Keep the acceptance definition the same across model candidates. Otherwise differences in the score may reflect the evaluation rather than the model.
Include every model call in the task
Add input and output usage from the first attempt, retries and any supporting model calls. Apply the current rate for each model used.
Divide the total measured API cost by the number of accepted tasks. Report the failure rate alongside that figure so a model that rejects most requests cannot look artificially attractive.
Run a controlled comparison
Use TextCortex to keep the connection consistent while testing different model families.
- Use the same representative task set.
- Record model identifiers and request settings.
- Apply the same output validation.
- Measure total usage, accepted outcomes and latency.
- Repeat when the prompt or model version changes.
Route where the result justifies it
A more capable model can be worthwhile for difficult requests even if its token rate is higher. A smaller or faster option may work well for a simpler feature.
Use the API cost calculator for token arithmetic and the routing guide to keep the production model choices explicit.
