Measure the request that reaches the model
Input tokens include the material your application sends: instructions, conversation history, retrieved passages and tool results. A short user question can therefore generate a much larger input.
Output tokens vary with the requested answer and model settings. Measure both sides over a representative sample instead of forecasting from a single prompt.
Count unsuccessful attempts
A malformed extraction, timeout or answer rejected by your validator may lead to another request. Include those attempts in your cost-per-result calculation.
Set a bounded retry policy. Retrying every failure immediately can increase both spend and load without improving the result, especially if the problem is an invalid request.
Compare models using the same task
Use the TextCortex API to run the same approved request set against candidate models.
- Hold the task and success criteria constant.
- Record model version and generation settings.
- Capture billed usage and current rates.
- Count valid results, repeated attempts and failures.
- Compare end-to-end time alongside cost.
Build the monthly forecast
Multiply the measured cost per completed task by expected task volume. Keep separate estimates for features that use different models or request shapes.
Use the token calculator for the basic arithmetic. Add charges outside that estimate according to the selected model’s billing rules, then review the forecast as actual traffic arrives.
