Skip to content

TextCortex API guides

The real cost of an LLM API

Calculate LLM API costs across input, output, retries and model selection. Compare your workload through the TextCortex multi-model API.

The real cost of an LLM API

The practical takeaway

Build on a multi-model API.

Forecast API costs from measured requests. A common API makes it easier to compare model choices without confusing provider integration differences with the cost of the model itself.

Measure the request that reaches the model

Input tokens include the material your application sends: instructions, conversation history, retrieved passages and tool results. A short user question can therefore generate a much larger input.

Output tokens vary with the requested answer and model settings. Measure both sides over a representative sample instead of forecasting from a single prompt.

Count unsuccessful attempts

A malformed extraction, timeout or answer rejected by your validator may lead to another request. Include those attempts in your cost-per-result calculation.

Set a bounded retry policy. Retrying every failure immediately can increase both spend and load without improving the result, especially if the problem is an invalid request.

Compare models using the same task

Use the TextCortex API to run the same approved request set against candidate models.

  • Hold the task and success criteria constant.
  • Record model version and generation settings.
  • Capture billed usage and current rates.
  • Count valid results, repeated attempts and failures.
  • Compare end-to-end time alongside cost.

Build the monthly forecast

Multiply the measured cost per completed task by expected task volume. Keep separate estimates for features that use different models or request shapes.

Use the token calculator for the basic arithmetic. Add charges outside that estimate according to the selected model’s billing rules, then review the forecast as actual traffic arrives.

Questions

Your next questions, answered.

How is API usage priced?

Use the rates for your selected model and account. Estimate input and output tokens separately and include any applicable cached-token, reasoning-token or other charges. API consumption is distinct from a workspace subscription.

How does model routing work?

Your application selects an available model in each API request, and TextCortex routes the request through its API. You can choose different models for different tasks. Automatic selection and fallback behavior should only be used where documented for your configuration.

Is TextCortex limited to open-source models?

No. The API provides access to proprietary model families such as GPT, Claude and Gemini alongside open-weight families such as Kimi, GLM, DeepSeek and MiMo. Use GET /models to list the identifiers available to your account.

TextCortex AI

One integration. Your choice of models.

Sign up for TextCortex, create an API key and start building with the models your application needs.

Sources and documentation

Product documentation and provider references used for this guide. Reviewed on 5 October 2026.