Skip to content

TextCortex API guides

Compare LLM APIs by cost per successful task

Compare model APIs using accepted outputs, retries, token usage and latency. Evaluate candidate models through the TextCortex API.

Compare LLM APIs by cost per successful task

The practical takeaway

Build on a multi-model API.

The cheapest request is not always the cheapest useful result. Compare total API consumption against tasks that meet your application’s acceptance criteria.

Define an accepted result

Write a rule or rubric for whether a task succeeded. An extraction can require valid fields and correct values; a generated change can require passing project checks.

Keep the acceptance definition the same across model candidates. Otherwise differences in the score may reflect the evaluation rather than the model.

Include every model call in the task

Add input and output usage from the first attempt, retries and any supporting model calls. Apply the current rate for each model used.

Divide the total measured API cost by the number of accepted tasks. Report the failure rate alongside that figure so a model that rejects most requests cannot look artificially attractive.

Run a controlled comparison

Use TextCortex to keep the connection consistent while testing different model families.

  • Use the same representative task set.
  • Record model identifiers and request settings.
  • Apply the same output validation.
  • Measure total usage, accepted outcomes and latency.
  • Repeat when the prompt or model version changes.

Route where the result justifies it

A more capable model can be worthwhile for difficult requests even if its token rate is higher. A smaller or faster option may work well for a simpler feature.

Use the API cost calculator for token arithmetic and the routing guide to keep the production model choices explicit.

Questions

Your next questions, answered.

How is API usage priced?

Use the rates for your selected model and account. Estimate input and output tokens separately and include any applicable cached-token, reasoning-token or other charges. API consumption is distinct from a workspace subscription.

How does model routing work?

Your application selects an available model in each API request, and TextCortex routes the request through its API. You can choose different models for different tasks. Automatic selection and fallback behavior should only be used where documented for your configuration.

Is TextCortex limited to open-source models?

No. The API provides access to proprietary model families such as GPT, Claude and Gemini alongside open-weight families such as Kimi, GLM, DeepSeek and MiMo. Use GET /models to list the identifiers available to your account.

TextCortex AI

One integration. Your choice of models.

Sign up for TextCortex, create an API key and start building with the models your application needs.

Sources and documentation

Product documentation and provider references used for this guide. Reviewed on 5 October 2026.