Define success before picking a model
For extraction, define required fields and how to handle missing values. For coding, define the checks a proposed change must pass. For conversation, define helpfulness, factual support and escalation behavior.
Use a representative input set with ordinary cases and difficult ones. Keep the test set separate from examples used to tune prompts.
Filter the catalog by requirements
Choose models with the input types, output behavior and tools your feature needs. Apply processing-region requirements before testing performance.
Get identifiers from the TextCortex model catalog. Keep version and route details in the experiment so results are not accidentally attributed to a different deployment.
Measure more than a headline score
Run the same task across candidate models using appropriate settings for each.
- Rate valid and useful results against a written rubric.
- Record first-token and completed-response latency.
- Measure token usage and repeated attempts.
- Inspect tool arguments and output parsing failures.
- Evaluate the chosen EU-hosted route where required.
Assign models to features, then revisit the choice
Different features can justify different models. Keep the task-to-model mapping in configuration and monitor outcomes after rollout.
Re-evaluate when prompts, traffic or model versions change. Use the routing guide to keep selection manageable and the pricing calculator to estimate usage.
