1. Inventory the current integration
Separate your serverless inference requests from any dedicated deployments or custom-model dependencies. A move to a common model API does not automatically transfer a deployment.
Record the model identifiers, request parameters, tools, streaming behavior and output formats your application relies on. Keep a copy of the existing client configuration for rollback.
2. Map destination model identifiers
Retrieve the TextCortex catalog using GET /models. Match each application task to an available model, including the EU-hosted route where required.
A provider-specific identifier is not portable by default. Set the destination identifier explicitly and review any model-specific reasoning or generation settings.
3. Configure the TextCortex client
Create a TextCortex account and generate an API key in account settings. Install the OpenAI Python package on your server. Store the key in TEXTCORTEX_API_KEY and set TEXTCORTEX_MODEL to an identifier returned by GET /models.
This example keeps credentials in environment variables. The same client configuration works when you choose another supported model. Read the API reference for endpoint-specific parameters.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["TEXTCORTEX_API_KEY"],
base_url="https://api.textcortex.com/v1",
)
for model in client.models.list():
print(model.id)
response = client.chat.completions.create(
model=os.environ["TEXTCORTEX_MODEL"],
messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)
4. Replay a representative request set
Run synthetic or approved test inputs through both integrations. Keep external actions such as emails or purchases disabled during the test.
- Compare answer quality and structured-output validity.
- Check streaming chunks, tool arguments and cancellation.
- Exercise authentication errors, rate limits and timeouts.
- Measure usage and latency under expected concurrency.
5. Switch gradually and retain the previous route
Introduce a small share of eligible traffic after the acceptance checks pass. Track failures and model usage by route so you can identify regressions quickly.
Keep a rollback configuration until the monitoring period is complete. Remove only credentials and services that are no longer used. Use the routing guide to organize future model choices.
