THOX Chat API
Connect to the ThoxLLM Cloud chat endpoint using an eligible inference key and a served model.
ThoxLLM Cloud supports an OpenAI-compatible chat-completions contract. Compatibility is endpoint-specific; it does not mean that every OpenAI API, SDK feature or parameter is available.
1. Choose the endpoint
The hosted ThoxLLM Cloud base URL is https://llm.thox.ai/v1. Requests to this endpoint leave your device. Use only data you are authorized to send to the selected service.
For a local runtime, use the base URL and authentication documented for that installation. Do not assume the hosted gateway and a local runtime expose identical capabilities.
2. Obtain an inference key
Sign in and open Dashboard > API Keys. Key creation depends on the account entitlement. Save a newly created secret securely; it is not shown again later.
For the hosted catalog and chat routes, use Authorization: Bearer with the inference key, or the documented x-thox-api-key header. Keep keys out of browser bundles, source control and support screenshots.
3. Discover a model
Set THOX_API_KEY in a trusted terminal environment using your credential-handling process, then request the hosted catalog with curl. This example reads model metadata; it does not run inference.
Choose an exact returned identifier and check its served status where present. A catalog entry can remain unserved, and the request must still satisfy the key scope and account entitlement.
curl --fail-with-body https://llm.thox.ai/v1/models \
-H "Authorization: Bearer $THOX_API_KEY"4. Send a small chat request
The following POSIX-shell example sends a synthetic prompt. Replace MODEL_ID_FROM_CATALOG with an available identifier you are entitled to use. Calling a hosted model can consume usage under your plan.
The website demos have separate server-managed credentials and limits; this example is for your own eligible ThoxLLM Cloud key.
curl --fail-with-body https://llm.thox.ai/v1/chat/completions \
-H "Authorization: Bearer $THOX_API_KEY" \
-H "Content-Type: application/json" \
--data '{"model":"MODEL_ID_FROM_CATALOG","messages":[{"role":"user","content":"Say hello in one sentence."}],"max_tokens":64,"stream":false}'5. Configure a compatible client
In a client that supports a custom OpenAI-compatible chat endpoint, set the base URL, inference key and exact model ID. Store the key using that client credential controls.
Confirm streaming, tool calling and other optional features against the chosen endpoint and model. Do not assume that /v1/completions, /v1/embeddings or the Responses API are available because chat completions works.
Interpret errors and route evidence
For 401 or 403, check key validity, entitlement and scope. For an unavailable model, refresh the catalog. For 429, follow the returned limit guidance instead of retrying continuously.
Record the timestamp and request identifier when available. Inspect returned route and served-model metadata when provided; a requested alias does not establish the model or processing location that served the request.