Skip to content

Running Your First AI Model

Choose a local or hosted endpoint, discover an available model and test a small request.

Back to Articles

Choose where inference should run before choosing a model. A model catalog entry does not mean the model is installed on your device or available to your account.

1. Choose local or hosted processing

For a local workflow, use the runtime installed on a machine you control and confirm its network and provider settings. For hosted ThoxLLM Cloud, the endpoint is https://llm.thox.ai/v1 and requests leave your device.

The local-AI guide explains the local setup path. Website live demos have separate provider disclosures and input rules.

2. Confirm access

Use the authentication method required by your chosen endpoint. Eligible ThoxLLM Cloud accounts can manage inference keys at Dashboard > API Keys. A website login or OAuth application is not an inference key.

3. Discover models on that endpoint

Use the runtime model list for a local installation. For ThoxLLM Cloud, authenticated GET /v1/models returns the catalog for the service and account.

Use an exact model identifier, check its served status where provided, and confirm that your key has the required scope. Planned or unserved entries cannot be treated as ready for inference.

4. Check requirements before downloading

For a local model, verify the source, license, artifact format, quantization and supported runtime. Check disk space separately from memory needed for weights, context and concurrent requests.

Download only after choosing a compatible artifact. This page does not promise a preinstalled model across all THOX devices.

5. Test a short synthetic prompt

Use a harmless example such as asking for a greeting. Set a small output limit if the endpoint supports it. Record the requested and served model, processing location and any error returned.

Review the answer before relying on it. A successful request confirms that request worked; it does not establish general model quality or device performance.

6. Connect your editor

Configure a client that supports your endpoint contract. Confirm what context it sends and where it sends it. Inline completion and file indexing depend on that client.

Related guidance