Skip to content

Using Coding Models

Select a served coding model and compare it on tasks that matter to your project.

Back to Articles

Model names, available artifacts and serving routes can change independently. Discover the model on your endpoint rather than assuming a fixed THOX Coder 7B, 14B or 32B lineup is installed or supported.

Check availability and licensing

Use the current endpoint catalog and its served status where available. A published weight file, a model catalog entry and a running endpoint are different things.

For a downloaded model, review its own model card and license. The license of a THOX application does not replace the model license.

Check local hardware fit

Record the artifact format, quantization, runtime version, supported accelerator and context limit. Allow memory for context and concurrency as well as model weights.

Parameter count does not determine a universal memory requirement or tokens-per-second rate. Measure the actual artifact on the intended device.

Compare a small set of project tasks

Use synthetic or appropriately approved examples for completion, refactoring, explanations and tests. Keep prompts and output limits consistent across candidates.

Check whether generated code compiles, whether tests exercise the requested behavior, and whether suggestions introduce unsafe dependencies. Do not treat a larger model as a security review guarantee.

Configure the client

Use the exact model ID from your endpoint. Confirm the client supports the required chat or completion format and understand which files it sends.

For a hosted route, verify the account entitlement and key scope before troubleshooting local hardware. An authorization failure is not a memory-capacity problem.

Keep a useful comparison record

Record the requested model, served model where reported, runtime, hardware, processing location, prompt size, latency and outcome. Separate a cold model load from warm inference.

Related guidance