Performance Optimization
Measure the actual model, endpoint and hardware before tuning an inference workload.
Performance depends on the model artifact, runtime, hardware, context, concurrency and network path. This guide does not promise a speedup based only on parameter count or a generic accelerator setting.
Establish a repeatable baseline
Use the same synthetic prompt, selected model and output limit for each comparison. Record the requested and served model where available, and whether the endpoint is local, mesh or hosted.
Measure first-token latency, total duration and generated token count separately. Keep cold-start and warmed-up runs separate.
Check model and context memory
Measure available memory and the runtime reported usage. Reduce context length or concurrency before increasing load. If needed, evaluate a smaller or differently quantized artifact supported by the runtime.
Disk space is not the same as inference memory. Leave room for context caches, runtime overhead and other applications.
Measure the network path
Compare connection latency with inference time. Where supported, test a wired connection to isolate a Wi-Fi issue; there is no fixed Ethernet-versus-Wi-Fi latency difference.
Keep required VPN and security controls in place. For MeshStack, verify the peer connection before measuring a workload routed to that peer.
Test sustained behavior
Follow the hardware operating and ventilation limits. Watch for reported thermal throttling or memory pressure during repeated work, not just a single short request.
Do not change thermal limits or use undocumented accelerator commands to reach a target benchmark.
Record enough to reproduce the result
Include the application/runtime version, hardware, model artifact, context and output settings, concurrency and timing method. Share synthetic inputs and redacted logs with support.
A local benchmark is specific to that configuration. It does not establish a general THOX hardware or hosted-service performance guarantee.