RigRoute LabsRigRoute ↗
Benchmarks

Accepted measurements

Every row below is one accepted standard-ai-v0.1 run: prompt-processing and generation throughput at a fixed context depth, measured with llama.cpp's own llama-bench. Exact values are shown here, not just charted — see /data for the full raw evidence behind each one.

HardwareModelContextBackendPrompt processingGenerationStatus
Apple M1 Prollama 8B Q4_08192BLAS,MTL190.3927.92PASS
Apple M1 Proqwen3 14B Q4_K - Medium8192BLAS,MTL97.0911.73PASS
AMD Radeon(TM) Graphicsllama 8B Q4_08192Vulkan2511.81102.31PASS
AMD Radeon(TM) Graphicsqwen3 14B Q4_K - Medium8192Vulkan1177.2256.38PASS

Units: tokens/second (tok/s), mean of 5 repetitions. Individual repetitions, standard deviation, and full raw llama-bench output are available per run on /data.