Accepted systems, side by side
Every system below ran the same standard-ai-v0.1 methodology: byte-identical model files (verified by SHA-256), the same llama.cpp source revision and build number, and the same context depth and repetition count. The runtime was compiled per platform — Metal on macOS, Vulkan on Windows — so the binaries are not identical, and the systems differ in host CPU, drivers and operating system (listed below). The charts show differences between these systems, not isolated component effects — see Methodology.
- Evidence level
- E1 Local observation: a performance run on one host per system, with incomplete host control. Not independently reproduced.
- Measured on
- 27–28 September 2026
- Tested configuration
- standard-ai-v0.1 · llama.cpp build 11146 · context depth 8192 · 512 prompt / 128 generated tokens · 5 timing repetitions in one session (repetitions, not independent replications)
- What was measured
- Prompt-processing and generation throughput, in tokens per second.
- What PASS means
- The run completed without a failure classification (model loaded, every layer offloaded to the targeted accelerator, no CPU fallback, timeout or crash) and was reviewed into the accepted registry. It is not a verdict on answer quality or usability.
- Not measured
- Answer quality (semantic capability)
- Useful context length
- Time to first token
- Interactive responsiveness
- Thermal behaviour
- Power draw
- Largest model that fits
- Multitasking alongside other work
8b-q4
Prompt processing (tok/s)
Generation (tok/s)
Each chart is scaled to its own largest value. Prompt processing and generation are deliberately not plotted on a shared scale — they differ by roughly an order of magnitude, and one scale would make the generation differences look like no difference at all.
14b-q4
Prompt processing (tok/s)
Generation (tok/s)
Each chart is scaled to its own largest value. Prompt processing and generation are deliberately not plotted on a shared scale — they differ by roughly an order of magnitude, and one scale would make the generation differences look like no difference at all.
| Apple M1 Pro (32 GB unified memory) | AMD Radeon RX 9070 (16 GB dedicated VRAM) | AMD Radeon RX 7800 XT (16 GB dedicated VRAM) | |
|---|---|---|---|
| Host OS | macOS | Microsoft Windows 11 Pro | Microsoft Windows 11 Pro |
| Host CPU | Apple M1 Pro | AMD Ryzen 7 9800X3D 8-Core Processor | AMD Ryzen 7 3700X 8-Core Processor |
| Benchmark target driver | UNKNOWN | 32.0.31021.5001 | 32.0.21030.2001 |
| Backend | BLAS,MTL | Vulkan | Vulkan |
| Methodology | standard-ai-v0.1 | standard-ai-v0.1 | standard-ai-v0.1 |
| Acceleration evidence | 33/33, 41/41 layers offloaded to GPU | 33/33, 41/41 layers offloaded to GPU | 33/33, 41/41 layers offloaded to GPU |
| Telemetry (power/temp) | UNAVAILABLE | UNAVAILABLE | UNAVAILABLE |
| Memory | see provenance → | see provenance → | see provenance → |
This is not a single score. RigRoute Labs does not name an overall winner — the numbers above describe two specific, measured dimensions of one specific workload, nothing broader. These systems also do not share a host platform; each system's host CPU and OS are listed above so that difference stays visible.