RigRoute LabsDAENRigRoute ↗
Compare

Accepted systems, side by side

Every system below ran the same standard-ai-v0.1 methodology: byte-identical model files (verified by SHA-256), the same llama.cpp source revision and build number, and the same context depth and repetition count. The runtime was compiled per platform — Metal on macOS, Vulkan on Windows — so the binaries are not identical, and the systems differ in host CPU, drivers and operating system (listed below). The charts show differences between these systems, not isolated component effects — see Methodology.

What this evidence covers
Evidence level
E1 Local observation: a performance run on one host per system, with incomplete host control. Not independently reproduced.
Measured on
27–28 September 2026
Tested configuration
standard-ai-v0.1 · llama.cpp build 11146 · context depth 8192 · 512 prompt / 128 generated tokens · 5 timing repetitions in one session (repetitions, not independent replications)
What was measured
Prompt-processing and generation throughput, in tokens per second.
What PASS means
The run completed without a failure classification (model loaded, every layer offloaded to the targeted accelerator, no CPU fallback, timeout or crash) and was reviewed into the accepted registry. It is not a verdict on answer quality or usability.
Not measured
  • Answer quality (semantic capability)
  • Useful context length
  • Time to first token
  • Interactive responsiveness
  • Thermal behaviour
  • Power draw
  • Largest model that fits
  • Multitasking alongside other work
llama 8B Q4_0

8b-q4

Prompt processing (tok/s)

Apple M1 Pro (32 GB unified memory)
190.39 tok/s
AMD Radeon RX 9070 (16 GB dedicated VRAM)
2511.81 tok/s
AMD Radeon RX 7800 XT (16 GB dedicated VRAM)
763.53 tok/s

Generation (tok/s)

Apple M1 Pro (32 GB unified memory)
27.92 tok/s
AMD Radeon RX 9070 (16 GB dedicated VRAM)
102.31 tok/s
AMD Radeon RX 7800 XT (16 GB dedicated VRAM)
74.81 tok/s

Each chart is scaled to its own largest value. Prompt processing and generation are deliberately not plotted on a shared scale — they differ by roughly an order of magnitude, and one scale would make the generation differences look like no difference at all.

qwen3 14B Q4_K - Medium

14b-q4

Prompt processing (tok/s)

Apple M1 Pro (32 GB unified memory)
97.09 tok/s
AMD Radeon RX 9070 (16 GB dedicated VRAM)
1177.22 tok/s
AMD Radeon RX 7800 XT (16 GB dedicated VRAM)
417.30 tok/s

Generation (tok/s)

Apple M1 Pro (32 GB unified memory)
11.73 tok/s
AMD Radeon RX 9070 (16 GB dedicated VRAM)
56.38 tok/s
AMD Radeon RX 7800 XT (16 GB dedicated VRAM)
43.51 tok/s

Each chart is scaled to its own largest value. Prompt processing and generation are deliberately not plotted on a shared scale — they differ by roughly an order of magnitude, and one scale would make the generation differences look like no difference at all.

Other dimensions
Apple M1 Pro (32 GB unified memory)AMD Radeon RX 9070 (16 GB dedicated VRAM)AMD Radeon RX 7800 XT (16 GB dedicated VRAM)
Host OSmacOSMicrosoft Windows 11 ProMicrosoft Windows 11 Pro
Host CPUApple M1 ProAMD Ryzen 7 9800X3D 8-Core ProcessorAMD Ryzen 7 3700X 8-Core Processor
Benchmark target driverUNKNOWN32.0.31021.500132.0.21030.2001
BackendBLAS,MTLVulkanVulkan
Methodologystandard-ai-v0.1standard-ai-v0.1standard-ai-v0.1
Acceleration evidence33/33, 41/41 layers offloaded to GPU33/33, 41/41 layers offloaded to GPU33/33, 41/41 layers offloaded to GPU
Telemetry (power/temp)UNAVAILABLEUNAVAILABLEUNAVAILABLE
Memorysee provenance →see provenance →see provenance →

This is not a single score. RigRoute Labs does not name an overall winner — the numbers above describe two specific, measured dimensions of one specific workload, nothing broader. These systems also do not share a host platform; each system's host CPU and OS are listed above so that difference stays visible.