RigRoute LabsDAENRigRoute ↗

← Hardware

System identity

Apple M1 Pro (32 GB unified memory)

Memory architecture

Apple M1 Pro — unified-memory (shared with system RAM, not a separate VRAM pool)

Memory architecture: Unified memory -- system RAM and GPU-visible memory are the same 32 GB pool, not a separate VRAM pool.
Source: docs/acceptance-m1-pro-v0.1.md; hardware.accelerators[0].memoryArchitecture in both accepted result.json files
What this evidence covers
Evidence level
E1 Local observation: a performance run on one host per system, with incomplete host control. Not independently reproduced.
Measured on
27 September 2026
Tested configuration
standard-ai-v0.1 · llama.cpp build 11146 · context depth 8192 · 512 prompt / 128 generated tokens · 5 timing repetitions in one session (repetitions, not independent replications)
What was measured
Prompt-processing and generation throughput, in tokens per second.
What PASS means
The run completed without a failure classification (model loaded, every layer offloaded to the targeted accelerator, no CPU fallback, timeout or crash) and was reviewed into the accepted registry. It is not a verdict on answer quality or usability.
Not measured
  • Answer quality (semantic capability)
  • Useful context length
  • Time to first token
  • Interactive responsiveness
  • Thermal behaviour
  • Power draw
  • Largest model that fits
  • Multitasking alongside other work
Accepted benchmark runs
ModelContextPrompt processingGenerationAcceleration evidenceStatus
llama 8B Q4_08192190.39 tok/s27.92 tok/s33/33 layers offloaded to GPUPASS
qwen3 14B Q4_K - Medium819297.09 tok/s11.73 tok/s41/41 layers offloaded to GPUPASS
Backend

BLAS,MTL — target device: BLAS,MTL

Methodology version

standard-ai-v0.1

Environment
Host OSmacOS 26.6.2
CPUApple M1 Pro
Installed memory32 GB
llama.cpp / llama-bench0.5.0 (build 11146, commit 7fe450e19)
Backends compiledBLAS,CPU,MTL
Telemetry availability
PowerUNAVAILABLE
TemperatureUNAVAILABLE
Accelerator memoryUNAVAILABLE

Missing telemetry is shown as UNAVAILABLE, never estimated or substituted. It does not invalidate the recorded timing values, but power and thermal behaviour were not characterised, so their influence on those values cannot be ruled out.

Provenance & limitations
  • Model identity (SHA-256, declared type/parameter count) comes from the model file's own bytes and llama-bench's own GGUF parsing — see /data for the exact hashes.
  • Acceleration evidence above is parsed directly from llama.cpp's own log output for each run, not assumed from the hardware being present.
  • These measurements describe throughput on the specific standard-ai-v0.1 workload only — see Methodology for what this does and does not establish.