System identityMemory architecture Accepted benchmark runs
Environment
Telemetry availability Provenance & limitations
Apple M1 Pro (32 GB unified memory)
Apple M1 Pro — unified-memory (shared with system RAM, not a separate VRAM pool)
Memory architecture: Unified memory -- system RAM and GPU-visible memory are the same 32 GB pool, not a separate VRAM pool.
Source: docs/acceptance-m1-pro-v0.1.md; hardware.accelerators[0].memoryArchitecture in both accepted result.json files
| Model | Context | Prompt processing | Generation | Acceleration evidence | Status |
|---|---|---|---|---|---|
| llama 8B Q4_0 | 8192 | 190.39 tok/s | 27.92 tok/s | 33/33 layers offloaded to GPU | PASS |
| qwen3 14B Q4_K - Medium | 8192 | 97.09 tok/s | 11.73 tok/s | 41/41 layers offloaded to GPU | PASS |
Backend
BLAS,MTL — target device: BLAS,MTL
Methodology version
standard-ai-v0.1
| Host OS | macOS 26.6.2 |
| CPU | Apple M1 Pro |
| Installed memory | 32 GB |
| llama.cpp / llama-bench | 0.5.0 (build 11146, commit 7fe450e19) |
| Backends compiled | BLAS,CPU,MTL |
PowerUNAVAILABLE
TemperatureUNAVAILABLE
Accelerator memoryUNAVAILABLE
Missing telemetry is shown as UNAVAILABLE, never estimated or substituted. Throughput measurements are unaffected by missing telemetry.
- Model identity (SHA-256, declared type/parameter count) comes from the model file's own bytes and llama-bench's own GGUF parsing — see /data for the exact hashes.
- Acceleration evidence above is parsed directly from llama.cpp's own log output for each run, not assumed from the hardware being present.
- These measurements describe throughput on the specific standard-ai-v0.1 workload only — see Methodology for what this does and does not establish.