BenchmarksWhat this evidence covers
Accepted measurements
Every row below is one accepted standard-ai-v0.1 run: prompt-processing and generation throughput at a fixed context depth, measured with llama.cpp's own llama-bench. Exact values are shown here, not just charted — see /data for the full raw evidence behind each one.
- Evidence level
- E1 Local observation: a performance run on one host per system, with incomplete host control. Not independently reproduced.
- Measured on
- 27–28 September 2026
- Tested configuration
- standard-ai-v0.1 · llama.cpp build 11146 · context depth 8192 · 512 prompt / 128 generated tokens · 5 timing repetitions in one session (repetitions, not independent replications)
- What was measured
- Prompt-processing and generation throughput, in tokens per second.
- What PASS means
- The run completed without a failure classification (model loaded, every layer offloaded to the targeted accelerator, no CPU fallback, timeout or crash) and was reviewed into the accepted registry. It is not a verdict on answer quality or usability.
- Not measured
- Answer quality (semantic capability)
- Useful context length
- Time to first token
- Interactive responsiveness
- Thermal behaviour
- Power draw
- Largest model that fits
- Multitasking alongside other work
| Hardware | Model | Context | Backend | Prompt processing (tok/s) | Generation (tok/s) | Status |
|---|---|---|---|---|---|---|
| Apple M1 Pro (32 GB unified memory) | llama 8B Q4_0 | 8192 | BLAS,MTL | 190.39 | 27.92 | PASS |
| Apple M1 Pro (32 GB unified memory) | qwen3 14B Q4_K - Medium | 8192 | BLAS,MTL | 97.09 | 11.73 | PASS |
| AMD Radeon RX 9070 (16 GB dedicated VRAM) | llama 8B Q4_0 | 8192 | Vulkan | 2511.81 | 102.31 | PASS |
| AMD Radeon RX 9070 (16 GB dedicated VRAM) | qwen3 14B Q4_K - Medium | 8192 | Vulkan | 1177.22 | 56.38 | PASS |
| AMD Radeon RX 7800 XT (16 GB dedicated VRAM) | llama 8B Q4_0 | 8192 | Vulkan | 763.53 | 74.81 | PASS |
| AMD Radeon RX 7800 XT (16 GB dedicated VRAM) | qwen3 14B Q4_K - Medium | 8192 | Vulkan | 417.30 | 43.51 | PASS |
Units: tokens/second (tok/s), mean of 5 timing repetitions within one session. Individual repetitions, standard deviation, and full raw llama-bench output are available per run on /data.