System identityMemory architecture What this evidence covers Accepted benchmark runs
Environment
Telemetry availability Provenance & limitations
AMD Radeon RX 7800 XT (16 GB dedicated VRAM)
AMD Radeon RX 7800 XT — unknown
Dedicated VRAM: 16 GB, dedicated. llama.cpp's own Vulkan device query reported 15405 MiB free on this device at model-load time, on an adapter whose Vulkan properties report `uma: 0` -- a dedicated pool, not memory shared with system RAM.
Source: llama.cpp's own device-selection line in both accepted runs' raw stderr (`llama_prepare_model_devices: using device Vulkan0 (AMD Radeon RX 7800 XT) ... - 15405 MiB free`), published verbatim at /data. NOT from Windows WMI's Win32_VideoController.AdapterRAM, a 32-bit field that truncates for cards >=4GB -- which is why the structured result.json's hardware.accelerators[0].dedicatedVramBytes is honestly "UNKNOWN". Note this is free memory measured with a desktop session already resident, so it is a lower bound on the pool, not the pool's total; the accepted RX 9070 reported 15416 MiB free by the identical mechanism and the same llama.cpp build.
Reported device name: Vulkan reports this card as AMD Radeon RX 7800 XT. On the RX 9070 host, Vulkan0 is the integrated AMD Radeon(TM) Graphics device; the benchmark selects Vulkan1, reported as AMD Radeon RX 9070.
Source: Device enumeration and llama_prepare_model_devices selection lines in both systems' accepted stderr.log files, available at /data.
- Evidence level
- E1 Local observation: a performance run on one host per system, with incomplete host control. Not independently reproduced.
- Measured on
- 28 September 2026
- Tested configuration
- standard-ai-v0.1 · llama.cpp build 11146 · context depth 8192 · 512 prompt / 128 generated tokens · 5 timing repetitions in one session (repetitions, not independent replications)
- What was measured
- Prompt-processing and generation throughput, in tokens per second.
- What PASS means
- The run completed without a failure classification (model loaded, every layer offloaded to the targeted accelerator, no CPU fallback, timeout or crash) and was reviewed into the accepted registry. It is not a verdict on answer quality or usability.
- Not measured
- Answer quality (semantic capability)
- Useful context length
- Time to first token
- Interactive responsiveness
- Thermal behaviour
- Power draw
- Largest model that fits
- Multitasking alongside other work
| Model | Context | Prompt processing | Generation | Acceleration evidence | Status |
|---|---|---|---|---|---|
| llama 8B Q4_0 | 8192 | 763.53 tok/s | 74.81 tok/s | 33/33 layers offloaded to GPU | PASS |
| qwen3 14B Q4_K - Medium | 8192 | 417.30 tok/s | 43.51 tok/s | 41/41 layers offloaded to GPU | PASS |
Backend
Vulkan — target device: Vulkan
Methodology version
standard-ai-v0.1
| Host OS | Microsoft Windows 11 Pro 10.0.26200 |
| CPU | AMD Ryzen 7 3700X 8-Core Processor |
| Installed memory | 32 GB |
| llama.cpp / llama-bench | 0.5.0-dev (build 11146, commit 7fe450e19) |
| Backends compiled | CPU,RPC,Vulkan |
PowerUNAVAILABLE
TemperatureUNAVAILABLE
Accelerator memoryUNAVAILABLE
Missing telemetry is shown as UNAVAILABLE, never estimated or substituted. It does not invalidate the recorded timing values, but power and thermal behaviour were not characterised, so their influence on those values cannot be ruled out.
- Model identity (SHA-256, declared type/parameter count) comes from the model file's own bytes and llama-bench's own GGUF parsing — see /data for the exact hashes.
- Acceleration evidence above is parsed directly from llama.cpp's own log output for each run, not assumed from the hardware being present.
- These measurements describe throughput on the specific standard-ai-v0.1 workload only — see Methodology for what this does and does not establish.