Two Radeon systems at 8K context depth: measured prefill and generation differences
Published 2026-09-28 · Updated 2026-09-30 · Measured under standard-ai-v0.1
Updated 2026-09-30: the title and several sentences were narrowed to what these runs measured. The measurements are unchanged. See errata.


Both cards have 16 GB of dedicated VRAM. This review reports what RigRoute Labs measured on the two systems they were tested in: prompt-processing (prefill) and generation throughput for two models, at one context depth. It does not say which card is better, and it does not isolate what one AMD generation changes.
Both systems ran the standard-ai-v0.1 methodology — the same llama.cpp source revision (build 11146, commit 7fe450e19), the Vulkan backend on Windows, a context depth of 8192 tokens, 512 prompt tokens and 128 generated tokens, 5 timing repetitions within one session, and byte-identical GGUF model files verified by SHA-256. llama.cpp's logs record every model layer offloaded to the GPU on both systems; that is layer placement, not proof that no work ran on the CPU.
- Evidence level
- E1 Local observation: a performance run on one host per system, with incomplete host control. Not independently reproduced.
- Measured on
- 27–28 September 2026
- Tested configuration
- standard-ai-v0.1 · llama.cpp build 11146 · context depth 8192 · 512 prompt / 128 generated tokens · 5 timing repetitions in one session (repetitions, not independent replications)
- What was measured
- Prompt-processing and generation throughput, in tokens per second.
- What PASS means
- The run completed without a failure classification (model loaded, every layer offloaded to the targeted accelerator, no CPU fallback, timeout or crash) and was reviewed into the accepted registry. It is not a verdict on answer quality or usability.
- Not measured
- Answer quality (semantic capability)
- Useful context length
- Time to first token
- Interactive responsiveness
- Thermal behaviour
- Power draw
- Largest model that fits
- Multitasking alongside other work
What RigRoute measured
8b-q4
Prompt processing (tok/s)
Generation (tok/s)
The RX 9070 system measured 3.29× the prompt-processing throughput and 1.37× the generation throughput of the RX 7800 XT system on this workload. Each bar is the mean of 5 timing repetitions in one session. The two charts are scaled independently on purpose.
14b-q4
Prompt processing (tok/s)
Generation (tok/s)
The RX 9070 system measured 2.82× the prompt-processing throughput and 1.30× the generation throughput of the RX 7800 XT system on this workload. Each bar is the mean of 5 timing repetitions in one session. The two charts are scaled independently on purpose.
| Workload | RX 7800 XT system (tok/s, mean ± SD) | RX 9070 system (tok/s, mean ± SD) | Ratio (RX 9070 ÷ RX 7800 XT) |
|---|---|---|---|
| 8b-q4 — prompt processing | 763.53 ± 0.81 | 2511.81 ± 18.42 | 3.29× |
| 8b-q4 — generation | 74.81 ± 0.03 | 102.31 ± 0.11 | 1.37× |
| 14b-q4 — prompt processing | 417.30 ± 0.46 | 1177.22 ± 3.14 | 2.82× |
| 14b-q4 — generation | 43.51 ± 0.04 | 56.38 ± 0.03 | 1.30× |
SD is the standard deviation across the 5 timing repetitions of one run. It shows spread within that session only; it is not a confidence interval across sessions or hosts.
The difference is not evenly distributed
The measured difference between the two systems is much larger in prefill than in generation. Prompt processing measured 3.29× and 2.82× on the two model sizes; generation measured 1.37× and 1.30×. The models, model files and llama.cpp source revision were the same on both systems.
The pattern is consistent with what the two phases ask of a GPU. Prompt ingestion (prefill) processes the whole input at once and is dominated by dense matrix work, which rewards raw compute throughput. Autoregressive generation produces one token at a time and must stream the model's weights through the memory system for every single token, which makes it far more sensitive to memory movement than to peak arithmetic throughput.
RigRoute measured that the gap between these two systems is much larger in one phase than in the other. It has not run an experiment that isolates the cause: the GPUs differ, and so do the host CPUs and GPU drivers (see the test-system table below). The explanation above is the standard reading of this kind of pattern, not a second result.

Faster is not the same as larger
Both cards have 16 GB of dedicated VRAM, so higher measured throughput here does not mean room for a larger model. This review did not measure how large a model fits on either card. What fits depends on the quantization, the context length, offload settings, runtime buffers and what else is using the card; at model load llama.cpp reported 15,405 MiB free on the RX 7800 XT and 15,416 MiB on the RX 9070 (see the hardware pages).
Speed and fit are separate questions. If a model is too slow, a faster card can help. If a model does not fit, the options include more memory, a smaller quantization, a shorter context or partial offload, each with its own trade-offs — none of which these runs measured.
Test-system differences
The two cards were tested in different host systems. Both used the same llama.cpp source revision, byte-identical model files and full layer offload, but the host CPU, GPU driver and installed memory differ. The ratios in this review are differences between these two systems; how much of each comes from the GPU and how much from the rest of the system was not measured or bounded.
| RX 7800 XT system | RX 9070 system | |
|---|---|---|
| Host CPU | AMD Ryzen 7 3700X 8-Core Processor | AMD Ryzen 7 9800X3D 8-Core Processor |
| Installed memory | 32 GB | 31 GB |
| Operating system | Microsoft Windows 11 Pro 10.0.26200 | Microsoft Windows 11 Pro 10.0.26200 |
| GPU driver (benchmark target) | 32.0.21030.2001 | 32.0.31021.5001 |
| Recorded benchmark target | Vulkan | AMD Radeon RX 9070 (targeted via -dev Vulkan1; Vulkan) |
| llama.cpp build | 11146 | 11146 |
Power and temperature telemetry was not available on either system, so Labs reports throughput only and does not estimate missing telemetry.

If you already own an RX 7800 XT
These runs measured throughput at one context depth (8192 tokens). They did not measure time to first token, a serving workload or other context lengths. Read conditionally: if your use is dominated by processing long inputs, the prefill ratio is the more relevant number; if it is dominated by generating long answers, the generation ratio is. Whether the same ratios hold on your system, at your context length and settings, was not tested.
Neither profile makes an upgrade universally right or wrong, and RigRoute does not issue a verdict on your behalf. The useful question is which phase you spend your time in.
Evidence
- RX 7800 XT hardware page — accepted runs, VRAM provenance and environment.
- RX 9070 hardware page — the same, for the comparison machine.
- Compare — every accepted system side by side.
- Methodology — what standard-ai-v0.1 does and does not establish.
- Data — raw llama-bench output, model hashes and full logs for every run cited here.
Every number in this review is read directly from RigRoute Labs' accepted result registry at build time, not transcribed by hand. Photographs are original RigRoute photography of the two tested cards; web copies have had unit-identifying labels obscured, and the original files were not modified.