RX 7800 XT vs RX 9070 for Local LLMs: What One AMD Generation Changes
Published 2026-09-28 · Measured under standard-ai-v0.1


Both of these cards carry 16 GB of dedicated VRAM, and both run local language models perfectly well. The question this review answers is narrower and more useful than “which is faster”: when you move one AMD generation forward, what part of local LLM inference actually gets faster?
RigRoute Labs measured both cards under the identical standard-ai-v0.1 methodology — the same llama.cpp build 11146 (commit 7fe450e19), the same Vulkan backend, the same 8192-token context, the same 512 prompt tokens and 128 generated tokens, 5 repetitions, and byte-identical GGUF model files verified by SHA-256. Every model layer ran on the GPU on both cards; neither run fell back to the CPU.
What RigRoute measured
8b-q4
Prompt processing (tok/s)
Generation (tok/s)
RX 9070 measured 3.29× the prompt-processing throughput and 1.37× the generation throughput of the RX 7800 XT on this workload. The two charts are scaled independently on purpose.
14b-q4
Prompt processing (tok/s)
Generation (tok/s)
RX 9070 measured 2.82× the prompt-processing throughput and 1.30× the generation throughput of the RX 7800 XT on this workload. The two charts are scaled independently on purpose.
| Workload | RX 7800 XT | RX 9070 | Uplift |
|---|---|---|---|
| 8b-q4 — prompt processing | 763.53 | 2511.81 | 3.29× |
| 8b-q4 — generation | 74.81 | 102.31 | 1.37× |
| 14b-q4 — prompt processing | 417.30 | 1177.22 | 2.82× |
| 14b-q4 — generation | 43.51 | 56.38 | 1.30× |
The uplift is not evenly distributed
The headline is not that the RX 9070 is faster — it is that the improvement lands almost entirely in one half of the work. Prompt processing improved by 3.29× and 2.82× on the two model sizes. Generation improved by 1.37× and 1.30×. That is a very large gap between two phases of the same benchmark, on the same models, on the same software build.
The pattern is consistent with what the two phases ask of a GPU. Prompt ingestion (prefill) processes the whole input at once and is dominated by dense matrix work, which rewards raw compute throughput. Autoregressive generation produces one token at a time and must stream the model's weights through the memory system for every single token, which makes it far more sensitive to memory movement than to peak arithmetic throughput.
RigRoute has measured that the two phases scale very differently across these two cards. It has not run an experiment that isolates any one architectural feature as the cause, and this review does not claim one. The measurement establishes the pattern; the explanation above is the standard reading of that pattern, not a second result.

Faster is not the same as larger
Both cards have 16 GB of dedicated VRAM, and this is where the distinction matters most for a buying decision. The RX 9070 moves through work faster, dramatically so during prompt processing. It does not hold a larger model. A model that does not fit in 16 GB on an RX 7800 XT does not fit in 16 GB on an RX 9070 either.
Speed and memory envelope are separate axes. If your constraint is “this model is too slow,” a faster card addresses it. If your constraint is “this model does not fit,” only more memory addresses it, and neither of these cards offers more than the other. RigRoute's Decision Engine treats those as different problems because they are.
Test-system differences
The RX 9070 and RX 7800 XT were tested in different host systems: Ryzen 7 9800X3D and Ryzen 7 3700X respectively. Both runs used the same llama.cpp build, identical model files, and full GPU offload, so Labs treats the results as directly comparable for this workload.
The host platforms were not identical, so some small system-level influence cannot be ruled out.
Power and temperature telemetry was not available on either system, so Labs reports throughput only and does not estimate missing telemetry.

If you already own an RX 7800 XT
The measured trade-off is specific. Moving to an RX 9070 buys a large reduction in the time spent ingesting a prompt, and a modest increase in how fast tokens come out afterwards. If your work involves long prompts, large pasted documents, RAG contexts, or repeatedly re-reading a big context, prefill is likely the part you actually wait on, and that is the part that improved most. If you mostly send short prompts and wait on a long answer, generation dominates your wall-clock time, and the measured improvement there is far more modest.
Neither profile makes the upgrade universally right or wrong, and RigRoute is not going to issue a verdict on your behalf. The useful question is which phase you spend your time in — and that is answerable from the numbers above plus how you actually use the machine.
Evidence
- RX 7800 XT hardware page — accepted runs, VRAM provenance and environment.
- RX 9070 hardware page — the same, for the comparison machine.
- Compare — every accepted system side by side.
- Methodology — what standard-ai-v0.1 does and does not establish.
- Data — raw llama-bench output, model hashes and full logs for every run cited here.
Every number in this review is read directly from RigRoute Labs' accepted result registry at build time, not transcribed by hand. Photographs are original RigRoute photography of the two tested cards; web copies have had unit-identifying labels obscured, and the original files were not modified.