RigRoute LabsRigRoute ↗
Review

RX 7800 XT vs RX 9070 for Local LLMs: What One AMD Generation Changes

Published 2026-09-28 · Measured under standard-ai-v0.1

AMD Radeon RX 7800 XT reference-design graphics card, dual-fan cooler, photographed from above on a wooden desk.
The tested Radeon RX 7800 XT, an AMD reference-design board. Photo: RigRoute.
GIGABYTE Radeon RX 9070 GAMING OC graphics card, triple-fan cooler, photographed from above on a wooden desk.
The tested Radeon RX 9070, a GIGABYTE GAMING OC board. Photo: RigRoute.

Both of these cards carry 16 GB of dedicated VRAM, and both run local language models perfectly well. The question this review answers is narrower and more useful than “which is faster”: when you move one AMD generation forward, what part of local LLM inference actually gets faster?

RigRoute Labs measured both cards under the identical standard-ai-v0.1 methodology — the same llama.cpp build 11146 (commit 7fe450e19), the same Vulkan backend, the same 8192-token context, the same 512 prompt tokens and 128 generated tokens, 5 repetitions, and byte-identical GGUF model files verified by SHA-256. Every model layer ran on the GPU on both cards; neither run fell back to the CPU.

What RigRoute measured

llama 8B Q4_0

8b-q4

Prompt processing (tok/s)

AMD Radeon RX 7800 XT (16 GB dedicated VRAM)
763.53 tok/s
AMD Radeon RX 9070 (16 GB dedicated VRAM)
2511.81 tok/s

Generation (tok/s)

AMD Radeon RX 7800 XT (16 GB dedicated VRAM)
74.81 tok/s
AMD Radeon RX 9070 (16 GB dedicated VRAM)
102.31 tok/s

RX 9070 measured 3.29× the prompt-processing throughput and 1.37× the generation throughput of the RX 7800 XT on this workload. The two charts are scaled independently on purpose.

qwen3 14B Q4_K - Medium

14b-q4

Prompt processing (tok/s)

AMD Radeon RX 7800 XT (16 GB dedicated VRAM)
417.30 tok/s
AMD Radeon RX 9070 (16 GB dedicated VRAM)
1177.22 tok/s

Generation (tok/s)

AMD Radeon RX 7800 XT (16 GB dedicated VRAM)
43.51 tok/s
AMD Radeon RX 9070 (16 GB dedicated VRAM)
56.38 tok/s

RX 9070 measured 2.82× the prompt-processing throughput and 1.30× the generation throughput of the RX 7800 XT on this workload. The two charts are scaled independently on purpose.

WorkloadRX 7800 XTRX 9070Uplift
8b-q4 — prompt processing763.532511.813.29×
8b-q4 — generation74.81102.311.37×
14b-q4 — prompt processing417.301177.222.82×
14b-q4 — generation43.5156.381.30×

The uplift is not evenly distributed

The headline is not that the RX 9070 is faster — it is that the improvement lands almost entirely in one half of the work. Prompt processing improved by 3.29× and 2.82× on the two model sizes. Generation improved by 1.37× and 1.30×. That is a very large gap between two phases of the same benchmark, on the same models, on the same software build.

The pattern is consistent with what the two phases ask of a GPU. Prompt ingestion (prefill) processes the whole input at once and is dominated by dense matrix work, which rewards raw compute throughput. Autoregressive generation produces one token at a time and must stream the model's weights through the memory system for every single token, which makes it far more sensitive to memory movement than to peak arithmetic throughput.

RigRoute has measured that the two phases scale very differently across these two cards. It has not run an experiment that isolates any one architectural feature as the cause, and this review does not claim one. The measurement establishes the pattern; the explanation above is the standard reading of that pattern, not a second result.

Backplate of the GIGABYTE Radeon RX 9070 GAMING OC, showing the AMD Radeon and GIGABYTE GAMING branding.
The tested RX 9070's backplate. Its unit serial label has been obscured in this web copy. Photo: RigRoute.

Faster is not the same as larger

Both cards have 16 GB of dedicated VRAM, and this is where the distinction matters most for a buying decision. The RX 9070 moves through work faster, dramatically so during prompt processing. It does not hold a larger model. A model that does not fit in 16 GB on an RX 7800 XT does not fit in 16 GB on an RX 9070 either.

Speed and memory envelope are separate axes. If your constraint is “this model is too slow,” a faster card addresses it. If your constraint is “this model does not fit,” only more memory addresses it, and neither of these cards offers more than the other. RigRoute's Decision Engine treats those as different problems because they are.

Test-system differences

The RX 9070 and RX 7800 XT were tested in different host systems: Ryzen 7 9800X3D and Ryzen 7 3700X respectively. Both runs used the same llama.cpp build, identical model files, and full GPU offload, so Labs treats the results as directly comparable for this workload.

The host platforms were not identical, so some small system-level influence cannot be ruled out.

Power and temperature telemetry was not available on either system, so Labs reports throughput only and does not estimate missing telemetry.

The RX 7800 XT seen from its PCB edge, showing the PCIe connector and the AMD board silkscreen.
The tested RX 7800 XT from its PCB edge. Its unit identifier labels have been obscured in this web copy. Photo: RigRoute.

If you already own an RX 7800 XT

The measured trade-off is specific. Moving to an RX 9070 buys a large reduction in the time spent ingesting a prompt, and a modest increase in how fast tokens come out afterwards. If your work involves long prompts, large pasted documents, RAG contexts, or repeatedly re-reading a big context, prefill is likely the part you actually wait on, and that is the part that improved most. If you mostly send short prompts and wait on a long answer, generation dominates your wall-clock time, and the measured improvement there is far more modest.

Neither profile makes the upgrade universally right or wrong, and RigRoute is not going to issue a verdict on your behalf. The useful question is which phase you spend your time in — and that is answerable from the numbers above plus how you actually use the machine.

Evidence

  • RX 7800 XT hardware page — accepted runs, VRAM provenance and environment.
  • RX 9070 hardware page — the same, for the comparison machine.
  • Compare — every accepted system side by side.
  • Methodology — what standard-ai-v0.1 does and does not establish.
  • Data — raw llama-bench output, model hashes and full logs for every run cited here.

Every number in this review is read directly from RigRoute Labs' accepted result registry at build time, not transcribed by hand. Photographs are original RigRoute photography of the two tested cards; web copies have had unit-identifying labels obscured, and the original files were not modified.