Two Radeon systems at 8K context depth: measured prefill and generation differences
RigRoute Labs measured prompt-processing and generation throughput for two models on an RX 7800 XT system and an RX 9070 system with llama.cpp Vulkan at 8,192 tokens of context depth. Evidence level E1; the host systems differ.
Read the article →