Skip to content
MoonWorkLab
  • Local LLM & GPU
  • AI-Assisted Dev
  • Hardware & Desktop
  • Self-Hosting
  • About
  • Contact
MoonWorkLab
  • Bar charts: prompt processing 343 vs 180 tokens per second favouring the 2080 Ti, generation 36 vs 47 favouring the 9070 XT
    Local LLM & GPU

    Does CUDA Beat VRAM? I Ran the Same 30B Model on a 2080 Ti and an RX 9070 XT

    ByMoonWork August 14, 2026August 15, 2026

    Same model, same build, same test — one on an RTX 2080 Ti (11GB, CUDA), one on an RX 9070 XT (16GB, Vulkan). The result split: 1.9x prompt processing one way, 1.3x generation the other.

    Read More Does CUDA Beat VRAM? I Ran the Same 30B Model on a 2080 Ti and an RX 9070 XTContinue

  • Line charts: generation rises 21 to 47 tokens per second and prompt processing 71 to 180 as n-cpu-moe drops from 48 to 14
    Local LLM & GPU

    I Ran Qwen3-Coder 30B on an AMD RX 9070 XT — 47 tok/s, No ROCm Required

    ByMoonWork August 13, 2026August 15, 2026

    Real llama-bench numbers for Qwen3-Coder-30B-A3B on an AMD RX 9070 XT (16GB, RDNA4) using llama.cpp with Vulkan — including the full n-cpu-moe sweep that took generation from 21 to 47 tok/s.

    Read More I Ran Qwen3-Coder 30B on an AMD RX 9070 XT — 47 tok/s, No ROCm RequiredContinue

Page navigation

Previous PagePrevious 1 2 3

© 2026 MoonWork Lab — every number measured on hardware I own.
About · Contact · Privacy · Affiliate Disclosure

  • Local LLM & GPU
  • AI-Assisted Dev
  • Hardware & Desktop
  • Self-Hosting
  • About
  • Contact