About Gemma 4 26B-A4B
Gemma 4 26B-A4B is Google's 26B-parameter model released in April 2026, with a 256K-token context window. It is a mixture-of-experts model: all 26B parameters must sit in memory, but only ~4B are active per token, which is what makes it fast for its size. It uses a hybrid-attention design — only a fraction of its layers cache the full context, so long conversations cost far less VRAM than a classic dense model: its KV cache is about 0.5 GB at an 8K context, 5.6 GB at 128K, and 10.9 GB at the full 256K window (FP16 cache).
For most people Q4_K_M is the sweet spot — the most popular quality/size trade-off — while Q8 is near-lossless if you have the memory. Totals above include the KV cache and a realistic framework overhead, so they are what you should expect to see in practice rather than just the download size. Weight sizes are calibrated against real GGUF files — see themethodology.
Frequently asked questions
How much VRAM does Gemma 4 26B-A4B need?
At Q4_K_M with an 8K context, Gemma 4 26B-A4B needs about 19 GB (weights 16 GB + KV cache + overhead). The smallest common hardware that fits is a RX 7900 XTX.
Can an RTX 4090 (24GB) run Gemma 4 26B-A4B?
Yes. An RTX 4090's 24 GB runs Gemma 4 26B-A4B at Q5_K_M (about 22 GB at 8K context).
Can a Mac run Gemma 4 26B-A4B?
Yes — Apple Silicon with 32 GB of unified memory or more (macOS lets the GPU use ~75% of it, ~24 GB) runs Gemma 4 26B-A4B at Q4_K_M.
Related
VRAM calculator for any model ·Token counter
Last updated 2026-08-03. Architecture figures from the model's published config.json; see themethodology.