Skip to main content
Vantaige

Gemma 4 on RTX 4090: Real Tokens/Sec, VRAM & Variant Pick (2026)

A
Aymen B
12 min read
Gemma 4 on RTX 4090: Real Tokens/Sec, VRAM & Variant Pick (2026)

Gemma 4 on RTX 4090: Real Tokens/Sec, VRAM & Variant Pick (2026)

Google released Gemma 4 on April 2, 2026 under a clean Apache 2.0 license, and the question every owner of a 24 GB consumer GPU is asking is: which variant actually runs well, how much VRAM does it eat, and what tokens/sec should I expect on an RTX 4090? This article ships verified per-variant numbers from public benchmarks plus the exact ollama and nvidia-smi commands to reproduce them on your own card. Where a row is not yet measured first-hand, it is flagged honestly rather than guessed at.

TL;DR

  • E4B at INT4 is the daily-driver pick on RTX 4090. tiny VRAM, room for everything else

  • 26B-A4B (MoE) is the sweet spot. quality of a 26B model at speeds near a 4B

  • 31B dense fits at INT4 but only just. expect ~7-8 tok/s and no headroom

  • Apache 2.0 license means no usage carve-outs, unlike Llama or Gemma 3

  • Ollama is the fastest path: ollama pull gemma4:e4b and you are running

· Founder, Vantaige · Published 2026-05-08 · 11 min read · Last reviewed 2026-05-08

Gemma 4 variants explained

Gemma 4 ships in four sizes: E2B, E4B, 26B-A4B, and 31B, all released April 2, 2026 on Google's open-source blog under Apache 2.0. The "E" prefix denotes "effective" parameters. these are the edge variants tuned for laptops and on-device deployment, per the Gemma 4 model overview.

The four variants in plain language:

  • E2B. effective ~2B-parameter edge model, mobile and laptop class

  • E4B. effective ~4B-parameter edge model, sweet spot for 8 GB GPUs

  • 26B-A4B. 26B-total Mixture-of-Experts with ~4B active per token; punches above its size

  • 31B. dense 31B-parameter flagship; highest quality, heaviest to run

Per Google, the 31B variant scores 84.3% on GPQA Diamond and 80.0% on LiveCodeBench v6, nearly doubling the 42.4% Gemma 3 IT 27B score on GPQA. Context window is 256K tokens native across the family. Audio input is supported on the smaller variants, and image and video input is supported across the line, per InfoQ's coverage of the launch.

The license matters more than usual. Gemma 1, 2, and 3 shipped under Google's custom "Gemma Terms of Use" with prohibited-use clauses. Gemma 4 is the first of the family on standard Apache 2.0. no Google-specific carve-outs, no "harmful use" supplementary terms.

Hardware tested: RTX 4090 setup details

The reference rig for the numbers below is a single NVIDIA RTX 4090 with 24 GB GDDR6X, driver 555.x or later, CUDA 12.4, on Ubuntu 24.04 with 64 GB system RAM and an SSD on the model directory. Public benchmarks cited in the table use comparable single-RTX-4090 setups; where you see "author to fill in" the row is reserved for first-hand measurement on this exact rig.

Three things that materially change the numbers:

  • Quantization format. Q4_K_M (a llama.cpp variant of INT4) is the default Ollama ships. INT8 and FP16 are the other tiers. Lower bits = less VRAM, more speed, slightly lower quality.

  • Context length used. Pulling the full 256K context costs serious KV-cache memory. The 31B at full context will not fit on 24 GB without aggressive cache compression.

  • Runtime. Ollama (built on llama.cpp) is easiest. vLLM is faster on batched requests. Gemma 4 E4B currently has a TRITON_ATTN fallback issue that drops it to ~9 tok/s on RTX 4090 in some configurations. known and tracked.

Tokens/sec per variant

The table below merges public single-RTX-4090 measurements with rows reserved for first-hand testing. Honest labeling beats fabricated speed claims.

Variant

Quantization

VRAM used

Tokens/sec (RTX 4090)

Context tested

Source

E2B

Q4_K_M (INT4)

~1.5 GB

untested. author to fill in

8K

Vantaige rig (pending)

E4B

Q4_K_M (INT4)

~3 GB

~9 tok/s (TRITON_ATTN fallback bug); untested at default. author to fill in

8K

turboquant-bench

E4B

INT8

~5 GB

untested. author to fill in

8K

Vantaige rig (pending)

26B-A4B (MoE)

Q4_K_M (INT4)

~14-16 GB

~85 tok/s typical; up to 129 tok/s with TurboQuant 3-bit KV cache at 262K context

8K-262K

Lushbinary, turboquant-bench

31B (dense)

Q4_K_M (INT4)

~18 GB

~7.8 tok/s

8K

Gemma 4 Wiki

31B (dense)

INT8

~31 GB (does NOT fit 24 GB; CPU offload required)

untested. author to fill in (expect <3 tok/s with offload)

8K

Vantaige rig (pending)

31B (dense)

FP16

~62 GB (does NOT fit; needs multi-GPU or offload)

not viable on single RTX 4090

n/a

Knightli VRAM table

The MoE row is the headline. 26B-A4B at INT4 delivers the quality of a 26B-parameter model at roughly 10x the throughput of the 31B dense model, because only ~3.8B parameters activate per token. That is the variant most RTX 4090 owners should default to.

VRAM headroom and what fits alongside

On a 24 GB RTX 4090, headroom matters because nobody runs only the model. your editor, browser, and any background CUDA workload (Stable Diffusion, video transcoding) compete for the same VRAM pool.

Practical headroom math at 8K context:

  • E4B INT4 (~3 GB) leaves ~21 GB free. room for ComfyUI / SDXL alongside

  • 26B-A4B INT4 (~14-16 GB) leaves ~8-10 GB free. fits a small image model

  • 31B INT4 (~18 GB) leaves ~6 GB free. model only, no co-tenants

  • Long context (~64K-256K) adds gigabytes of KV cache. budget accordingly

If you push 26B-A4B to 256K context, plan on losing 4-6 GB to the KV cache unless you compress it (TurboQuant or Ollama's OLLAMA_KV_CACHE_TYPE=q4_0 flag).

Quantization tradeoffs (FP16 vs INT8 vs INT4)

INT4 is the right default for a single RTX 4090, INT8 is the upgrade if your variant fits, and FP16 is reserved for multi-GPU or research workflows. The quality loss from INT4 is small enough that almost no daily workload notices.

Documented quality retention, per the Knightli VRAM table and Unsloth's Gemma 4 guide:

  • FP16. full quality baseline, ~2x VRAM of INT8

  • INT8. preserves ~98-99% of FP16 quality, ~2x VRAM of INT4

  • INT4 (Q4_K_M). preserves ~93-95% of FP16 quality, smallest footprint

  • INT3 / 2-bit. only for memory-starved systems, expect noticeable degradation

For coding, summarization, and agent loops on an RTX 4090, INT4 of the largest variant that fits is almost always better than INT8 of a smaller one. Parameter count beats precision at this tier.

Which variant to pick for which job

Pick by job, not by ego. The 31B is not always better; the smaller MoE is faster and frees VRAM for the other tools you actually use.

  • Daily chat, code completion, low-latency. E4B INT4. Tiny, fast, fits anywhere.

  • General assistant + agent loops. 26B-A4B INT4. Best quality-per-token-per-second on this GPU.

  • Heavy reasoning, single-shot quality. 31B INT4. Slow but the best Gemma 4 quality you can run on 24 GB.

  • Edge / laptop / phone testing. E2B INT4. Designed for it.

  • Long-context retrieval (>64K tokens). 26B-A4B with KV cache compression. The 31B will OOM long before you fill the context.

Gemma 4 vs Llama 4 / Qwen 3 / DeepSeek on same RTX 4090

On a single 24 GB GPU, the realistic 2026 contenders are Gemma 4's 26B-A4B and 31B, Qwen3 30B-A3B, and Llama-derived 30B-class models. The frontier-tier models. DeepSeek V4-Pro (1.6T), Llama 4 Maverick (400B), Qwen3.6 Plus (1M context). are not consumer-runnable per Codersera's 2026 landscape review.

Model

Active params

RTX 4090 tok/s (INT4)

License

Best at

Gemma 4 26B-A4B

~3.8B (MoE)

~85 tok/s

Apache 2.0

Balanced general use, 256K context

Gemma 4 31B

31B (dense)

~7.8 tok/s

Apache 2.0

Single-shot reasoning quality

Qwen3 30B-A3B

3B (MoE)

120-196 tok/s

Apache 2.0

Speed; agent loops

Llama 3.3 70B

70B (dense)

~8 tok/s (CPU offload)

Llama Community License

Knowledge depth (offload-tolerant)

DeepSeek V4-Flash

varies

not single-RTX-4090 viable in full form

Custom

Cloud / multi-GPU coding

The honest summary: Qwen3 30B-A3B is faster on this hardware, but Gemma 4 26B-A4B's Apache 2.0 license, native 256K context, and audio-on-edge make it a stronger pick for commercial product work where licensing and modality breadth matter.

How do you set up walkthrough. 5 minutes with Ollama?

The fastest way to verify any of these numbers on your own RTX 4090 is Ollama. The full path from blank machine to a measured tokens/sec number is five commands.

  1. Install Ollama. On Linux: curl -fsSL https://ollama.com/install.sh | sh. On Windows or Mac, grab the installer from ollama.com.

  2. Pull a Gemma 4 variant. Start small to validate: ollama pull gemma4:e4b. Then go bigger: ollama pull gemma4:26b or ollama pull gemma4:31b. Tags are listed on the official Ollama Gemma 4 page.

  3. Confirm it pulled. ollama list should show the model with its on-disk size. ollama ps shows what is currently loaded.

  4. Run a quick generation with verbose timing. ollama run --verbose gemma4:26b "Write a 200-word explanation of attention mechanisms." Ollama prints eval rate in tokens/sec at the end.

  5. Watch VRAM live. In another terminal: nvidia-smi -l 1. The Memory-Usage column shows real-time consumption while the model runs.

Success looks like: model name in ollama list, single-digit-GB or low-double-digit-GB in nvidia-smi, and an eval rate printed at the end of the generation in the 7-130 tok/s range depending on the variant you picked.

The artifact to ship in your own write-up:

# 1. Capture baseline VRAM
nvidia-smi --query-gpu=memory.used,memory.total --format=csv > before.csv

# 2. Run the benchmark prompt
ollama run --verbose gemma4:26b \
  "Summarize this 2026 release in 5 bullets: Gemma 4 launched April 2 under Apache 2.0..."

# 3. Capture peak VRAM during the run (in second terminal)
nvidia-smi --query-gpu=memory.used --format=csv -l 1 | tee during.csv

# 4. Save the output
ollama show gemma4:26b > model-card.txt

Ship the before.csv, during.csv, and the eval rate line from the verbose output. That triple is the credible first-hand artifact.

FAQ

Can the RTX 4090 actually run Gemma 4 31B?

Yes, but only at INT4 quantization, and only with no other VRAM-hungry workload running. The 31B dense at Q4_K_M consumes roughly 18 GB of the RTX 4090's 24 GB, leaving ~6 GB for the KV cache and OS. Throughput is approximately 7.8 tokens/sec at 8K context per the Gemma 4 Speed Benchmark wiki. INT8 and FP16 of the 31B do not fit a single RTX 4090.

Why is Gemma 4 26B-A4B so much faster than the 31B dense model?

The 26B-A4B is a Mixture-of-Experts model with roughly 3.8B parameters active per token, even though the total model weight is 26B. Ollama only pushes the active-expert subset through the GPU per generated token, so throughput resembles a 4B model while quality reflects the full 26B. The 31B dense model activates all 31B parameters every token, multiplying the per-token compute and memory bandwidth cost.

Is Gemma 4 actually fully open under Apache 2.0?

Yes. The Gemma 4 release on April 2, 2026 is the first in the Gemmaverse to ship under standard Apache 2.0, replacing the custom "Gemma Terms of Use" used in Gemma 1, 2, and 3. Per Google's open-source blog announcement, there are no Google-specific carve-outs and no supplementary "harmful use" clauses. it is the same Apache 2.0 license used by Apache HTTP Server and Kubernetes.

What is the best Gemma 4 variant for an 8 GB GPU?

E4B at Q4_K_M (INT4). It uses approximately 3 GB of VRAM, leaving 5 GB for the KV cache, OS reservation, and any co-tenant workload. Per Unsloth's Gemma 4 documentation, E4B Q4_K_M is the recommended starting variant for any GPU under 16 GB. The 26B-A4B will not fit even quantized on an 8 GB card without aggressive offload.

Does Ollama support Gemma 4 multimodal input on RTX 4090?

Image input is supported across the Gemma 4 family in the underlying weights, and audio input is supported on the smaller E2B and E4B variants. Ollama's official gemma4 library page lists modality support per tag. confirm the specific tag you pull lists vision or audio capability before relying on it. As of May 2026, text-only is the most thoroughly tested path.

Should I use Ollama or vLLM for Gemma 4?

Ollama is faster to set up and easier for single-user interactive use; vLLM is faster for concurrent batched requests. For a developer running interactive prompts on a single RTX 4090, Ollama wins on developer experience and is within 10-20% of vLLM on single-stream throughput. For an inference server fielding multiple simultaneous requests, vLLM's continuous batching pulls ahead. see Allen Kuo's vLLM-vs-Ollama Gemma 4 benchmark for numbers on a 96 GB Blackwell GPU.

References

  1. Google, "Gemma 4: Byte for byte, the most capable open models". blog.google

  2. Google Open Source Blog, "Gemma 4: Expanding the Gemmaverse with Apache 2.0". opensource.googleblog.com

  3. Google AI for Developers, "Gemma 4 model overview". ai.google.dev

  4. InfoQ, "Google Opens Gemma 4 Under Apache 2.0 with Multimodal and Agentic Capabilities". infoq.com

  5. Knightli, "Running Gemma 4 Locally: VRAM Requirements". knightli.com

  6. Gemma 4 Wiki, "Gemma 4 Speed Benchmark". gemma4.wiki

  7. Conor Seabrook, "gemma4-turboquant-bench". github.com

  8. Unsloth Documentation, "Gemma 4 - How to Run Locally". unsloth.ai

  9. Ollama, "gemma4 model library". ollama.com

Get the best new AI tools and guides, weekly

One short email a week. The tools worth trying, the guides worth reading, nothing else.

No spam. Unsubscribe anytime.

A

Aymen B

Contributing writer at Vantaige, covering the AI tools ecosystem.