Skip to main content
FLAGSHIP · FREEVRAM Calculator, which GPU runs which model? →
Featured AI Models & Fashion Pack →
Vantaige
SMALL HARDWARE. BIG POSSIBILITIES.

Big ideas.
Small memory footprint.LLM VRAM Calculator: What Models Can Your GPU Run?

Pick a real open model, not just a parameter count, and get a transparent memory estimate, compatible local hardware, and matching GPU rentals.

Your next AI setup
might already be on your desk.

Let’s find out
64 ready models · search the daily cached catalog
01

Configure

Choose what you want to run

LIVE ESTIMATE
Popular, budget-friendly models
Apache-2.0 · Open source: Permissive open-source license (e.g. Apache-2.0 or MIT). Generally allows commercial use, modification, and self-hosting subject to standard copyright and patent notices.
Apache-2.0 · Open source: Permissive open-source license (e.g. Apache-2.0 or MIT). Generally allows commercial use, modification, and self-hosting subject to standard copyright and patent notices.
Apache-2.0 · Open source: Permissive open-source license (e.g. Apache-2.0 or MIT). Generally allows commercial use, modification, and self-hosting subject to standard copyright and patent notices.
Apache-2.0 · Open source: Permissive open-source license (e.g. Apache-2.0 or MIT). Generally allows commercial use, modification, and self-hosting subject to standard copyright and patent notices.
Apache-2.0 · Open source: Permissive open-source license (e.g. Apache-2.0 or MIT). Generally allows commercial use, modification, and self-hosting subject to standard copyright and patent notices.
MIT · Open source: Permissive open-source license (e.g. Apache-2.0 or MIT). Generally allows commercial use, modification, and self-hosting subject to standard copyright and patent notices.
8.2B paramsGQAUp to 128K contextOpen sourceModel card How to inspect specs

Commercial use notice: Permissive open-source license (e.g. Apache-2.0 or MIT). Generally allows commercial use, modification, and self-hosting subject to standard copyright and patent notices.

Quantization glossary

Pick any card to compare VRAM fit, headroom, and estimated speed for the model above.

8K tokens
Advanced assumptions
KV guide
YOUR MEMORY ESTIMATE
CompatibleExact file size

Yes, Qwen3 8B should run on RTX 4090 with 17.3 GiB raw headroom.

6.7 GiB

Qwen3 8B · Q4_K_M · 8K tokens

Required accelerator memory estimateRTX 4090 · 24 GiB capacity
Weights
4.7 GiB
KV cache
1.1 GiB
Runtime
0.9 GiB
Estimated speed:
~140 tok/s(1008 GB/s bandwidth heuristic)

RTX 4060

8 GiB · 1.3 GiB headroom

Buy hardwareCompatible

Affiliate link · Vantaige may earn a commission.

RTX 5060

8 GiB · 1.3 GiB headroom

Buy hardwareCompatible

Affiliate link · Vantaige may earn a commission.

RTX 3060

12 GiB · 5.3 GiB headroom

Buy hardwareCompatible

Affiliate link · Vantaige may earn a commission.

RTX 4070 Super

12 GiB · 5.3 GiB headroom

Buy hardwareCompatible

Affiliate link · Vantaige may earn a commission.

RTX 5070

12 GiB · 5.3 GiB headroom

Buy hardwareCompatible

Affiliate link · Vantaige may earn a commission.

RTX 4060 Ti

16 GiB · 9.3 GiB headroom

Buy hardwareCompatible

Affiliate link · Vantaige may earn a commission.

RTX 4070 Ti Super

16 GiB · 9.3 GiB headroom

Buy hardwareCompatible

Affiliate link · Vantaige may earn a commission.

RTX 4080 Super

16 GiB · 9.3 GiB headroom

Buy hardwareCompatible

Affiliate link · Vantaige may earn a commission.

RTX 5070 Ti

16 GiB · 9.3 GiB headroom

Buy hardwareCompatible

Affiliate link · Vantaige may earn a commission.

RTX 5080

16 GiB · 9.3 GiB headroom

Buy hardwareCompatible

Affiliate link · Vantaige may earn a commission.

RTX 5060 Ti

16 GiB · 9.3 GiB headroom

Buy hardwareCompatible

Affiliate link · Vantaige may earn a commission.

RX 7800 XT

16 GiB · 9.3 GiB headroom

Buy hardwareCompatible

Affiliate link · Vantaige may earn a commission.

Safe-fit results include an OS or accelerator reserve; raw capacity alone is not treated as enough.

THE RIGHT FIT, WITHOUT THE OVERKILL

Meet your model’s next home.

46 local options fit Qwen3 8B at these settings. Start small. Scale when you need to.

Memory fit checked against your settings
MORE ROOM TO EXPERIMENTS

Apple 512 GB

Apple Silicon 512 GB unified memory

Comfortably fits+505 GiB
~111 tok/sdecode est.512 GiB435 GiB GiB usable
Check pricecomplete computer · reference
Check price
MORE ROOM TO EXPERIMENTS

Apple 256 GB

Apple Silicon 256 GB unified memory

Comfortably fits+249 GiB
~111 tok/sdecode est.256 GiB218 GiB GiB usable
Check pricecomplete computer · reference
Check price
MORE ROOM TO EXPERIMENTS

Apple 192 GB

Apple Silicon 192 GB unified memory

Comfortably fits+185 GiB
~111 tok/sdecode est.192 GiB163 GiB GiB usable
Check pricecomplete computer · reference
Check price
PRO / CLOUDS

H200

NVIDIA H200 141 GB

Comfortably fits+134 GiB
~666 tok/sdecode est.141 GiB134 GiB GiB usable
Check priceGPU only · reference
Check price
MORE ROOM TO EXPERIMENTS

Apple 128 GB

Apple Silicon 128 GB unified memory

Comfortably fits+121 GiB
~111 tok/sdecode est.128 GiB109 GiB GiB usable
Amazon · Apple MacBook Pro 16 M4 Max 128GB unified memorycomplete computer · reference
Check price
CUDA READYS

RTX PRO 6000

NVIDIA RTX PRO 6000 Blackwell 96 GB

Comfortably fits+89.3 GiB
~250 tok/sdecode est.96 GiB90.2 GiB GiB usable
Newegg · PNY RTX PRO 6000 Blackwell 96GBGPU only · reference
Check price
MORE ROOM TO EXPERIMENTS

Apple 96 GB

Apple Silicon 96 GB unified memory

Comfortably fits+89.3 GiB
~56 tok/sdecode est.96 GiB81.6 GiB GiB usable
Amazon · Apple Mac Studio M3 Ultra 96GB unified memorycomplete computer · reference
Check price
PRO / CLOUDS

H100 NVL

NVIDIA H100 NVL 94 GB

Comfortably fits+87.3 GiB
~541 tok/sdecode est.94 GiB89.3 GiB GiB usable
Check priceGPU only · reference
Check price
PRO / CLOUDS

A100 80 GB

NVIDIA A100 80 GB

Comfortably fits+73.3 GiB
~283 tok/sdecode est.80 GiB76.0 GiB GiB usable
Check priceGPU only · reference
Check price
PRO / CLOUDS

H100 80 GB

NVIDIA H100 80 GB

Comfortably fits+73.3 GiB
~465 tok/sdecode est.80 GiB76.0 GiB GiB usable
Check priceGPU only · reference
Check price
MORE ROOM TO EXPERIMENTS

Ryzen Mini 64 GB

AMD Ryzen Mini PC 64 GB unified memory

Comfortably fits+57.3 GiB
~12 tok/sdecode est.64 GiB54.4 GiB GiB usable
Amazon · Minisforum UM890 Pro Mini PC 64GB DDR5 OCuLinkcomplete computer · reference
Check price
MORE ROOM TO EXPERIMENTS

Apple 64 GB

Apple Silicon 64 GB unified memory

Comfortably fits+57.3 GiB
~56 tok/sdecode est.64 GiB54.4 GiB GiB usable
Amazon · Apple MacBook Pro 16 M4 Max 64GB unified memorycomplete computer · reference
Check price
CUDA READYS

RTX 6000 Ada

NVIDIA RTX 6000 Ada 48 GB

Comfortably fits+41.3 GiB
~133 tok/sdecode est.48 GiB45.1 GiB GiB usable
Newegg · PNY RTX 6000 Ada Generation 48GBGPU only · reference
Check price
CUDA READYS

RTX A6000

NVIDIA RTX A6000 48 GB

Comfortably fits+41.3 GiB
~107 tok/sdecode est.48 GiB45.1 GiB GiB usable
Amazon · PNY RTX A6000 48GB GDDR6GPU only · reference
Check price
PRO / CLOUDS

A40

NVIDIA A40 48 GB

Comfortably fits+41.3 GiB
~97 tok/sdecode est.48 GiB45.6 GiB GiB usable
Check priceGPU only · reference
Check price
PRO / CLOUDS

L40S

NVIDIA L40S 48 GB

Comfortably fits+41.3 GiB
~120 tok/sdecode est.48 GiB45.6 GiB GiB usable
Check priceGPU only · reference
Check price
MORE ROOM TO EXPERIMENTS

Apple 48 GB

Apple Silicon 48 GB unified memory

Comfortably fits+41.3 GiB
~42 tok/sdecode est.48 GiB40.8 GiB GiB usable
Amazon · Apple Mac mini M4 Pro 48GB unified memorycomplete computer · reference
Check price
PRO / CLOUDS

A100 40 GB

NVIDIA A100 40 GB

Comfortably fits+33.3 GiB
~216 tok/sdecode est.40 GiB38.0 GiB GiB usable
Check priceGPU only · reference
Check price
MORE ROOM TO EXPERIMENTS

Apple 36 GB

Apple Silicon 36 GB unified memory

Comfortably fits+29.3 GiB
~21 tok/sdecode est.36 GiB30.6 GiB GiB usable
Amazon · Apple MacBook Pro 16 M4 36GB unified memorycomplete computer · reference
Check price
CUDA READYS

RTX 5090

NVIDIA RTX 5090 32 GB

Comfortably fits+25.3 GiB
~249 tok/sdecode est.32 GiB29.4 GiB GiB usable
Amazon · ASUS TUF Gaming RTX 5090 32GBGPU only · reference
Check price
CUDA READYS

RTX 5000 Ada

NVIDIA RTX 5000 Ada 32 GB

Comfortably fits+25.3 GiB
~80 tok/sdecode est.32 GiB30.1 GiB GiB usable
Amazon · PNY RTX 5000 Ada Generation 32GBGPU only · reference
Check price
MORE ROOM TO EXPERIMENTS

Ryzen Mini 32 GB

AMD Ryzen Mini PC 32 GB unified memory

Comfortably fits+25.3 GiB
~12 tok/sdecode est.32 GiB27.2 GiB GiB usable
Amazon · Beelink SER8 Mini PC AMD Ryzen 7 32GB DDR5complete computer · reference
Check price
CUDA READYS

RTX 3090

NVIDIA RTX 3090 24 GB

Comfortably fits+17.3 GiB
~130 tok/sdecode est.24 GiB22.1 GiB GiB usable
Amazon · ASUS TUF Gaming OC RTX 3090 24GBGPU only · reference
Check price
CUDA READYS

RTX 4090

NVIDIA RTX 4090 24 GB

Comfortably fits+17.3 GiB
~140 tok/sdecode est.24 GiB22.1 GiB GiB usable
Amazon · ASUS ROG Strix LC RTX 4090 24GBGPU only · reference
Check price
MORE ROOM TO EXPERIMENTS

RX 7900 XTX

AMD Radeon RX 7900 XTX 24 GB

Comfortably fits+17.3 GiB
~133 tok/sdecode est.24 GiB22.1 GiB GiB usable
Amazon · PowerColor Red Devil RX 7900 XTX 24GBGPU only · reference
Check price
PRO / CLOUDS

L4

NVIDIA L4 24 GB

Comfortably fits+17.3 GiB
~42 tok/sdecode est.24 GiB22.8 GiB GiB usable
Check priceGPU only · reference
Check price
PRO / CLOUDS

A10

NVIDIA A10 24 GB

Comfortably fits+17.3 GiB
~83 tok/sdecode est.24 GiB22.8 GiB GiB usable
Check priceGPU only · reference
Check price
PRO / CLOUDS

A30

NVIDIA A30 24 GB

Comfortably fits+17.3 GiB
~130 tok/sdecode est.24 GiB22.8 GiB GiB usable
Check priceGPU only · reference
Check price
MORE ROOM TO EXPERIMENTS

Apple 24 GB

Apple Silicon 24 GB unified memory

Comfortably fits+17.3 GiB
~21 tok/sdecode est.24 GiB20.0 GiB GiB usable
Amazon · Apple Mac mini M4 24GB unified memorycomplete computer · reference
Check price
MORE ROOM TO EXPERIMENTS

RX 7900 XT

AMD Radeon RX 7900 XT 20 GB

Comfortably fits+13.3 GiB
~111 tok/sdecode est.20 GiB18.4 GiB GiB usable
Amazon · Sapphire Pulse RX 7900 XT 20GBGPU only · reference
Check price
CUDA READYS

RTX 4060 Ti

NVIDIA RTX 4060 Ti 16 GB

Comfortably fits+9.3 GiB
~40 tok/sdecode est.16 GiB14.7 GiB GiB usable
Amazon · GIGABYTE Gaming OC RTX 4060 Ti 16GBGPU only · reference
Check price
CUDA READYS

RTX 4070 Ti Super

NVIDIA RTX 4070 Ti Super 16 GB

Comfortably fits+9.3 GiB
~93 tok/sdecode est.16 GiB14.7 GiB GiB usable
Amazon · GIGABYTE Gaming OC RTX 4070 Ti Super 16GBGPU only · reference
Check price
CUDA READYS

RTX 4080 Super

NVIDIA RTX 4080 Super 16 GB

Comfortably fits+9.3 GiB
~102 tok/sdecode est.16 GiB14.7 GiB GiB usable
Amazon · GIGABYTE Windforce OC RTX 4080 Super 16GBGPU only · reference
Check price
CUDA READYS

RTX 5070 Ti

NVIDIA RTX 5070 Ti 16 GB

Comfortably fits+9.3 GiB
~124 tok/sdecode est.16 GiB14.7 GiB GiB usable
Amazon · ASUS Prime RTX 5070 Ti 16GBGPU only · reference
Check price
CUDA READYS

RTX 5080

NVIDIA RTX 5080 16 GB

Comfortably fits+9.3 GiB
~139 tok/sdecode est.16 GiB14.7 GiB GiB usable
Amazon · ASUS Prime RTX 5080 16GBGPU only · reference
Check price
CUDA READYS

RTX 5060 Ti

NVIDIA RTX 5060 Ti 16 GB

Comfortably fits+9.3 GiB
~62 tok/sdecode est.16 GiB14.7 GiB GiB usable
Amazon · ASUS Prime RTX 5060 Ti 16GBGPU only · reference
Check price
MORE ROOM TO EXPERIMENTS

RX 7800 XT

AMD Radeon RX 7800 XT 16 GB

Comfortably fits+9.3 GiB
~87 tok/sdecode est.16 GiB14.7 GiB GiB usable
Amazon · ASRock Steel Legend RX 7800 XT 16GB OCGPU only · reference
Check price
MORE ROOM TO EXPERIMENTS

RX 9060 XT

AMD Radeon RX 9060 XT 16 GB

Comfortably fits+9.3 GiB
~67 tok/sdecode est.16 GiB14.7 GiB GiB usable
Amazon · ASUS Dual RX 9060 XT 16GBGPU only · reference
Check price
MORE ROOM TO EXPERIMENTS

RX 9070

AMD Radeon RX 9070 16 GB

Comfortably fits+9.3 GiB
~89 tok/sdecode est.16 GiB14.7 GiB GiB usable
Amazon · ASRock Challenger RX 9070 16GB OCGPU only · reference
Check price
MORE ROOM TO EXPERIMENTS

RX 9070 XT

AMD Radeon RX 9070 XT 16 GB

Comfortably fits+9.3 GiB
~107 tok/sdecode est.16 GiB14.7 GiB GiB usable
Amazon · PowerColor Red Devil RX 9070 XT 16GBGPU only · reference
Check price
THE EVERYDAY PICKS

Apple 16 GB

Apple Silicon 16 GB unified memory

Comfortably fits+9.3 GiB
~17 tok/sdecode est.16 GiB12.0 GiB GiB usable
Amazon · Apple Mac mini M4 16GB unified memorycomplete computer · reference
Check price
CUDA READYS

RTX 3060

NVIDIA RTX 3060 12 GB

Comfortably fits+5.3 GiB
~50 tok/sdecode est.12 GiB11.0 GiB GiB usable
Amazon · MSI Gaming GeForce RTX 3060 12GB Ventus 2X OCGPU only · reference
Check price
CUDA READYS

RTX 4070 Super

NVIDIA RTX 4070 Super 12 GB

Comfortably fits+5.3 GiB
~70 tok/sdecode est.12 GiB11.0 GiB GiB usable
Amazon · GIGABYTE Windforce OC RTX 4070 Super 12GBGPU only · reference
Check price
CUDA READYS

RTX 5070

NVIDIA RTX 5070 12 GB

Comfortably fits+5.3 GiB
~93 tok/sdecode est.12 GiB11.0 GiB GiB usable
Amazon · GIGABYTE Windforce OC RTX 5070 12GBGPU only · reference
Check price
CUDA READYC

RTX 4060

NVIDIA RTX 4060 8 GB

Fits · limited headroom+1.3 GiB
~38 tok/sdecode est.8 GiB7.0 GiB GiB usable
Amazon · MSI Ventus 2X Black RTX 4060 8GB OCGPU only · reference
Check price
CUDA READYC

RTX 5060

NVIDIA RTX 5060 8 GB

Fits · limited headroom+1.3 GiB
~40 tok/sdecode est.8 GiB7.0 GiB GiB usable
Amazon · GIGABYTE AERO OC RTX 5060 8GBGPU only · reference
Check price
CUDA READYF

RTX 2060

NVIDIA RTX 2060 6 GB

Exceeds safe budget-0.7 GiB
Offload requiredspeed not predicted6 GiB5.0 GiB GiB usable
Check priceGPU only · reference
Check price

Reference prices in USD, not live quotes. Check retailer pricing and availability. Some purchase links may earn us a commission.

From “will it fit?” to your first prompt.

Start with Ollama’s default quantization. Confirm its actual memory with ollama ps.

Speed means estimated decode (generating tokens), not prompt processing. It is based on memory bandwidth, active weight size, and an efficiency assumption—not a benchmark. CPU offloading can be dramatically slower. AMD runtime support varies; check compatibility before buying.

A LITTLE CONTEXT GOES A LONG WAY

Make sense of local AI.

For your first model, your next upgrade, and everything in between.

Knowledge & Deep Dives

Local LLM & VRAM Knowledge Hub

In-depth guides, architectural breakdowns, and hardware sizing manuals to help you build, run, and scale offline AI models with zero guesswork.

50+ Terms8 min reference
Master technical terminology: GGUF vs. Safetensors, K-quants, KV cache, GQA, MoE total vs. active parameters, and Apple unified memory.
Interactive search and category filtering
Practical hardware takeaways for every term
Comprehensive definitions with zero jargon
Alphabetical quick reference index
Read Guide
Interactive Tool6 min read
Understand why context length causes sudden CUDA Out of Memory errors, the exact mathematical formula, and how to reduce memory by 50% to 75%.
Interactive context memory calculator
MHA vs. GQA vs. DeepSeek MLA architectures
FP8 and INT4 cache quantization savings
Multi-turn conversation memory overhead
Read Guide
Config Decoder7 min read
Inspect Hugging Face repositories like a machine learning engineer: decode config.json, spot real context limits, and avoid the MoE parameter trap.
Interactive config.json parameter inspector
GGUF quantization naming conventions decoded
MoE total vs. active parameter memory rules
Context window and RoPE scaling parameters
Read Guide
License Matrix5 min read
The legal and practical differences between true OSI Open Source AI and open-weight models like Llama 3, DeepSeek, Qwen 2.5, and Gemma 2.
Interactive model license comparison matrix
Commercial revenue and user count thresholds
Synthetic distillation and training restrictions
Enterprise legal compliance checklist
Read Guide
Hardware Guide8 min read
Learn why memory bandwidth dictates generation speed, how to pick between Apple Silicon and Nvidia, and the exact VRAM needed for 8B to 70B models.
VRAM capacity tiers (8GB to 128GB+)
Memory bandwidth vs. compute formula
Apple Silicon Mac vs. Nvidia PC comparison
Dual-GPU 70B workstation build blueprints
Read Guide

A few good questions.

Estimates are useful. Knowing their limits is better.

Budget three pieces: model weights, the KV cache for your context length, and a bit of runtime overhead. At Q4, rough floors are about 5 to 8 GiB for 7B/8B, 8 to 12 GiB for 13B/14B, 18 to 24 GiB for 32B, and 42 to 50 GiB for dense 70B, then add context. Use the calculator above with your exact model, quant, and context instead of a napkin estimate.

At Q4_K_M with a moderate context (4K to 8K), most 7B and 8B models land around 5 to 8 GiB total. An 8GB card can run them if you keep context in check. FP16 still wants roughly 14 to 16 GiB of weights before cache. Pick the model and Q4 in the calculator to see the breakdown.

Q4_K_M weights for a dense 70B are already in the ~40 GiB class. With short context and overhead, plan on about 45 to 50 GiB on-GPU. FP16 is ~140 GiB and is data-center territory. A single 24GB card cannot hold a usable Q4 70B without heavy CPU offload.

Not fully on-GPU at Q4 quality. A 4090 or 3090 (24GB) can hybrid-offload 70B, but expect slow single-digit tok/s. Fit is not the same as fast. Practical paths: dual 24GB, a 48GB+ card, a high-memory Mac, a smaller MoE, or a cloud GPU.

Map by VRAM, not TFLOPS. About 8GB covers 7B/8B Q4; 12GB covers 13B/14B; 16 to 24GB covers many 20B to 32B and strong MoE options; ~48GB total is the common consumer path for dense 70B Q4. Always include KV cache in the fit check, then confirm with the calculator.

Yes. After model size, quant is the biggest lever. FP16 is ~2 bytes per parameter; Q4_K_M is roughly 0.5 to 0.6. A 7B drops from ~14 GiB of weights at FP16 to ~4 to 5 GiB at Q4. Weight quant does not shrink the KV cache unless you also quantize the cache.

Yes. The KV cache grows with tokens and concurrent sequences. At 32K or 128K it can rival or exceed the weight size. A model that fits at 4K can OOM at 32K. Set max context to what you actually use, or quantize the KV cache before buying a bigger GPU.

The KV cache stores attention keys and values so the model does not recompute the whole prompt every token. It lives in VRAM, scales with context length (and batch/chats), and is why GGUF file size is not the VRAM you need. GQA/MLA architectures shrink it versus naive multi-head attention.

VRAM is GPU memory and is what makes local inference fast. System RAM holds offloaded layers, the OS, and CPU-only paths. More RAM lets you run a model that does not fit, but PCIe/RAM bandwidth tanks tok/s. Buy VRAM for speed; use RAM as a fallback, not a substitute.

Yes, with GGUF/llama.cpp-style loaders if VRAM plus RAM covers the model. Speed often drops to a few tok/s, which is fine for batch work and poor for snappy chat. Prefer a smaller quant or a model that fits fully on GPU if you care about tokens per second.

Treat 8B Llama 3.x like other 8B dense models: Q4 with modest context usually fits in 8GB with a little care, and is comfortable on 12GB+. FP16 still wants a 16GB-class card. Select the exact Llama 3.x 8B artifact above to see weights versus KV for your context.

Plan ~40 GiB of weights plus cache and overhead, so about 45 to 50 GiB on-GPU for a usable Q4_K_M 70B. Dual 24GB (2x 3090/4090) is the usual consumer build. One 24GB card will offload and feel slow. Confirm with the 70B Q4_K_M row in the calculator.

Enough to learn and to run 7B/8B at Q4 with short context. Tight for modern coding/agent workloads and for long context. If you are buying new, 12 to 16GB+ is a safer floor. 8GB laptops can work with tiny models or hybrid offload, not dense 13B+ at speed.

Yes for 7B to 14B Q4 with room for context, and for some MoE setups with care. Dense 32B and 70B are offload territory. For local LLMs, 12GB almost always beats 8GB of faster compute (3060 12GB over 3070 8GB).

There is no single winner: 24GB (used 3090 or 4090) is the popular single-card sweet spot for ~27B to 32B Q4. Dense 70B wants ~48GB total. High-memory Macs win on capacity, NVIDIA wins on CUDA ecosystem and tok/s. Use the Buy and Rent tabs after you size the model.

Pick the highest quant that still fits with context headroom. Q4_K_M is the community sweet spot for chat and coding. Use Q5/Q6/Q8 when you have spare VRAM. Avoid Q2/Q3 unless nothing else fits. Quality loss shows first on hard reasoning, not on everyday chat.

Yes. Unified memory is one pool, so large quants can load without a discrete VRAM ceiling. Leave headroom for macOS. Rough tiers: 16GB for small models, 32GB for comfortable mid-size, 64GB+ before 70B-class is practical. Bandwidth still sets tok/s; a 64GB Mac can load what a 24GB NVIDIA cannot.

Ollama and llama.cpp still need weights plus KV plus overhead. File size is only the weights. Partial GPU offload (`n-gpu-layers`) changes how much lands in VRAM versus RAM. After you pick a model here, load it and measure with `nvidia-smi` or `ollama ps`, because context growth is real.

The GGUF is weights only. KV cache, activations, the desktop/display reservation, fragmentation, and long prompts all sit on top. Leave headroom, lower context, quantize the KV cache, or pick a smaller quant. Fitting the file size is not the same as fitting a conversation.

Buy when you run many hours a week and want privacy and always-on. Rent for bursts, 48 to 80GB experiments, or to test a 70B setup before you spend on dual-GPU hardware. A hybrid is common: a modest local card for daily 8B to 32B work, cloud when you need more VRAM.

How to size a GPU for a local LLM

Running a large language model locally means fitting two things in GPU memory at once: the model's weights, and a KV cache that grows with how much context is active. Weight size is straightforward, parameter count times bytes per parameter, so quantizing a model down from FP16 to Q4 roughly quarters its footprint. The KV cache is easy to overlook: a model that just barely fits at a 4K context can run out of memory at 32K, since the cache scales with context length. This calculator estimates both and adds a small overhead factor for framework and activation memory, then checks the total against common GPUs from the 12GB RTX 3060 up to an 80GB H100.

Who is this for?

See if your role, business, or workflow matches how people actually use this tool.

Local LLM hobbyist / self-hoster

You found a model on Hugging Face and want to know before downloading gigabytes of weights whether your exact GPU can run it, so you search can I run this LLM on my GPU.

PC builder shopping for an AI GPU

You are about to spend $800-2000 on a graphics card and want proof of which quantized models it will actually handle before buying, searching best GPU for local LLM VRAM.

ML engineer / data scientist sizing a deployment

You need to pick the right quantization level to fit a model on available inference hardware without OOM errors, searching GGUF quantization VRAM calculator.

Indie developer running a local AI coding assistant

You want Continue.dev or a similar VS Code plugin talking to Ollama instead of paying for Copilot, and search how much VRAM for local coding assistant to see what fits on your dev machine.

More free tools

Building a prompt instead of sizing hardware? Try the AI Prompt Builder or browse the full Free AI Tools hub.