Apple 512 GB
Apple Silicon 512 GB unified memory
Pick a real open model, not just a parameter count, and get a transparent memory estimate, compatible local hardware, and matching GPU rentals.
Configure
Commercial use notice: Permissive open-source license (e.g. Apache-2.0 or MIT). Generally allows commercial use, modification, and self-hosting subject to standard copyright and patent notices.
Pick any card to compare VRAM fit, headroom, and estimated speed for the model above.
Qwen3 8B · Q4_K_M · 8K tokens
RTX 4060
8 GiB · 1.3 GiB headroom
Affiliate link · Vantaige may earn a commission.
RTX 5060
8 GiB · 1.3 GiB headroom
Affiliate link · Vantaige may earn a commission.
RTX 3060
12 GiB · 5.3 GiB headroom
Affiliate link · Vantaige may earn a commission.
RTX 4070 Super
12 GiB · 5.3 GiB headroom
Affiliate link · Vantaige may earn a commission.
RTX 5070
12 GiB · 5.3 GiB headroom
Affiliate link · Vantaige may earn a commission.
RTX 4060 Ti
16 GiB · 9.3 GiB headroom
Affiliate link · Vantaige may earn a commission.
RTX 4070 Ti Super
16 GiB · 9.3 GiB headroom
Affiliate link · Vantaige may earn a commission.
RTX 4080 Super
16 GiB · 9.3 GiB headroom
Affiliate link · Vantaige may earn a commission.
RTX 5070 Ti
16 GiB · 9.3 GiB headroom
Affiliate link · Vantaige may earn a commission.
RTX 5080
16 GiB · 9.3 GiB headroom
Affiliate link · Vantaige may earn a commission.
RTX 5060 Ti
16 GiB · 9.3 GiB headroom
Affiliate link · Vantaige may earn a commission.
RX 7800 XT
16 GiB · 9.3 GiB headroom
Affiliate link · Vantaige may earn a commission.
Safe-fit results include an OS or accelerator reserve; raw capacity alone is not treated as enough.
46 local options fit Qwen3 8B at these settings. Start small. Scale when you need to.
Apple Silicon 512 GB unified memory
Apple Silicon 256 GB unified memory
Apple Silicon 192 GB unified memory
NVIDIA H200 141 GB
Apple Silicon 128 GB unified memory
NVIDIA RTX PRO 6000 Blackwell 96 GB
Apple Silicon 96 GB unified memory
NVIDIA H100 NVL 94 GB
NVIDIA A100 80 GB
NVIDIA H100 80 GB
AMD Ryzen Mini PC 64 GB unified memory
Apple Silicon 64 GB unified memory
NVIDIA RTX 6000 Ada 48 GB
NVIDIA RTX A6000 48 GB
NVIDIA A40 48 GB
NVIDIA L40S 48 GB
Apple Silicon 48 GB unified memory
NVIDIA A100 40 GB
Apple Silicon 36 GB unified memory
NVIDIA RTX 5090 32 GB
NVIDIA RTX 5000 Ada 32 GB
AMD Ryzen Mini PC 32 GB unified memory
NVIDIA RTX 3090 24 GB
NVIDIA RTX 4090 24 GB
AMD Radeon RX 7900 XTX 24 GB
NVIDIA L4 24 GB
NVIDIA A10 24 GB
NVIDIA A30 24 GB
Apple Silicon 24 GB unified memory
AMD Radeon RX 7900 XT 20 GB
NVIDIA RTX 4060 Ti 16 GB
NVIDIA RTX 4070 Ti Super 16 GB
NVIDIA RTX 4080 Super 16 GB
NVIDIA RTX 5070 Ti 16 GB
NVIDIA RTX 5080 16 GB
NVIDIA RTX 5060 Ti 16 GB
AMD Radeon RX 7800 XT 16 GB
AMD Radeon RX 9060 XT 16 GB
AMD Radeon RX 9070 16 GB
AMD Radeon RX 9070 XT 16 GB
Apple Silicon 16 GB unified memory
NVIDIA RTX 3060 12 GB
NVIDIA RTX 4070 Super 12 GB
NVIDIA RTX 5070 12 GB
NVIDIA RTX 4060 8 GB
NVIDIA RTX 5060 8 GB
NVIDIA RTX 2060 6 GB
Reference prices in USD, not live quotes. Check retailer pricing and availability. Some purchase links may earn us a commission.
Start with Ollama’s default quantization. Confirm its actual memory with ollama ps.
Speed means estimated decode (generating tokens), not prompt processing. It is based on memory bandwidth, active weight size, and an efficiency assumption—not a benchmark. CPU offloading can be dramatically slower. AMD runtime support varies; check compatibility before buying.
For your first model, your next upgrade, and everything in between.
Understand GPU memory, Apple unified memory, and what is worth paying for.
Read the guide UNDERSTAND THE NUMBERSSee why longer conversations need more memory, even with the same model.
Read the guide KNOW YOUR LICENSEKnow what you can download, modify, and use commercially before you build.
Read the guideKnowledge & Deep Dives
In-depth guides, architectural breakdowns, and hardware sizing manuals to help you build, run, and scale offline AI models with zero guesswork.
Estimates are useful. Knowing their limits is better.
Budget three pieces: model weights, the KV cache for your context length, and a bit of runtime overhead. At Q4, rough floors are about 5 to 8 GiB for 7B/8B, 8 to 12 GiB for 13B/14B, 18 to 24 GiB for 32B, and 42 to 50 GiB for dense 70B, then add context. Use the calculator above with your exact model, quant, and context instead of a napkin estimate.
At Q4_K_M with a moderate context (4K to 8K), most 7B and 8B models land around 5 to 8 GiB total. An 8GB card can run them if you keep context in check. FP16 still wants roughly 14 to 16 GiB of weights before cache. Pick the model and Q4 in the calculator to see the breakdown.
Q4_K_M weights for a dense 70B are already in the ~40 GiB class. With short context and overhead, plan on about 45 to 50 GiB on-GPU. FP16 is ~140 GiB and is data-center territory. A single 24GB card cannot hold a usable Q4 70B without heavy CPU offload.
Not fully on-GPU at Q4 quality. A 4090 or 3090 (24GB) can hybrid-offload 70B, but expect slow single-digit tok/s. Fit is not the same as fast. Practical paths: dual 24GB, a 48GB+ card, a high-memory Mac, a smaller MoE, or a cloud GPU.
Map by VRAM, not TFLOPS. About 8GB covers 7B/8B Q4; 12GB covers 13B/14B; 16 to 24GB covers many 20B to 32B and strong MoE options; ~48GB total is the common consumer path for dense 70B Q4. Always include KV cache in the fit check, then confirm with the calculator.
Yes. After model size, quant is the biggest lever. FP16 is ~2 bytes per parameter; Q4_K_M is roughly 0.5 to 0.6. A 7B drops from ~14 GiB of weights at FP16 to ~4 to 5 GiB at Q4. Weight quant does not shrink the KV cache unless you also quantize the cache.
Yes. The KV cache grows with tokens and concurrent sequences. At 32K or 128K it can rival or exceed the weight size. A model that fits at 4K can OOM at 32K. Set max context to what you actually use, or quantize the KV cache before buying a bigger GPU.
The KV cache stores attention keys and values so the model does not recompute the whole prompt every token. It lives in VRAM, scales with context length (and batch/chats), and is why GGUF file size is not the VRAM you need. GQA/MLA architectures shrink it versus naive multi-head attention.
VRAM is GPU memory and is what makes local inference fast. System RAM holds offloaded layers, the OS, and CPU-only paths. More RAM lets you run a model that does not fit, but PCIe/RAM bandwidth tanks tok/s. Buy VRAM for speed; use RAM as a fallback, not a substitute.
Yes, with GGUF/llama.cpp-style loaders if VRAM plus RAM covers the model. Speed often drops to a few tok/s, which is fine for batch work and poor for snappy chat. Prefer a smaller quant or a model that fits fully on GPU if you care about tokens per second.
Treat 8B Llama 3.x like other 8B dense models: Q4 with modest context usually fits in 8GB with a little care, and is comfortable on 12GB+. FP16 still wants a 16GB-class card. Select the exact Llama 3.x 8B artifact above to see weights versus KV for your context.
Plan ~40 GiB of weights plus cache and overhead, so about 45 to 50 GiB on-GPU for a usable Q4_K_M 70B. Dual 24GB (2x 3090/4090) is the usual consumer build. One 24GB card will offload and feel slow. Confirm with the 70B Q4_K_M row in the calculator.
Enough to learn and to run 7B/8B at Q4 with short context. Tight for modern coding/agent workloads and for long context. If you are buying new, 12 to 16GB+ is a safer floor. 8GB laptops can work with tiny models or hybrid offload, not dense 13B+ at speed.
Yes for 7B to 14B Q4 with room for context, and for some MoE setups with care. Dense 32B and 70B are offload territory. For local LLMs, 12GB almost always beats 8GB of faster compute (3060 12GB over 3070 8GB).
There is no single winner: 24GB (used 3090 or 4090) is the popular single-card sweet spot for ~27B to 32B Q4. Dense 70B wants ~48GB total. High-memory Macs win on capacity, NVIDIA wins on CUDA ecosystem and tok/s. Use the Buy and Rent tabs after you size the model.
Pick the highest quant that still fits with context headroom. Q4_K_M is the community sweet spot for chat and coding. Use Q5/Q6/Q8 when you have spare VRAM. Avoid Q2/Q3 unless nothing else fits. Quality loss shows first on hard reasoning, not on everyday chat.
Yes. Unified memory is one pool, so large quants can load without a discrete VRAM ceiling. Leave headroom for macOS. Rough tiers: 16GB for small models, 32GB for comfortable mid-size, 64GB+ before 70B-class is practical. Bandwidth still sets tok/s; a 64GB Mac can load what a 24GB NVIDIA cannot.
Ollama and llama.cpp still need weights plus KV plus overhead. File size is only the weights. Partial GPU offload (`n-gpu-layers`) changes how much lands in VRAM versus RAM. After you pick a model here, load it and measure with `nvidia-smi` or `ollama ps`, because context growth is real.
The GGUF is weights only. KV cache, activations, the desktop/display reservation, fragmentation, and long prompts all sit on top. Leave headroom, lower context, quantize the KV cache, or pick a smaller quant. Fitting the file size is not the same as fitting a conversation.
Buy when you run many hours a week and want privacy and always-on. Rent for bursts, 48 to 80GB experiments, or to test a 70B setup before you spend on dual-GPU hardware. A hybrid is common: a modest local card for daily 8B to 32B work, cloud when you need more VRAM.
Running a large language model locally means fitting two things in GPU memory at once: the model's weights, and a KV cache that grows with how much context is active. Weight size is straightforward, parameter count times bytes per parameter, so quantizing a model down from FP16 to Q4 roughly quarters its footprint. The KV cache is easy to overlook: a model that just barely fits at a 4K context can run out of memory at 32K, since the cache scales with context length. This calculator estimates both and adds a small overhead factor for framework and activation memory, then checks the total against common GPUs from the 12GB RTX 3060 up to an 80GB H100.
See if your role, business, or workflow matches how people actually use this tool.
Local LLM hobbyist / self-hoster
You found a model on Hugging Face and want to know before downloading gigabytes of weights whether your exact GPU can run it, so you search can I run this LLM on my GPU.
PC builder shopping for an AI GPU
You are about to spend $800-2000 on a graphics card and want proof of which quantized models it will actually handle before buying, searching best GPU for local LLM VRAM.
ML engineer / data scientist sizing a deployment
You need to pick the right quantization level to fit a model on available inference hardware without OOM errors, searching GGUF quantization VRAM calculator.
Indie developer running a local AI coding assistant
You want Continue.dev or a similar VS Code plugin talking to Ollama instead of paying for Copilot, and search how much VRAM for local coding assistant to see what fits on your dev machine.
More free tools
Building a prompt instead of sizing hardware? Try the AI Prompt Builder or browse the full Free AI Tools hub.