Can my PC run it?
Pick your graphics card and RAM. We'll show which local AI models run, roughly how fast, and the one command to start each in Ollama.
4 of 19 models run on the card, 12 run with RAM helping, 3 won't fit.
Runs on your graphics card 4
Fits in video memory. This is where local AI feels instant.
- Qwen3.5 9BEveryday assistant on 8GB cards · 6GB downloadFast≈101 words/s*
ollama run qwen3.5:9b - Gemma 4 E4BLaptops and small PCs · 3GB downloadFast≈165 words/s*
ollama run gemma4:e4b - DeepSeek-R1 14BStep-by-step reasoning · 9GB downloadFast≈71 words/s*
ollama run deepseek-r1:14b - gpt-oss 20B (MoE)Fast reasoning on a 16GB card · 14GB downloadFast≈101 words/s*
ollama run gpt-oss:20b
Runs with system RAM helping 12
Part of the model sits in RAM. Fine for mixture-of-experts models, slow for big dense ones.
- Qwen3.8 27BLatest Qwen; Ollama’s MTP roughly doubles its speed · 17GB downloadComfortable≈14 words/s*
ollama run qwen3.8:27b - Qwen3.5 27BStrong all-rounder · 16GB downloadComfortable≈18 words/s*
ollama run qwen3.5:27b - Qwen3.5 35B-A3B (MoE)Fast for its size, great with RAM offload · 21GB downloadFast≈43 words/s*
ollama run qwen3.5:35b-a3b - Qwen3.5 122B-A10B (MoE)Big-model quality on a 16GB card plus RAM · 67GB downloadComfortable≈11 words/s*
ollama run qwen3.5:122b-a10b - Qwen3 30B-A3B (MoE)Still a speed favourite · 19GB downloadFast≈43 words/s*
ollama run qwen3:30b - Qwen3-Coder 30B-A3B (MoE)Coding agents · 19GB downloadFast≈43 words/s*
ollama run qwen3-coder:30b - Gemma 4 26B-A4B (MoE)Fast, text and images · 16GB downloadFast≈56 words/s*
ollama run gemma4:26b - Gemma 4 31BGoogle’s best open model · 18GB downloadComfortable≈12 words/s*
ollama run gemma4:31b - DeepSeek-R1 32BReasoning · 20GB downloadUsable≈7.8 words/s*
ollama run deepseek-r1:32b - DeepSeek-R1 70BHeavy reasoning · 43GB downloadPainful≈1.8 words/s*
ollama run deepseek-r1:70b - gpt-oss 120B (MoE)Big model, runs well with RAM offload · 65GB downloadComfortable≈18 words/s*
ollama run gpt-oss:120b - Llama 3.3 70BBig dense model · 43GB downloadPainful≈1.8 words/s*
ollama run llama3.3:70b
Won’t fit 3
Not enough memory between the card and RAM. The gap is shown.
- Qwen3.5 397B-A17B (MoE)Flagship; workstation memory needed · 214GB downloadNeeds 146GB more
ollama run qwen3.5:397b-a17b - Qwen3 235B-A22B (MoE)Near-frontier, needs lots of RAM · 142GB downloadNeeds 74GB more
ollama run qwen3:235b - DeepSeek V4 Flash (MoE)Frontier-class; Q3 fits in about 110GB · 155GB downloadNeeds 87GB more
ollama run deepseek-v4-flash (GGUF)
How this works
Will it fit? We add the model's download size to the memory it needs to hold a conversation (about 8,000 words of context), then compare that with your graphics card's memory and the RAM left after Windows and your apps (we reserve 8GB).
How fast? Generating each word means reading the model's active weights from memory, so speed is set by memory bandwidth: your card's for the part on the GPU, your RAM's for the rest. Mixture-of-experts models (Qwen3 30B-A3B, gpt-oss, Qwen3 235B) only read a fraction of their weights per word, which is why they stay quick even when most of them sit in RAM.
*Estimates, rounded, for one person chatting. Real speed varies with drivers, context length, and settings like how many layers you put on the GPU. "Words per second" here means tokens per second, which is close enough.
Buying for local AI?
Video memory decides what's pleasant; system RAM decides what's possible. A 16GB card plus plenty of fast DDR5 runs mixture-of-experts models far better than its VRAM alone suggests.
See the best graphics cards for local AI and the best RAM for local AI, ranked at today's prices with used options from eBay.
Questions people ask
Can I run Qwen 3.8 locally?
Yes. Qwen3.8 27B needs about 17GB at the default Q4 quality, so it sits entirely on a 24GB card (RTX 3090, RTX 4090) and runs at roughly 45 words a second on a 4090, and close to double that with Ollama’s multi-token prediction switched on. On a 16GB card, pick the Q3 version above or let system RAM take the overflow.
What is the best local LLM for a 16GB graphics card?
For models that live entirely on the card: gpt-oss 20B, Qwen3.5 9B and DeepSeek-R1 14B. With RAM helping, mixture-of-experts models stay quick because they only read a few billion parameters per word: Qwen3.5 35B-A3B, Qwen3-Coder 30B-A3B and Gemma 4 26B-A4B are the ones to try.
How much VRAM do I need to run a local LLM?
At Q4, roughly 0.6GB per billion parameters plus 1 to 3GB for the conversation. A 9B model needs about 6GB, a 27B model about 18GB and a 70B model about 45GB. Anything that doesn’t fit spills into system RAM, which is fine for mixture-of-experts models and slow for big dense ones.
How much RAM do I need for local AI?
32GB if your models fit on the graphics card. 64GB to run gpt-oss 120B or Qwen3.5 122B with a 16GB card. 128GB for Qwen3 235B at Q3, and around 160GB or more for DeepSeek V4 Flash at Q4.
Can I run DeepSeek V4 at home?
V4 Flash, yes, with lots of memory: it has 284B parameters but uses only 13B per word. The 3-bit version needs about 110GB of RAM and VRAM combined; 4-bit needs about 162GB. V4 Pro, at 1.6 trillion parameters, is a server model.
What is the best graphics card for local AI on a budget?
For 16GB, the RX 9060 XT is the cheapest new card, though slower; the RTX 5070 Ti is the sensible NVIDIA pick. For 24GB, a used RTX 3090 is still the cheapest route: it holds 27B to 32B models entirely on the card. AI demand has pushed used 3090s to around £1,000, so check live prices on our ranking.
Do AMD and Intel cards work for local AI?
Yes, with caveats. LM Studio and Ollama run on recent Radeon cards, and llama.cpp’s Vulkan backend covers Intel Arc. NVIDIA is still the least fiddly, and most guides and tools assume CUDA.
Does RAM speed matter for local AI?
Only for the part of the model that doesn’t fit on the graphics card, but then it matters a lot: speed is set by memory bandwidth. Dual-channel DDR5-6000 moves roughly twice as much data as DDR4-3200, so offloaded models run about twice as fast.