Local AI by VRAM Tier - 8GB | 12GB | 16GB | 24GB | 32GB
Deep dives: Chat | Coding | Translation | Vision | Speech | Agents
Updated July 2026: Gemma 3 gave way to Gemma 4, Qwen 3.5 to Qwen 3.6, and GLM-4.7-Flash joined the high tiers. GPT-OSS 20B is still the 16GB pick. Tiers below are rebuilt around the current models; where a source did not publish a firm VRAM number we say so rather than guess.
A local chatbot that runs on your own hardware, answers your questions without sending data to anyone, and costs nothing per query. That’s the pitch. The reality depends entirely on your GPU.
We track the major open-weight chat models across five VRAM tiers to find what actually works - not just what looks good on a benchmark chart, but what feels responsive in daily use. Every model here runs through Ollama at Q4_K_M quantization unless noted otherwise.
How We Compare
The signals that matter for chat:
- Artificial Analysis Intelligence Index - a blended cross-benchmark score, useful for ranking general capability
- GPQA / AIME - graduate-science and competition-math questions, the ceiling tests for reasoning
- VRAM at Q4 - whether it actually fits your card with room for context
- License - whether you can use it commercially (it varies more than people expect)
- Tokens/second - how fast it talks back
Anything above 15 tok/s feels responsive for chat. Above 30 tok/s is fast enough that you forget it’s local.
8GB VRAM {#8gb}
GPUs: RTX 4060, RTX 3060 8GB, RTX 3070
The entry tier. You’re limited to models around 4-9B parameters, or a small MoE, at Q4. In 2026 that still means genuinely useful conversations.
Candidates
| Model | Params | VRAM (Q4) | License | Note |
|---|---|---|---|---|
| Qwen3.5-9B | 9B | ~6.96 GB @ 32K ctx | Apache 2.0 | 54-58 tok/s, offloads to 200K+ context |
| Gemma 4 E4B | ~4B effective | ~4.98 GB | Apache 2.0 | 128K context, min 8GB / safer 12GB |
| LFM2.5-8B-A1B | 8.3B / 1.5B active (MoE) | under ~6 GB | (see model card) | Built for tool calling, 128K context |
Winner: Qwen3.5-9B
Qwen3.5-9B remains the best all-round 8GB chat model. At Q4 it uses about 6.96GB at 32K context and runs 54-58 tok/s, leaving just enough room for normal conversations - and it can offload to reach 200K+ context if you need it. The Qwen 3.6 generation (below) is stronger, but its smallest members are 27B and up, so they do not fit here; the 3.5 9B is still the right pick at this tier.
Efficiency pick: Gemma 4 E4B. Google’s Gemma 4 uses an effective-parameter design; the E4B build is about 5GB at Q4 with 128K context, leaving generous headroom on an 8GB card. Good when you want to run something else alongside it.
Tool-calling pick: LFM2.5-8B-A1B. Liquid AI’s on-device mixture-of-experts activates only ~1.5B of 8.3B parameters per token, stays under ~6GB, and was built to emit function calls by default - the pick if your 8GB machine needs an agent rather than a chatbot.
The honest take: 8GB chat models handle everyday questions, brainstorming, and simple tasks well. They struggle with long multi-step reasoning and deep world knowledge. Ask one to plan detailed logistics and you’ll see the limits.
For all use cases at this level, see our 8GB VRAM complete guide.
12GB VRAM {#12gb}
GPUs: RTX 3060 12GB, RTX 4070
The comfortable tier. You can run a 9B with lots of context headroom, or reach into 15B dense reasoning models.
Winner: Qwen3.5-9B (with headroom), or Apriel-1.6-15B for reasoning
Running Qwen3.5-9B on 12GB gives you plenty of room for long context and fast responses. When you want stronger step-by-step reasoning, step up to Apriel-1.6-15B-Thinker - a 15B dense model under an MIT license with 131K context and a high Artificial Analysis Intelligence Index (57). Its predecessor ran roughly 10-13GB in independent testing, so on a 12GB card expect it to fit with tight context; drop to 16GB if you want breathing room.
Gemma 4 also ships a 12B dense model listed in Ollama’s library; its exact Q4 VRAM was not published in the sources we checked, so test it on your card rather than trusting a round number.
For all use cases at this level, see our 12GB VRAM complete guide.
16GB VRAM {#16gb}
GPUs: RTX 4060 Ti 16GB, RTX 5060, Intel Arc A770, AMD RX 7800 XT
The sweet spot. 15-20B models fit comfortably, and small MoEs become practical.
Candidates
| Model | Params | VRAM (Q4) | License | Key figure |
|---|---|---|---|---|
| GPT-OSS 20B | 20B (MoE) | ~13.7 GB | Apache 2.0 | AA Intelligence Index 52.1, ~42 tok/s |
| Apriel-1.6-15B-Thinker | 15B dense | ~10-13 GB | MIT | AA Intelligence Index 57 |
| Qwen3.6-35B-A3B | 35B / 3B active (MoE) | ~16.6-22 GB (Q3-Q4) | Apache 2.0 | Newer generation, tight fit |
Winner: GPT-OSS 20B
OpenAI’s open-weight model is still the top 16GB pick in the most recent tier-specific testing we found. At ~13.7GB Q4 it leaves room for context and runs around 42 tok/s. It now has real competition, but nothing at this tier clearly beats it as an all-rounder yet.
Reasoning alternative: Apriel-1.6-15B-Thinker. Its Intelligence Index (57) actually edges GPT-OSS on blended benchmarks, it’s MIT-licensed, and 16GB gives it comfortable context room. Try it if GPT-OSS feels shallow on hard reasoning.
New generation, if it fits: Qwen3.6-35B-A3B is a mixture-of-experts model from the current Qwen line. VRAM figures for it are disputed across sources (roughly 16.6GB at Q3 up to 22GB at Q4) - and because an MoE keeps all experts resident, the lower end is optimistic on a 16GB card. Treat it as a “fits at aggressive quant” option, not a comfortable default, until you’ve tested it.
For all use cases at this level, see our 16GB VRAM complete guide.
24GB VRAM {#24gb}
GPUs: RTX 3090, RTX 4090
This is where local models start competing with cloud APIs on quality. The 27-31B tier fits on a single consumer GPU.
Candidates
| Model | Params | VRAM (Q4) | License | Note |
|---|---|---|---|---|
| Qwen3.6-27B | 27B dense | ~19.8 GB @ 64K ctx | Apache 2.0 | New default at this tier |
| Gemma 4 26B-A4B | 26B / 4B active (MoE) | ~18.7 GB @ 64K | Apache 2.0 | Sustained full-GPU 8K-64K in testing |
| Gemma 4 31B | 31B dense | ~18.32 GB (file) | Apache 2.0 | Min 24GB / safer 32GB for full GPU |
Winner: Qwen3.6-27B
The current-generation dense model is the new 24GB default. At Q4 it uses about 19.8GB at 64K context, leaving headroom on a 24GB card, and it leads the tier on general capability. This is the model that makes “local can replace my chat subscription for most tasks” a defensible claim.
MoE efficiency: Gemma 4 26B-A4B activates only ~4B parameters per token and was the model in that testing that sustained full-GPU execution from 8K to 64K context in both Ollama and llama.cpp - the smoothest high-context experience at this tier.
Dense alternative: Gemma 4 31B fits at Q4 but the source lists 24GB as the minimum and 32GB as the comfortable target for full-GPU use, so expect a tight fit with limited context on a 24GB card.
For all use cases at this level, see our 24GB VRAM complete guide.
32GB VRAM {#32gb}
GPUs: RTX 5090
The frontier. 32GB gives comfortable high-context 27-35B models and MoE headroom.
Candidates
| Model | Params | VRAM | License | Key figure |
|---|---|---|---|---|
| GLM-4.7-Flash | 30B / 3B active (MoE) | ~24 GB (32GB full precision) | MIT | AIME 2025 91.6, GPQA 75.2 |
| Qwen3.6-35B-A3B | 35B / 3B active (MoE) | ~22.1 GB | Apache 2.0 | Full in-VRAM, long context |
| Qwen3.6-27B | 27B dense | ~19.8 GB | Apache 2.0 | Max quant / max context |
Winner: GLM-4.7-Flash
GLM-4.7-Flash is a MIT-licensed 30B-A3B mixture-of-experts model that runs on 24GB and uses the full 32GB comfortably at full precision. Its reasoning scores are the standout at this tier: AIME 2025 at 91.6 and GPQA at 75.2 per its model card. Unsloth’s guide covers the exact quant options.
Same-tier alternative: Qwen3.6-35B-A3B fully resident (~22GB) gives 32GB owners generous context headroom, staying inside the well-supported Qwen tooling.
A note on the giants: the mid-2026 open-weight leaderboard is topped by models like GLM-5.2, DeepSeek V4, and Qwen3.5-397B - but those are hundreds of billions of parameters and do not fit any consumer card. This guide sticks to what a single 8-32GB GPU can actually run.
For all use cases at this level, see our 32GB VRAM complete guide.
Cross-Tier Summary
| Tier | Best Pick | Reasoning Alt | Efficiency Pick |
|---|---|---|---|
| 8GB | Qwen3.5-9B | LFM2.5-8B-A1B (tools) | Gemma 4 E4B |
| 12GB | Qwen3.5-9B | Apriel-1.6-15B-Thinker | Gemma 4 E4B |
| 16GB | GPT-OSS 20B | Apriel-1.6-15B-Thinker | Gemma 4 12B |
| 24GB | Qwen3.6-27B | - | Gemma 4 26B-A4B (MoE) |
| 32GB | GLM-4.7-Flash | Qwen3.6-35B-A3B | Qwen3.6-27B |
Quick Start
Every model listed above runs through Ollama. Install it, pull a model, and start chatting:
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
# Pick your model (choose based on your VRAM tier above)
ollama pull qwen3.5:9b # 8-12GB tier
ollama pull gpt-oss:20b # 16GB tier
ollama pull qwen3.6:27b # 24GB tier
ollama pull glm-4.7-flash # 32GB tier
# Start chatting
ollama run qwen3.5:9b
Model tags change as libraries update - check the Ollama library for the exact current tag. For a ChatGPT-style web interface, pair Ollama with Open WebUI; our full setup guide walks through it.
What These Models Can’t Do
Local chat models have real limits compared to frontier cloud models like Claude, GPT, or Gemini:
- Long, complex instructions - multi-step tasks with many constraints still fray below the 27B tier
- Very recent knowledge - these models were trained months ago; they don’t know about last week
- Creative writing at scale - short-form is fine, novel-length coherence degrades
- Reliable factual accuracy - they hallucinate, especially on obscure topics. Always verify claims
The 27-35B models at this tier come closest to closing the gap. But for now, local chat is best as a private, instant, free alternative for everyday conversations - not a wholesale replacement for frontier models on the hardest tasks.