What It Takes to Run Qwen3.8-27B on Your Own Hardware
Five stock GGUF quants for Qwen3.8-27B span 9.01 GB to 29.05 GB. A file size is not a VRAM requirement, and a 262,144-token ceiling is not free to fill.
Category
Five stock GGUF quants for Qwen3.8-27B span 9.01 GB to 29.05 GB. A file size is not a VRAM requirement, and a 262,144-token ceiling is not free to fill.
The compute capability, driver and ROCm facts to verify on a second-hand card before it has to run Ollama, llama.cpp or a current CUDA toolkit.
There is no single crossover figure. The cost components on each side, the arithmetic shape, and why every price in it carries a date.
AMD's ROCm support splits by card and by OS, Intel archived both of its own LLM libraries this year, and llama.cpp is the path still standing.
None of the mainstream inference servers ask for a password, and vLLM binds every interface by default. The safe pattern for household serving.
8-bit measures close to lossless. 4-bit ranges from a 0.6% gain to a 59% drop on one benchmark, depending on the model. The figures, fully attributed.
Open weights is not an open licence. Verified Hugging Face licence tags for 46 model repos, and the size and version traps that block shipping.
Qwen's 27B vision-language model is on Hugging Face ungated under Apache 2.0, and AMD says it needs roughly 24GB of VRAM to run comfortably.
Which open models run on CPU and RAM alone, how to size them, why mixture-of-experts helps, and what a machine with no discrete GPU cannot do.
What sits at the top of the field right now, hosted or downloadable, with every licence, price and index score re-read from a primary source today.
VRAM decides what runs at all, bandwidth decides how fast it types. A buying guide by tier across NVIDIA, AMD, Intel, Apple and the used market.
Ollama, LM Studio, Jan, llama.cpp and vLLM compared: licences, engines, OpenAI-compatible APIs, multi-user serving, and which to install first.
Six self-hosted RAG stacks compared: AnythingLLM, Open WebUI, LibreChat, Msty, RAGFlow, Cherry Studio. Licences, embedders, and who can rerank.
Seventeen downloadable-weight models ranked by capability and licence, with every parameter count and licence re-checked against the Hugging Face API.
Liquid AI's 2.6B model ships via Setapp for offline Mac agents, while iCloud Private Relay leaks real IPs through passkeys in the same 24 hours.
Local embedding and reranking models for RAG, sized by VRAM. Real GGUF file sizes, licences, and the runtime gap that stops Ollama reranking.
Open image models sized for real GPUs, with the text encoder counted in and the licence checked. FLUX, Z-Image, Qwen-Image and SD3.5 compared.
Local AI on a 6GB GPU: GTX 1660, RTX 2060, RTX 3050 and laptop cards. Real weight sizes for chat, coding, vision, speech, translation and RAG.
Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Eighteen guides, one index.
Apple Silicon has no discrete VRAM, so tier guides mislead Mac owners. The real ceilings are bandwidth and the GPU-usable slice of unified memory.
Apple caps its on-device model at 4096 tokens per session, and Google's own tables put Gemma 4 E2B at 25.0 decode tokens/sec on an iPhone 17 Pro CPU.
Quantization formats compared with real file sizes and bits per weight. Why Q4 does not halve a model, and why low quants generate faster.
Alibaba announced 2.4T-parameter Qwen3.8-Max open weights and a 27B sibling that Unsloth says will run in 17GB of RAM or VRAM.
NVIDIA's Cosmos-H-Dreams runs at 160 fps on one RTX PRO 6000, with weights, code, dataset, and recipe open. Not a robot controller, NVIDIA warns.