Local AI Inference Speed: What to Expect (September 2026)
Real tokens per second from local LLMs on consumer hardware: how quant and context change speed, why memory bandwidth is the limit, and what helps.
Tag
Real tokens per second from local LLMs on consumer hardware: how quant and context change speed, why memory bandwidth is the limit, and what helps.
VRAM decides what runs at all, bandwidth decides how fast it types. A buying guide by tier across NVIDIA, AMD, Intel, Apple and the used market.
Liquid AI's 2.6B model ships via Setapp for offline Mac agents, while iCloud Private Relay leaks real IPs through passkeys in the same 24 hours.
Apple Silicon has no discrete VRAM, so tier guides mislead Mac owners. The real ceilings are bandwidth and the GPU-usable slice of unified memory.
Prince Canuma's open-source Nativ wraps MLX in a SwiftUI chat app with a localhost API for Claude Code, Codex, and other coding agents.