Best Local Image Generation Models by VRAM (2026)
Open image models sized for real GPUs, with the text encoder counted in and the licence checked. FLUX, Z-Image, Qwen-Image and SD3.5 compared.
Articles
Reporting and explainers on how AI actually works, who it affects, and what to do about it.
Open image models sized for real GPUs, with the text encoder counted in and the licence checked. FLUX, Z-Image, Qwen-Image and SD3.5 compared.
Anthropic's July 8, 2026 policy trains on consumer chats unless you opt out, and Gemini keeps human-reviewed chats for up to three years.
Local AI on a 6GB GPU: GTX 1660, RTX 2060, RTX 3050 and laptop cards. Real weight sizes for chat, coding, vision, speech, translation and RAG.
Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Eighteen guides, one index.
Apple Silicon has no discrete VRAM, so tier guides mislead Mac owners. The real ceilings are bandwidth and the GPU-usable slice of unified memory.
Apple caps its on-device model at 4096 tokens per session, and Google's own tables put Gemma 4 E2B at 25.0 decode tokens/sec on an iPhone 17 Pro CPU.
Quantization formats compared with real file sizes and bits per weight. Why Q4 does not halve a model, and why low quants generate faster.
An OpenAI agent broke out of a sandbox and into Hugging Face to read test answers. It's the clearest case yet of why reward hacking is getting worse.
An FCC rule puts foreign humanoids, quadrupeds, and wheeled robots on the Covered List. Most US university robotics labs rely on Unitree. The bill is due now.
Alibaba announced 2.4T-parameter Qwen3.8-Max open weights and a 27B sibling that Unsloth says will run in 17GB of RAM or VRAM.
Anthropic disclosed three Claude models breached real customer networks during cybersecurity tests. Existing hacking law was written for humans.
Microsoft's Q4 FY26 earnings show a $3.2B Anthropic gain and $4.96B from OpenAI. Nadella says customers should swap models at will.
Håkon Måløy disclosed a self-replicating prompt injection in Word's Copilot after 144 days of coordination. Two mitigations did not close it.
NVIDIA's Cosmos-H-Dreams runs at 160 fps on one RTX PRO 6000, with weights, code, dataset, and recipe open. Not a robot controller, NVIDIA warns.
Moonshot released the full Kimi K3 weights on July 27, 2026: 2.8T params, 1M context, MXFP4. Read the license before you plan a deployment.
Two back-to-back merges add MiniMax Sparse Attention and a Qwen2.5-VL style vision tower to llama.cpp, but every existing MiniMax-M3 GGUF must be regenerated.
A new industry letter asks Washington to protect open-weight AI while separating legitimate distillation from alleged theft of closed models.
Accomplish's Oren Yomtov chained CVE-2026-46331 to escape Claude Cowork's macOS sandbox in a single short message and reach the host filesystem.
Cisco released Antares-350M and Antares-1B open-weight models for vulnerability localization. They run locally and cost about 172x less than GPT-5.5.
A UK safety test found five frontier models used forbidden shortcuts in cyber evaluations, exposing limits in self-reporting and reasoning traces.
Apple patched a Hide My Email flaw, but aliases created before July 7 may have exposed the real addresses they were meant to conceal.
Prince Canuma's open-source Nativ wraps MLX in a SwiftUI chat app with a localhost API for Claude Code, Codex, and other coding agents.
A Princeton/Chicago ICML study found OpenAI's o3 segregated fictional groups 65% more than humans. Diversity bonuses helped; 'be fair' prompts did not.
California's DROP tool wipes a resident's data from 614 brokers with one form. Here's what's covered, what isn't, and what to try if you don't live there.