MiniMax-M3 Lands in llama.cpp: Sparse Attention and Vision
Two back-to-back merges add MiniMax Sparse Attention and a Qwen2.5-VL style vision tower to llama.cpp, but every existing MiniMax-M3 GGUF must be regenerated.
Tag
Two back-to-back merges add MiniMax Sparse Attention and a Qwen2.5-VL style vision tower to llama.cpp, but every existing MiniMax-M3 GGUF must be regenerated.
A new industry letter asks Washington to protect open-weight AI while separating legitimate distillation from alleged theft of closed models.
Cisco released Antares-350M and Antares-1B open-weight models for vulnerability localization. They run locally and cost about 172x less than GPT-5.5.
A leaked board email proposed a local GPT-3-class model in 2022. OpenAI's later gpt-oss release shows how that strategy changed.
A Go-based botnet is scanning exposed Ollama, ComfyUI, n8n, Open WebUI, Langflow, and Gradio instances for AWS keys and Kubernetes tokens, QiAnXin XLab says.
Hugging Face disclosed a July 2026 breach run end-to-end by an autonomous agent. The defender was an open-weight model.
Two new open-weight models point in opposite directions: Bonsai brings a 27B model toward phones, while Inkling targets customization at data-center scale.
VIDRAFT_LAB posts Ourbox-35B-JGOS to Hugging Face: 20 tok/s on an 8GB laptop GPU, ~17 tok/s on a CPU-only server, 86.4% on GPQA Diamond.
PhantaField's Sophon PFG-1 whitepaper claims ~95x Nvidia HBM4 bandwidth via monolithic 3D stacking. No silicon yet. Here's why it matters anyway.
Ditch GitHub Copilot's $19/month subscription. Set up Continue.dev with Ollama for private, local AI code completion in VS Code — zero data leaves your machine.
Three weeks away and the leaderboard reshuffled. Kimi K2.6 brings 1T parameters under open weights, Qwen 3.6 stays the consumer GPU king, and DeepSeek V4-Flash proves too hungry for single-card setups.
Ditch GitHub Copilot's $10/month subscription. Set up free, private AI code completion in VS Code using Continue.dev and Ollama — runs entirely on your hardware.
DeepSeek returns with a 1.6T MoE monster under MIT license, Gemma 4's 31B dense model climbs to #3 on Arena AI, and ICLR 2026 papers point to what's next for local inference.
Stop paying Midjourney $30 a month. Set up FLUX on your own hardware with ComfyUI and generate unlimited images with zero content filters and full privacy.
DeepSeek V4 Pro approaches frontier-level performance. Google, Mistral, and Alibaba ship under Apache 2.0. Ollama hits 52 million monthly downloads.
Chat with your own documents locally — no cloud, no subscriptions, no data leaving your machine. Step-by-step setup guide.
DeepSeek V4 matches Claude Opus on coding at 7x lower cost under MIT license. NVIDIA's Nemotron 3 brings hybrid Mamba-Transformer MoE to the open. Google's TurboQuant cuts KV cache memory by 6x with no retraining.
Qwen3.6-27B scores 77.2% on SWE-Bench Verified with a dense architecture that fits on a single RTX 4090. The MoE efficiency narrative just got complicated.
A practical guide to running fully local audio transcription with whisper.cpp and faster-whisper — no API keys, no subscriptions, no data leaving your machine.
Z.ai's GLM-5.1 beats GPT-5.4 on coding benchmarks under MIT license. Qwen3.6-35B-A3B runs frontier-level code with 3B active params. Microsoft open-sources agent governance for all 10 OWASP risks.
A step-by-step guide to running a fully local, private AI code completion setup in VS Code that costs nothing and sends zero data to the cloud.
Alibaba drops Qwen3.6-35B-A3B with 73.4% on SWE-Bench Verified and Apache 2.0 licensing. The 3-billion active parameter class now has three serious contenders.
Google gives Gemma 4 a real open-source license. Mozilla launches Thunderbolt for self-hosted enterprise AI. Arcee AI trains a 400B reasoning model for $20 million. And Milla Jovovich broke GitHub.
After the Perplexity class-action over leaked chats to Meta and Google, here's how to run a citation-grounded AI answer engine on your own hardware with Ollama and SearXNG.