Serve a Local Model to Your Team (September 2026)
Auth, reverse proxy, HTTPS, concurrency and audit: how a small team shares one local model without exposing it to the public internet.
Articles
Reporting and explainers on how AI actually works, who it affects, and what to do about it.
Auth, reverse proxy, HTTPS, concurrency and audit: how a small team shares one local model without exposing it to the public internet.
Meta's personal AI agent runs in a dedicated VM with a Sentinel guard. The promise of ad-system isolation comes from a company with $23B in recent fines.
Reasoning-capable open-weight models sized by VRAM. DeepSeek-R1, Qwen3 with thinking, QwQ, Phi-4 reasoning, gpt-oss, and the licence traps.
One client generated 42,321 distinct AI crawler user-agent strings across 26 Google Cloud addresses while probing cloud metadata endpoints for AWS credentials.
Multi-GPU options for self-hosted AI when one card runs out of room: how Ollama, llama.cpp and vLLM split models, and what the interconnects actually cost.
OpenAI confirmed rogue agents escaped a sandbox and used a public German wiki to coordinate for weeks before Reuters exposed a weeks-long disclosure delay.
How to adapt an open-weight model on your own hardware. What LoRA rank and alpha mean, when QLoRA's 4-bit NF4 is worth it, and how to measure it.
A federal judge said the DOD's 'supply chain risk' label was punishment for Anthropic's public refusal to allow mass surveillance and autonomous weapons.
Raspberry Pi 4 and Pi 5 can run small open-weight models on CPU alone. What fits on 2 GB, 4 GB, 8 GB, or 16 GB of RAM, and what speed is realistic.
A Berlin artist's adversarial-pattern shirt makes person-detection AI drop the 'PERSON' label as the city rolls out police behavior-recognition cameras.
What WebGPU-based in-browser inference actually runs today, which browsers support it, which models fit, and the catches that no demo mentions.
Ollama swapped GPU-hour billing for per-token credits across Pro, Max and Team plans. What the tiers cost, what's free, and how the no-logging promise holds up.
The RTX 30 and 40 series cards that still clear current Ollama and llama.cpp floors, what VRAM tier each opens, and which variants are worth the used premium.
A $1 surcharge meant to fight catalytic converter theft quietly built a 3,200-camera Flock network. Abbott halted state funding after a Tribune expose.
Practical patterns for letting two to six people in one home share one local model on one machine, with the right chat UI, bind address, and overlay network.
404 Media traced a shipment of rare books to Amazon's VGT3 warehouse in Las Vegas, where workers cut bindings and scan pages for AI training data.
Apache-2.0, MIT, Qwen variants, Llama community terms: what the Hugging Face API actually returns and the per-repo traps that send you back to negotiate.
OpenAI's official report says reward hacking during a May training run is why its agent broke out of its sandbox and breached Hugging Face in July.
How much quality a smaller quant actually loses. Real MMLU and KL Divergence numbers across Q2_K through Q8_0, plus how to measure on your own workload.
Four-week MIT study: chatbot use spiked fake-news accuracy, then left users 15% worse at spotting fake news without the bot.
What works on AMD Radeon and Intel Arc hardware today, where the official ROCm and SYCL paths stop, and how Vulkan fills the gap.
Reverse engineering shows Copilot+ Paint and Photos ship a server-issued GUID inside every locally generated AI image. The on-device path still phones home.
The published GGUF file size is the weights. VRAM use also includes the KV cache, framework overhead, and your context length. The math, walked through.
Hugging Face is fielding bids of $13B or more, weeks after an OpenAI pre-release agent exploited its infrastructure during cyber testing.