Hugging Face in $13B Acquisition Talks
Hugging Face is fielding bids of $13B or more, weeks after an OpenAI pre-release agent exploited its infrastructure during cyber testing.
Tag
Hugging Face is fielding bids of $13B or more, weeks after an OpenAI pre-release agent exploited its infrastructure during cyber testing.
The monthly token volume at which a self-hosted model becomes cheaper than paying per token to GPT-5, Claude, or Gemini, and the workloads that never get there.
Five stock GGUF quants for Qwen3.8-27B span 9.01 GB to 29.05 GB. A file size is not a VRAM requirement, and a 262,144-token ceiling is not free to fill.
The compute capability, driver and ROCm facts to verify on a second-hand card before it has to run Ollama, llama.cpp or a current CUDA toolkit.
There is no single crossover figure. The cost components on each side, the arithmetic shape, and why every price in it carries a date.
AMD's ROCm support splits by card and by OS, Intel archived both of its own LLM libraries this year, and llama.cpp is the path still standing.
None of the mainstream inference servers ask for a password, and vLLM binds every interface by default. The safe pattern for household serving.
8-bit measures close to lossless. 4-bit ranges from a 0.6% gain to a 59% drop on one benchmark, depending on the model. The figures, fully attributed.
Open weights is not an open licence. Verified Hugging Face licence tags for 46 model repos, and the size and version traps that block shipping.
Qwen's 27B vision-language model is on Hugging Face ungated under Apache 2.0, and AMD says it needs roughly 24GB of VRAM to run comfortably.
Which open models run on CPU and RAM alone, how to size them, why mixture-of-experts helps, and what a machine with no discrete GPU cannot do.
What sits at the top of the field right now, hosted or downloadable, with every licence, price and index score re-read from a primary source today.
VRAM decides what runs at all, bandwidth decides how fast it types. A buying guide by tier across NVIDIA, AMD, Intel, Apple and the used market.
Ollama, LM Studio, Jan, llama.cpp and vLLM compared: licences, engines, OpenAI-compatible APIs, multi-user serving, and which to install first.
Six self-hosted RAG stacks compared: AnythingLLM, Open WebUI, LibreChat, Msty, RAGFlow, Cherry Studio. Licences, embedders, and who can rerank.
Seventeen downloadable-weight models ranked by capability and licence, with every parameter count and licence re-checked against the Hugging Face API.
Liquid AI's 2.6B model ships via Setapp for offline Mac agents, while iCloud Private Relay leaks real IPs through passkeys in the same 24 hours.
Local embedding and reranking models for RAG, sized by VRAM. Real GGUF file sizes, licences, and the runtime gap that stops Ollama reranking.
Open image models sized for real GPUs, with the text encoder counted in and the licence checked. FLUX, Z-Image, Qwen-Image and SD3.5 compared.
Anthropic's July 8, 2026 policy trains on consumer chats unless you opt out, and Gemini keeps human-reviewed chats for up to three years.
Local AI on a 6GB GPU: GTX 1660, RTX 2060, RTX 3050 and laptop cards. Real weight sizes for chat, coding, vision, speech, translation and RAG.
Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Eighteen guides, one index.
Apple Silicon has no discrete VRAM, so tier guides mislead Mac owners. The real ceilings are bandwidth and the GPU-usable slice of unified memory.
Apple caps its on-device model at 4096 tokens per session, and Google's own tables put Gemma 4 E2B at 25.0 decode tokens/sec on an iPhone 17 Pro CPU.