Skip to content
Intelligibberish
  • News
  • Articles
  • Guides
  • Tools
  • About

Tag

#mlx

← All articles

Local AI Sep 23, 2026

Local AI Inference Speed: What to Expect (September 2026)

Real tokens per second from local LLMs on consumer hardware: how quant and context change speed, why memory bandwidth is the limit, and what helps.

Local AI Aug 27, 2026

What Quantization Costs You in Quality (August 2026)

How much quality a smaller quant actually loses. Real MMLU and KL Divergence numbers across Q2_K through Q8_0, plus how to measure on your own workload.

Local AI Aug 6, 2026

Run LLMs Locally on a Mac: What Actually Fits (September 2026)

Apple Silicon has no discrete VRAM, so tier guides mislead Mac owners. The real ceilings are bandwidth and the GPU-usable slice of unified memory.

Local AI Jul 22, 2026

Nativ: A Mac-Native Local AI App Built by the MLX-VLM Maintainer

Prince Canuma's open-source Nativ wraps MLX in a SwiftUI chat app with a localhost API for Claude Code, Codex, and other coding agents.

Local AI Mar 15, 2026

Apple M5 Max Makes 70B Parameter LLMs a Laptop Reality

The new MacBook Pro with M5 Max can run large language models entirely on-device, keeping your AI interactions private and offline

Intelligibberish

Making sense of AI overwhelm. Independent, self-hosted, no trackers.

News Articles Guides Tools About Disclosure Sponsor Privacy RSS

© 2026 Intelligibberish. Making sense of AI overwhelm.