MiniMax-M3 Lands in llama.cpp: Sparse Attention and Vision
Two back-to-back merges add MiniMax Sparse Attention and a Qwen2.5-VL style vision tower to llama.cpp, but every existing MiniMax-M3 GGUF must be regenerated.
Tag
Two back-to-back merges add MiniMax Sparse Attention and a Qwen2.5-VL style vision tower to llama.cpp, but every existing MiniMax-M3 GGUF must be regenerated.
Run image analysis, document OCR, and visual reasoning locally. Qwen3-VL, InternVL3.5, Molmo2, and MiniCPM-V tested from 8GB to 32GB VRAM.