Z.ai's ZCode silently uploaded developer code 564 times
Z.ai's ZCode assistant tried 564 times to upload 313MB of workspace code to Alibaba Cloud without consent. Timeline, what was exposed, what to check.
Tag
Z.ai's ZCode assistant tried 564 times to upload 313MB of workspace code to Alibaba Cloud without consent. Timeline, what was exposed, what to check.
Open-weight coding LLMs sized by VRAM: Qwen2.5-Coder, Qwen3-Coder, Devstral Small 2, KAT-Coder, GLM-4.7-Flash, and the trade-offs at each tier.
Seventeen downloadable-weight models ranked by capability and licence, with every parameter count and licence re-checked against the Hugging Face API.
Z.ai's GLM-5.1 beats GPT-5.4 on coding benchmarks under MIT license. Qwen3.6-35B-A3B runs frontier-level code with 3B active params. Microsoft open-sources agent governance for all 10 OWASP risks.
An open-weight model tops the hardest coding benchmark for the first time. A 1-bit LLM runs on a phone. And the protocol connecting AI to everything just passed React's adoption curve.
Google, Alibaba, Meta, Mistral, OpenAI, and Zhipu all ship competitive open-weight models under permissive licenses. The battleground shifts from benchmarks to inference speed on your actual GPU.
Qwen 3.5's MoE models hit S-tier benchmarks, NVIDIA's Nemotron 3 Super delivers 5x throughput gains, and GLM-4.7-Flash brings frontier coding to consumer GPUs. The open-weight race just accelerated.
Forget the marketing - here's how the latest open-weight models actually perform on your GPU, from 8GB budget cards to 24GB workstations.