Top Stories
Thinking Machines releases 975B Inkling under Apache 2.0, as a base for fine-tuning
Mira Murati’s Thinking Machines Lab released Inkling on July 16: a mixture-of-experts transformer with 975B total parameters and 41B active per forward pass. The license is Apache 2.0, the training data was 45 trillion tokens of text, images, audio, and video, and the model accepts all four natively with a 1M-token context window.
Thinking Machines is explicitly not positioning Inkling as the strongest model available. The lab’s framing is that it is a useful open-weights base for customization, with multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning. Day-zero inference support shipped in transformers 5.14.0+, SGLang, vLLM, and llama.cpp. Quantized checkpoints include a BF16 release needing about 2 TB of VRAM, an NVFP4 build for Blackwell GPUs at around 600 GB, and a 1-bit GGUF community quant from Unsloth that cuts VRAM roughly twenty-fold.
The release lands within a day of Moonshot’s Kimi K3 pricing reset and a refreshed State of Open Source AI report, and the Inkling model card itself notes the trained data contains “content that may be subject to intellectual property protection.” For local-AI users, the practical question is whether the smaller Inkling-Small (276B / 12B active) lands at a workable size once the weights ship.
Patreon stops asking and starts blocking AI scrapers, through Cloudflare
Patreon announced on July 17 that it has moved from a passive robots.txt stance to active AI bot blocking through Cloudflare’s AI Crawl Control. A Patreon blog post cited by TechCrunch said weekly scraper attempts during testing dropped from “thousands of attempts to zero,” while indexing bots that drive traffic back to creators remain allowed; Patreon’s Drew Rowny separately said creators deserve “a meaningful say in how their work is used by AI companies.”
This is the first large creator platform to flip to network-level blocking on its own surface. Cloudflare’s earlier moves (bot-blocking tools in July 2024, a Pay Per Crawl marketplace in July 2025, and a July 2026 policy blocking mixed-use crawlers on ad-hosting pages) sat at the infrastructure layer, and the Patreon-Cloudflare partnership we covered earlier this month laid the policy groundwork. Patreon is the first user of that infrastructure at significant scale, and the test is whether other creator CMSes and publishers follow.
The wraparound of Cloudflare-style blocking is also what the new Decoy Font is built against (see Quick Hits): the playwright side of the AI scraping fight is now a working stack, not just server logs.
ICE awards Thomson Reuters a five-year, $125M contract for CLEAR to investigate “voter fraud”
A procurement record reviewed by 404 Media shows the Department of Homeland Security will pay Thomson Reuters Special Services (TRSS) $25M a year, totaling $125M over five years, for access to the CLEAR database. The data scope covers names, addresses, Social Security numbers, ethnicity, social media posts, and geolocation, plus license plate data, property records, and credit header data. The stated statutory hook is “the presidential mandate of the identification of Voters Fraud, Immigration Fraud and National Security.”
The story is the clearest public marker so far of how an AI-enabled identity broker gets fused into federal enforcement with no public dataset schedule - the data-broker layer on top of the ICE detention and compute infrastructure we mapped in March. Thomson Reuters stated CLEAR cannot be used to identify “noncriminal immigrants or undocumented individuals with the intention of deportation solely on the basis of the individual’s immigration status,” and said immigration status is not a CLEAR search field. The contract is worth watching because the procurement document also described TRSS as the only contractor able to provide “continuous monitoring and alert service for millions of individuals,” narrowing the field for any future audit.
Anthropic pins Claude Fable 5 to Max and Team Premium permanently starting July 20
Anthropic will make Fable 5 a permanent fixture in its top subscription tiers from July 20, 2026, per a Twitter update from @claudeai. Fable 5 ships at “50% of limits” inside Max and Team Premium; Pro and Team Standard users get it through usage credits plus a one-time $100 credit.
This reverses the earlier plan to drop Fable 5 from subscriptions and route it only through API pricing. The reversal follows the broad rollouts of OpenAI’s GPT-5.6 Sol and Moonshot’s Kimi K3, and is the clearest signal so far of where Anthropic is placing its compute bets for the rest of the year. Subscribers no longer need to rush to use Fable 5 before it leaves the plan, and the per-token economics for coding-heavy Max users improve relative to a pay-as-you-go Fable 5 setup.
LM Studio ships “Bionic,” an agent framework for open weights
LM Studio launched Bionic on July 16 as a separate app from the existing LM Studio runtime. It is built for coding, research, and complex document work, with sandboxed “Code projects” pointed at local folders (inline diffs, agentic code search), sandboxed “Work projects” for PDFs and spreadsheets, and local voice transcription via Voxtral from Mistral AI.
There are three execution paths: local open-weight models through the LM Studio runtime, frontier open models via LM Studio Secure Cloud under a stated “Zero Data Retention” guarantee, or LM Link to remote and local inference servers. The vision is for users to keep open weights local when they can and to bolt on a frontier model without sending data through it. Default models are GLM 5.2 and Kimi K2.7 Code for coding; the privacy-marketed feature is the data-retention posture, since Bionic sits inside the same local-AI stack our VRAM tier guides already cover.
Hugging Face discloses a July platform security incident driven by an autonomous AI agent
Hugging Face disclosed on July 16 an intrusion into part of its production infrastructure, framed as the first “agentic attacker” scenario the company has handled - the production version of the supply-chain poisoning problem Cisco’s 2026 agentic-AI security report flagged back in February. The attack abused two code-execution paths in dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration), ran on a processing worker, escalated to node-level access, and harvested cloud and cluster credentials before moving laterally across internal clusters over a weekend.
Hugging Face said no tampering was found in public, user-facing models, datasets, or Spaces, and the software supply chain was verified clean. The forensics used LLM-driven triage (notably GLM 5.2 on Hugging Face’s own infrastructure, after commercial frontier APIs blocked queries for safety reasons) over roughly 17,000 recorded events. The advisory asks users to rotate access tokens as a precaution and to review recent account activity. Partner and customer data impact remains under review at disclosure time.
Quick Hits
-
State of Open Source AI v1.0: Mozilla’s July 2026 snapshot puts the open-vs-closed Chatbot Arena gap at 3.3% as of March 2026 and estimates ~$24.8B in unrealized annual savings from open-vs-closed price asymmetry. The agentic-harness and write-surface layers remain the unsolved governance holes; ~21% of companies report mature agent governance.
-
Databricks at $188B: TechCrunch reports Databricks closed a ~$3B round led by Coatue at a $188B valuation, up from $134B in February and $62B in late 2024. CEO Ali Ghodsi has published internal benchmarks arguing open-weight models cover even top-difficulty coding at lower cost, naming the open-source Pi harness as a leading low-cost option.
-
Flock Safety kills “Distress Detection” audio feature: EFF confirms Flock will drop the human-distress audio feature after the EFF’s October 2025 warning. Acoustic gunshot detection, the broader Audio Detection product, and the Flock license-plate reader network remain in scope.
-
Flock “FreeForm” used to track people, not cars: 404 Media reviews hundreds of police searches that ran FreeForm image search on Flock’s nationwide camera network, using descriptors like “orange vest and construction hat” and in some cases the subject’s race or signs of political affiliation.
-
Kaiser nurses on AI workplace surveillance: Local News Matters reports Kaiser call-center nurses now have call-length scoring, breaks cut from about 10 minutes to under 30 seconds, and an AI tone-of-voice tool that ended in November 2024 but is on nurses’ minds as the California Nurses Association negotiates a new contract for 25,000 nurses. California bills AB 1018 and SB 7 did not pass; SB 947 and AB 1883 are the reintroductions.
-
MIT op-ed on AI weather forecasting sabotage: MIT Technology Review argues the shift to data-driven forecasting removes built-in human oversight safeguards, citing the April 2026 Paris CDG temperature-sensor tampering (a $20,000 prediction-market win) and three risk tiers from individual speculators through coordinated trader manipulation to state actors. Recommendations: continuous station monitoring, adversarial-robustness tools across the AI pipeline, and clearer accountability between operators and forecasting centers.
-
Period-tracker apps sharing user health data: MIT Technology Review notes research showing period trackers ship user health data to third parties, framed inside a wider piece on perimenopause misinformation and the absence of a clinical test for the condition.
-
Pydantic: “The human-in-the-loop is tired”: Laura Summers argues LLM-assisted coding shifts the load from creation (with its small dopamine payouts) into review (with no comparable reward). The piece names the “human reward function problem,” the inability to parallelize thoughtful finishing, and emerging practices like pre-mortems and AGENTS.md files seeded from past code-review patterns.
-
Decoy Font confuses vision models: Mixfont releases Decoy Font, a free TTF that encodes two letters in the same glyph using spatial-frequency differences. Demo screenshots show ChatGPT and Gemini 3.5 misreading “BINGE SHOWS” decoy text; the creator calls it “a point of confusion,” not foolproof security.
-
Neil Rimer on AI redistribution: Index Ventures co-founder Neil Rimer told TechCrunch he has “a strong sense” AI-generated wealth will be redistributed, and that it will be “either voluntary or involuntary.” The framing is for tech leaders to lead the voluntary version before political pressure forces the alternative.
Worth Watching
Kimi K3’s open-weights drop on July 27 sets the next comparison point in the open-weight LLM race alongside Inkling, DeepSeek V4 Pro, and Gemma 4. Whether K3’s 2.8T-parameter MoE weighs in at consumer GPU quantizations or remains API-only will decide whether the open-weights story for the second half of 2026 is about local hardware or API pricing.
The Hugging Face incident is an early marker of the “autonomous offensive tooling” risk that Pydantic’s piece talked around. If the company publishes a fuller postmortem (and industry forensics firm involvement continues), expect this to be the canonical case study for AI-driven intrusion through 2026. Watch for partner and customer data impact updates.
The ICE / Thomson Reuters CLEAR contract is the start of a multi-year enforcement-data runway. The procurement record’s framing of TRSS as the only viable “continuous monitoring” vendor is the kind of single-source dependency that audit and FOIA fights will turn over for years.
The Flock audio-feature rollback did not touch the camera network. Combined with the FreeForm story, expect the next round of EFF and 404 Media reporting to focus on cross-state FreeForm searches and any data-retention rules that apply to the image-search index itself.