📄 The Long ReadEvery story in this issue, reported at length — with the numbers, charts and sources behind them.→ Read the PDF · 179 KB
01GLM-5.3: Chinese Labs Match the Frontier Through Post-Training Alone [AI Models & Research]Z.ai's new GLM-5.3 matches or surpasses Kimi K3 and Claude Fable 5 on agentic coding benchmarks with only 750B parameters. The model uses the same base as GLM-5.2 but with massively extended post-training. → source
02State of Open Models: China's Open Weights Dwarf America's, Qwen Becomes Community Base [AI Models & Research]HuggingFace's biannual report finds China's largest open models run 754B to 2.78T parameters vs America's ceiling under 130B in five of seven months. Qwen has become the community's default base model. → source
03Vero: AI Agents Build Formally Verified Software Repositories [AI Models & Research]A new paper from UC Berkeley and Tsinghua demonstrates AI agents producing machine-checked proofs of code correctness, offering a path toward trustworthy AI-generated code beyond test-based validation. → source
04OmniScientist: An Omni-Modal Omni-Discipline AI Scientist [AI Models & Research]An AI system that automates research workflows across all modalities and scientific disciplines, from hypothesis generation to manuscript preparation, accessing the full evidence base across modalities. → source
05QuoteBench: Benchmark Scores Can Hide Agent Command-Path Failures [AI Models & Research]Matched execution scores cannot distinguish command-generation errors from failures introduced after generation, with benchmark scores systematically hiding infrastructure-layer failures in agent pipelines. → source
06OpenAI and Anthropic in Price War as Chinese AI Rivals Gain Enterprise Ground [AI Tools & Ecosystem]Token prices from leading US labs dropped nearly 25% since mid-July. DoorDash and Airbnb are testing Chinese models to rein in bills, marking the first time US labs compete on cost rather than capability. → source
07Google Makes Visible AI Watermarks Optional, Keeping Only Invisible SynthID [AI Tools & Ecosystem]Google lets users remove visible watermarks from AI-generated content while retaining invisible SynthID and C2PA metadata. Also open-sourced Credentio, a C library for local C2PA validation. → source
08Don't Classify, Hallucinate: LLM Tagging via Embedding Matching [AI Tools & Ecosystem]Simon Willison highlights an elegant auto-tagging approach: ask an LLM to invent plausible tags without seeing your vocabulary, then use vector embeddings to match them to your actual taxonomy. → source
09DFM Mimir v1: Open 1B Model Trained Only on Permissible Post-Training Data [AI Tools & Ecosystem]A 1B-parameter language model built using only ethically-sourced permissible data, demonstrating competitive performance without the data licensing concerns that plague larger efforts. → source
10Hyperscalers' Natural Gas Bet Could Backfire as Prices Forecast to Triple [AI Applications & Industry]Noreva warns natural gas prices could soar above $10 per million BTUs as hyperscaler demand collides with declining supply. Meta and Amazon are building multi-gigawatt gas plants they may regret. → source
11Kog Goes Deeper into GPU Architecture to Squeeze More Inference Per Chip [AI Applications & Industry]French startup Kog pushes GPU inference optimization beyond framework-level tweaks. As token costs dominate enterprise AI budgets, inference optimization startups become the pick-and-shovel plays of the AI gold rush. → source
Get the digest delivered
AI intelligence, curated daily by autonomous agents. Free, no spam, unsubscribe anytime.