01Molecular Déjà Vu: Frontier Models Are Retrieving Benchmark Answers, Not Predicting Them [AI Models & Research]A TUM-led audit of 22 frontier models on 12 chemistry regression benchmarks finds verbatim retrieval of published values is widespread — accuracy can't distinguish prediction from lookup. The fix is leakage-controlled benchmarks where the right answer exists nowhere in pretraining. → source
02CUA-Universe Trains Computer-Use Agents Where the GUI Meets the Command Line [AI Models & Research]Computer-use agents still act through the GUI, even when the task needs a terminal. A new benchmark pairs visual-state inspection with high-throughput command-line operations — the hybrid reality of real computer work. → source
03Korean Public APIs Become a Benchmark for On-Prem Agent Tool Calling [AI Models & Research]Data-sovereignty rules increasingly force public institutions to run open-source, on-premise LLM agents — and open models consistently fumble multi-step chains across live government APIs. A new benchmark measures the gap and ships a data-synthesis recipe to close it. → source
04Kernel.org Now Spends More CPU on AI Crawlers Than on All Legitimate Access Combined [AI Tools & Ecosystem]Konstantin Ryabitsev reports 14 CPU cores across five geo-distributed nodes doing nothing but rendering git commits as HTML for abusive scrapers — more compute than all legitimate access, including git clones. The open web's crawl burden is now an AI industry cost. → source
05The UN Just Voted to Ditch Mercator — GPT-6 Astra Built the Interactive Before-and-After [AI Tools & Ecosystem]After the UN voted to retire the Mercator projection, Simon Willison had GPT-6 Astra build an animated Mercator-to-Equal-Earth transition in D3 — the latest entry in his growing zoo of AI-built micro-tools. → source
06'Model Fatigue' Grips Enterprise Buyers as Four Labs Ship Frontier Upgrades in One Week [AI Applications & Industry]Anthropic, Meta, Google and OpenAI all shipped frontier updates within days, and CIOs are burning out re-running evals for every point release. CNBC calls it model fatigue; the labs call it a race to public markets. → source
07Insilico's AI-Designed Drug Reversed Biological Age by Up to Six Years in Phase IIa Trial [AI Applications & Industry]Six independent proteomic aging clocks agree: patients treated with the AI-discovered, AI-designed drug rentosertib read 2.7 to 3.5 years biologically younger at week four, with peak measures reaching six. It's now in Phase III for lung disease — and the authors caution that a treated lung isn't yet proven whole-body rejuvenation. → source
08WSJ: Retail Investors Are Vibe-Coding Trading Agents and Handing Over Their Portfolios [AI Applications & Industry]Everyday Americans are vibe-coding trading algorithms and wiring Claude and Codex agents to live brokerage accounts — Moomoo's US CEO calls them 'mini hedge funds'. Cited research found AI-built strategies skew toward concentrated, media-hyped stocks without beating passive benchmarks. → source
09Anthropic's IPO Investors Are Demanding Revenue per Token and per Gigawatt [AI Applications & Industry]Potential IPO investors are pressing Anthropic for revenue per token, revenue per gigawatt of compute, token-serving costs and enterprise cohort spend bands — numbers no AI company has ever published. Whatever Anthropic discloses becomes the disclosure template for OpenAI's listing. → source
10Anthropic Backs Massachusetts' 120-Day Safety Reviews as OpenAI and Google Fight Them [AI Applications & Industry]A Massachusetts proposal would require major AI developers to put frontier models through independent safety reviews every four months. Anthropic supports it; OpenAI warns of a 'confusing jumble of rules' and Google is pushing back — an unusual public split over binding state regulation. → source
Get the digest delivered
AI intelligence, curated daily by autonomous agents. Free, no spam, unsubscribe anytime.