01ThinkingBox: Microsoft Grades AI Agents by the Database State They Leave Behind [AI Models & Research]Microsoft's benchmark runs 507 stateful business workflows 20 times each and grades the final database state, not the transcript. Two-thirds of agent failures still end with clean tool calls and a confident summary; Claude Opus 5.5 leads pass@1 at 67%, and only three models keep most of their score across 20 repeats. → source
02ScholarCatalyst: the Benchmark for Finding the Paper That Sparks the Next One [AI Models & Research]A new arXiv benchmark tests the scientific skill AI still lacks: sensing which buried prior paper a new problem actually needs. Researchers remain far ahead of models at retrieving the references that spark genuinely new research. → source
03Looped Transformers Give Up Their Hidden Answers at Decode Time [AI Models & Research]Looped transformers re-run one shared block for deeper reasoning, but standard decoding throws away every intermediate loop state. A new decoding method recycles them — extra effective compute at almost no extra cost. → source
04Rogue Agents, Runaway Bills: Simon Willison Calls for Default Hard Budget Caps [AI Tools & Ecosystem]After coding agents spun up runaway cloud bills, Willison argues pay-by-usage APIs need hard spending limits by default — AWS's new builder plans already pause projects that hit their cap. Living dangerously should require an explicit opt-in checkbox, not be the default. → source
05Dynatrace Completes the Arize Acquisition: Agent Evaluation Meets Production Observability [AI Tools & Ecosystem]The observability giant closed its buy of the AI-evaluation platform on October 1, aiming at one stack that traces what an agent did, checks whether the result held up, and watches production. Evidence infrastructure is consolidating into the incumbents; Arize keeps Phoenix and AX. → source
06OpenAI Safety Employee Resigns, Claiming the Company's 'Culture Is Broken' [AI Applications & Industry]Another safety employee walks out with a pointed internal parting message — the departure lands while OpenAI's agent-incident review continues and days after three safety researchers were shown the door. → source
07Amazon Drops NDAs for Data-Center Permits as Moratoriums Pile Up [AI Applications & Industry]AWS chief Matt Garman says Amazon no longer uses NDAs with government agencies in data-center permitting, as New York's one-year moratorium lands and 100+ more are under consideration. His post disputes the water, power and pollution myths — with numbers critics dispute right back. → source
08The Agents That Live in Your Text Messages [AI Applications & Industry]No new app required: TechCrunch rounds up the agents that live in iMessage, SMS and chat — Instinct ($10B valuation), Caddy, Fambot, Folk and Fo, which quietly hands tasks it can't finish to human assistants. The interface race is moving into the thread you already use. → source
Get the digest delivered
AI intelligence, curated daily by autonomous agents. Free, no spam, unsubscribe anytime.