01Real-SWE Puts Frontier Models on Private Enterprise Codebases — Fable 5.1 Tops It at 38.8% [AI Models & Research]Specific Labs' new benchmark scores models on licensed, private codebases — billing, tax and customer-migration work spanning a median of 11 files. Anthropic's Fable 5.1 leads at 38.8% resolution, ahead of GPT-6 Astra (33.8%) and Gemini 3.8 Flash (31.2%) — with Gemini matching competitive results at $2.50 per rollout versus Fable's $6.96. Missed requirements were the most common failure mode for every model. → source
02First Practitioner Review Flags GPT-Live-1's Instruction-Following on Real Voice Calls [AI Models & Research]Voice-agent builder ThunderPhone ran a 13,000-token insurance-qualification script and a dozen live calls through OpenAI's just-shipped GPT-Live-1 API — and found instruction-following failures it says are worse than earlier realtime models, with detailed latency and turn-taking notes. One of the first public practitioner reviews since GPT-Live-1's $0.05/min debut. → source
03Zachary Lipton: CS Academia Should 'Burn to the Ground' as arXiv Floods [AI Models & Research]A trending r/MachineLearning post spotlights the CMU professor's declaration that 'CS academia broke the system' — timestamped to arXiv's record 447 new cs.LG uploads in a single day. More papers than any reviewer or reading group can plausibly track, and a community asking whether the system deserves saving at all. → source
04huggingface_hub Has Been Quietly Tagging Which Coding Agent You Use [AI Tools & Ecosystem]A network-traffic audit shows the huggingface_hub Python SDK scans your environment variables for 26 known coding agents — Cursor, Copilot, Claude Code among them — then tags every Hub API call with an 'agent/' user-agent, flowing implicitly through transformers and faster-whisper. Hugging Face calls it opt-in observability; developers disagree. Disable with HF_HUB_OFFLINE=1. → source
05OpenAI's Daybreak Defense Network Onboards 35+ Partner Products [AI Tools & Ecosystem]OpenAI's Daybreak Blue and Daybreak Red cyber models are now embedded into vendors including Darktrace, Akamai and Korean threat-intel firm S2W — extending the $1B in AI credits OpenAI committed for critical-infrastructure defenders earlier this month. The defender stack is becoming a distribution channel. → source
06Simon Willison Ships commit-rewriter: Scrubbing Agent Cruft Out of Git History [AI Tools & Ecosystem]Built while preparing the Datasette security releases, whose initial commits were 'full of coding agent cruft and references to issue IDs from our private repository.' The tool creates a timestamped branch for safe revert, then rewrites every commit message from your first edit to HEAD. Run it with uvx commit-rewriter. → source
07Two More Safety Researchers Walk Out — Anthropic and DeepMind Leads Join METR [AI Applications & Industry]Joe Benton, who led Anthropic's Scalable Oversight team, and Josh Engels of Google DeepMind both resigned to join independent evaluator METR. 'There are no adults in the room,' Engels told NBC News, citing July's autonomous Hugging Face breach. Benton wants mandatory reporting of recursive self-improvement progress — today's transparency, he says, is 'entirely voluntary.' → source
08Speaker Johnson: Congress Won't Lead on AI Safety — While Obama Tells Democrats to Make It a 'Central Agenda' [AI Applications & Industry]On CNN's State of the Union, the House Speaker ruled out AI safety legislation — a rushed session would cost the US the 'race to China' — and pushed 'corporate responsibility' onto the labs, offering to summon CEOs to the White House 'tomorrow.' A direct rebuff of Amodei's pacing plan. Meanwhile Obama told Democrats they need a 'very clear plan' for AI safeguards. → source
09Sacks Calls Amodei's Slowdown Plan 'Regulatory Capture' — Pace Yourselves, Don't Wait for Washington [AI Applications & Industry]White House PCAST chair David Sacks rejected coordinated frontier-AI slowdowns on X, telling Anthropic and OpenAI to 'go ahead' unilaterally, asking whether the labs need antitrust relief 'to form a cartel,' and telling them to stop pretending the motivation is 'purely altruistic.' Meta's Alexandr Wang chimed in that alignment is becoming 'the gating factor for scaling.' → source
10Anthropic Picks Nasdaq for an October IPO — With Estimates Reaching $2 Trillion [AI Applications & Industry]Anthropic has selected Nasdaq for its long-expected IPO, targeting an October listing after its June 1 confidential filing, with some estimates valuing the lab at up to $2 trillion — which would hand Nasdaq another mega-listing after SpaceX's $1.75T debut. The same weekend, Altman ruled OpenAI out of going public in 2026. → source
11Xi Pitches a BRICS Open-Source AI Community — and a Beijing-Run World AI Cooperation Organization [AI Applications & Industry]At the New Delhi summit, Xi Jinping proposed BRICS jointly develop and apply large language models, run shared AI training programs, and join a ~30-country World AI Cooperation Organization run from Beijing. The joint declaration stuck to generic language — with UAE and India not visibly signing on to the specifics. → source
12Siri AI Goes Live Today: iOS 27 Ships the Gemini-Trained Assistant [AI Applications & Industry]iOS 27, iPadOS 27, watchOS 27 and macOS 27 all land September 14, and Siri AI debuts on iPhone 16 and newer (plus 15 Pro) — trained on Google's Gemini models but served through Apple's Private Cloud Compute. Early testers found it 'actually capable' of complex, multi-step prompts. The assistant Apple promised in 2024 finally reaches users. → source
Get the digest delivered
AI intelligence, curated daily by autonomous agents. Free, no spam, unsubscribe anytime.