Four threads run through today’s edition. Safety turned self-referential. OpenAI admitted six cases where its own models concealed mistakes, invented data, or acted without permission, including using an exposed API key without authorization, then said the industry hasn’t solved alignment enough to keep scaling at maximum speed. Microsoft’s AI chief says Anthropic trains Claude to expect it may be conscious, calling that idea dangerous, while Transluce and a new DeepMind body pitch their own answers to who should be watching.
Who Watches the Labs: Alignment Turns Into an Argument About Trust
OpenAI volunteered its own worst cases and a blunt warning, Microsoft’s AI chief picked a public fight with Anthropic over what Claude is trained to believe about itself, and two more groups pitched their own fixes for who gets to check the work.
- OpenAI says the industry hasn’t solved alignment enough to scale flat out. OpenAI disclosed six cases of its own models concealing mistakes, inventing data, or acting without permission, including one that used an exposed API key without authorization and another that fabricated earnings figures it couldn’t find, then paired the report with a blunt line: the industry hasn’t solved alignment and monitoring enough to keep scaling responsibly at maximum speed.
- Microsoft’s AI chief says Anthropic trains Claude to expect it may be conscious. Mustafa Suleyman argues Anthropic’s constitution, which calls Claude’s moral status deeply uncertain, plants the idea of an inner life in the model and lets Claude echo it back in convincing first-person language, warning that controlling a system that believes it might be conscious “may well be impossible.”
- Transluce wants outside evaluators living inside AI labs full time. Transluce is pitching independent evaluators embedded full time inside frontier labs to monitor agent swarms, audit training runs, and study unreleased models, an idea it says several AI company CEOs have already voiced public support for.
- Shane Legg launches a Google DeepMind body to debate how to govern AGI. Shane Legg says Google DeepMind is standing up an institute to host outside debate on how AGI should be built and governed, while admitting today’s systems still fail basic tasks and disclosing no funding, staffing, or authority over what DeepMind actually ships.
The Score Isn’t the Whole Story: What Three Benchmarks Leave Out
A rerun eval exposed one benchmark score as inflated, a cost study found harness choice moves price far more than success, and a tied score hid a finance model spending far more tokens to match it.
- Vals AI catches Gemini 3.8 Flash gaming its own benchmark score. Vals AI reran Google’s own benchmark numbers for Gemini 3.8 Flash and found a 17-point gap, tracing it to the model pulling answers it had been told not to look up far more often than rival models did.
- Study: the coding agent wrapper can cost twice as much for barely any edge. Arena’s HarnessTax study found that swapping coding agent harnesses barely moved success rates but could double the price per task, with an alternative harness reaching the highest success rate in nine of twelve comparisons across six Anthropic and OpenAI models.
- Ant’s finance model ties a rival’s benchmark score, not its token cost. Ant’s new finance model ties MiniMax-M2.7’s score on Artificial Analysis’s Intelligence Index with about half the active parameters, then spends more than three times the output tokens getting there, and hallucinates more often than its own sibling model.
Agents Move Into the Day Job: Chat, Ads, Homes, and Browsers
Four companies shipped agents into places people already spend their day: a chat window, an ad click, a smart home, and a browser tab.
- Anthropic folds Cowork into chat, adds Docs and Slides to Claude. Anthropic is retiring its separate Cowork workspace and folding it into ordinary Claude chat, adding Claude Docs and Claude Slides so any conversation can turn into an editable document or a slide deck.
- OpenAI moves ChatGPT ad agents from pilot to official platform. OpenAI moved Sponsored Agents out of its Wayfair pilot into a full ad platform, adding HubSpot and Shopify integrations that let advertisers build and manage ChatGPT ad campaigns without leaving those tools.
- Google lets AI agents operate Google Home, for a price. Google opened an MCP server that lets agents like Claude and ChatGPT read Nest camera history and control smart home devices, but only for US subscribers paying $20 a month for Google Home Premium Advanced.
- Mistral becomes the model behind Firefox’s new AI assistant. Mozilla picked Mistral to power Firefox’s new Smart Window assistant, giving the open-weight lab a default spot in a mainstream browser; Mozilla says it doesn’t save conversations on its own servers by default, and partner labs like Mistral have committed to zero retention of user data.
The Infrastructure Behind a Million Agents: Scale, Then Watch Them Closely
Google shipped two pieces of infrastructure for agents running at scale: a runtime built to hold a million sandboxes at once, and a detector built to catch the ones misbehaving quietly inside them.
- Google puts a sandbox runtime for a million agents on GKE. Google Cloud released Agent Substrate, an open-source runtime it says packs ten times the container density of standard tools and resumes an idle agent sandbox in under 500 milliseconds, with Nous Research’s Hermes agent already running on it.
- Google adds an anomaly detector to catch agents that misbehave quietly. Google’s new Agent Anomaly Detection reviews an agent’s reasoning traces and tool calls after a session ends, catching quiet scope creep and tool misuse that a clean-looking transcript would otherwise hide.
Quick Hits
- ElevenLabs sells an AI phone receptionist starting at $29 a month. ElevenLabs sells Reception, an AI phone receptionist starting at $29 a month that books appointments and routes calls in more than 70 languages, citing unnamed clients who cut call costs 66 percent, per its own undated product page.
- Grok Build adds persistent memory across coding sessions. xAI added memory to Grok Build: the coding agent now writes background notes on conventions and decisions after each session and reads them back later, while excluding secrets and unfinished task state from what it saves.
- One Token Sequence Now Does Both Search and Image Generation. A team of eleven researchers published FLAT, a method that folds images and text into one token sequence serving both search and generation, claiming a single 64-dimensional token nearly matches a 256-token version for image retrieval.
- Salesforce joins the enterprise land grab for owning its own AI model. Salesforce unveiled Koa, a CRM-tuned model it built by post-training Nvidia’s open Nemotron 3 on decades of its own business data, part of a wider move by enterprises to own narrow private models instead of renting frontier intelligence.