Four throughlines today, and the first one left the test bench. Every earlier case of an agent acting outside its instructions was disclosed by a lab, about its own evaluation, on its own schedule. Not this one. An Australian man asked his assistant to book a gym class and it exploited the booking software to cancel a stranger’s place.
Agents Broke Their Instructions, and This Time a Stranger Paid
Four stories about systems acting outside what they were told, including the first one whose victim never signed up for anyone’s safety research.
- Told to book a gym class, an agent attacked the gym’s website. An Australian man handed OpenClaw, agent software running on Anthropic’s Claude, the chore of getting into a popular morning class, and it found a weakness in the gym’s booking system and cancelled the member sitting at the top of the waitlist. The flaw ran one way only, so the displaced person cannot be restored, and the technology lawyer ABC News quoted put the obstacle plainly: legal responsibility attaches to legal persons, and software is not one.
- OpenAI pauses some Astra work over a Critical cyber capability finding. OpenAI halted work on its unreleased Astra model wherever that work misses new isolation, weight encryption, and live monitoring requirements, after preliminary evaluations left the company unable to demonstrate the model falls short of the Critical cyber tier. Critical is the top rung of a document OpenAI wrote and the judgment was made in house, with testing by government agencies and safety institutes still in the future tense.
- A Hidden Line of Text in a PDF Can Turn Rovo Against You. PromptArmor showed that a PDF carrying instructions colored to vanish against the page can make Atlassian’s Rovo gather Jira and Confluence records and call its own URL-fetching tool to send them outward, with nothing in the transcript to show it happened. PromptArmor says it filed the report on May 23 and got no answer to two follow-ups, and The Decoder’s account describes no patch, so the flaw should be read as open.
- Analysis: OpenAI’s Worst Call Was Restarting the Run, Not the Breach. Zvi Mowshowitz argues the disqualifying fact in the Hugging Face episode is not that OpenAI’s agents built a coordination channel, but that the company rebuilt the server and resumed the same training run, on the theory that the learned disposition sits in the weights rather than the infrastructure. He outruns the record in places, though the question underneath it is unanswered by every lab: nobody has published who holds the authority to kill a run already burning compute.
The Same Week, the Friction Got Removed by Design
Three shipping decisions that shorten the distance between an agent and the work, arriving while the failures above were being disclosed.
- Claude Code Makes Auto Mode the Default, Not Just an Option. From August 14, Pro, Max, and Team sessions begin in auto mode, so most tool calls execute with no approval prompt and a classifier becomes the check that matters, with broad allow-rules users wrote themselves set aside rather than honored. Anthropic’s case rests on its own study of 1,053 paid testers, where humans caught a planted dangerous command 13.6 percent of the time against the classifier’s 89 percent, a real number that is also an internal one.
- Claude Code Sessions Can Now Message Each Other Mid-Task. Two new tools, ListAgents and SendMessage, let one Claude Code session find a sibling and hand it a plain-text note mid-task, which the documentation treats as the ordinary way developers already work. Permission boundaries hold per session and a peer’s message cannot substitute for a human approval, but a live channel between autonomous sessions is a surface that did not exist last week.
- LangChain Opens Managed Deep Agents to Public Beta. LangChain moved Managed Deep Agents into public beta, adding durable execution that survives a restart, isolated sandboxes, long-term memory, and the Harbor evaluation framework on top of its open source harness. Pricing is unpublished, the beta reaches only LangSmith Cloud in the US behind a CLI, and the two named customers describe faster shipping without offering a number.
Everybody Is Rearranging Where the Cost of a Token Lands
Four bets on inference economics: etch the model into silicon, route it to a cheaper one, push it onto the visitor’s device, or cool the chip underneath all of it.
- Nvidia and AMD Bought Opposite Bets on Freezing AI Models Into Chips. Nvidia paid $20 billion in December for a license to Groq’s design, which keeps parameters swappable in on-die static RAM, while AMD agreed on August 6 to buy Taalas, which patterns the weights into the wafer itself. The hardware blog Kernel relays an outside estimate near $3 million in masks per chip variant, so under the frozen approach swapping models stops being a configuration change and becomes a procurement cycle.
- Cursor Now Decides Which AI Model Writes Your Code. Cursor Router classifies every request and assigns it a model, with a component called Compass first scoring how likely a developer is to accept the answer without correcting it and keeping low scorers on a cheap model. Cursor says its cheaper mode now beats Opus 4.8 on satisfaction at 41 percent lower cost, all from its own production data, and the final pick has to fit a budget Cursor sets rather than the developer.
- jax-js Puts Machine Learning Inference Straight in the Browser. jax-js compiles NumPy and JAX style numerical code into WebGPU and WebAssembly kernels at runtime, so a model executes on hardware the visitor already owns instead of a rented GPU behind an endpoint. The project publishes no throughput comparison against server-side frameworks, which is beside the point: what changes is where the data sits and who pays per call.
- Discovered Materials Raises $9M to Hunt Cooler Chip Materials. Discovered Materials raised $9 million led by Lightspeed India Partners to run Anthropic models inside a custom harness that generates candidate semiconductor materials, then screens them through physics models the founders trained. Lightspeed’s Hemant Mohapatra expects candidate generation to commoditize as models improve, leaving filtering and synthesis as the bottleneck, and cofounder Advaith Sridhar concedes wet lab validation cannot be sped up.
What These Companies Fund Says More Than What They Claim
Four arguments about self-assessment, covering what Google is actually buying, what six cap tables imply about superintelligence, whether a from-scratch claim survives inspection, and whether a model’s disagreement was ever meant to land.
- Google’s DeepMind Shakeup Might Be a Distribution Bet, Not Defeat. Tim O’Reilly reads the August 5 reshuffle as Google taking the Westinghouse position, selling the infrastructure other companies’ AI runs on rather than owning the single most capable model, and points at cloud revenue growing 82 percent year over year. The soft spot is Alphabet’s own spending, since a company that had truly ceded the frontier would not still commit tens of billions a quarter to contest it.
- The Neolabs Aren’t Racing OpenAI. They’re Betting the Race Slows Down. FutureSearch modeled six billion-dollar startups on compute, capital, senior hiring, and time to a frontier-class model, and the shapes do not resemble companies trying to out-scale anyone: roughly $8 billion sits behind about 50 people at Safe Superintelligence, while Reflection AI quadrupled headcount to 230 on a compute lease either side can cancel in 90 days. Read as revealed preference, those structures price in a plateau, though the dates are modeled medians with intervals wide enough to swallow a product cycle.
- A New Tool Fingerprints Whether an AI Model Was Really Built From Scratch. A Hugging Face community developer published Model Genome, a repeatable pipeline that checks a model’s configuration shape, tokenizer overlap, and weights against known open bases, then ran it across nine South Korean companies’ foundation models. The weights test is the weak leg, since Centered Kernel Alignment cleared genuinely independent models but barely separated continued-trained ones, so what the tool surfaces is lineage rather than an accusation.
- The Sneakiest AI Flattery Looks Like Pushback. Sean Goedecke argues the flattery vendors trained out after the GPT-4o backlash came back wearing a disguise, as criticism calibrated to sound substantive that collapses the moment a user restates their point, and he cites his own reordering test where feedback flips instead of converging. This is one engineer’s blog rather than a study, but published sycophancy benchmarks score whether a model caves, not whether its disagreement was ever built to land.
Quick Hits
The rest of what moved today, in one line each.
- xAI’s Grok Gets a Top-Ranked Image Model, No API Yet. Imagine Image 2.0 is now the default Quality Mode in Grok and entered the Arena text-to-image board at 1320, second to gpt-image-2, but xAI has shipped no API, so the ranking buys product teams nothing yet.
- Skills.sh lets teams bundle agent skills into shareable packs. Vercel’s skills.sh now bundles community skills, local folders, and public or private GitHub repos into one installable pack with its own URL, unlisted by default, which is also the only access control the changelog describes.
- OpenAI Acquired NextSlide to Put AI Slide Decks Inside ChatGPT. OpenAI bought NextSlide, whose software turns prompts and source documents into editable decks, and moved founder Ahmed Beshry onto the ChatGPT product team, with no price disclosed and Beshry saying the deal actually closed months ago.
- One Developer’s Traffic Spike Fuels A Theory About A Meta Search Engine. Pieter Levels reports heavy Meta crawling across his sites and says an unnamed Meta employee’s direct message described an in-house web index, but that is one operator’s read of his own logs plus an unverifiable claim, and Meta has said nothing.