Four throughlines run through today’s 19 stories: a reversal on agentic autonomy, a fight over who controls compute, a market-size claim bigger than reality, and a turn toward local intelligence.
The Agents That Didn’t Deliver: Autonomy Under Review
A reversed restructuring, a shared memory default, and a robotics field still short of reliability all point to the same question: what agents can actually be trusted to do on their own.
- Meta halted a plan to cut 60 percent of some teams with AI agents. Internal documents describe Project OT, a restructuring reversed after the agents assigned to do the work underdelivered and staff revolted over tracking software meant to monitor the transition.
- Anthropic merges Claude’s memory across chat and Cowork, on by default. The consolidation, live for consumer plans, turns memory into a governance question for any team running Cowork as an autonomous agent rather than a chat assistant.
- OpenAI’s Sottiaux says ChatGPT Work bet on “magic,” not choice. The product lead overseeing Codex and ChatGPT Work told TechCrunch the platform passed 20 million users by minimizing user decisions, not maximizing them.
- Robot brain builders say they are moving past their GPT-2 stage. Founders at last week’s Actuate conference told TechCrunch that physical AI is advancing beyond its earliest phase, though data and reliability gaps persist.
Who Owns the Compute: Concentration at the Top
A departing executive, a defensive vendor strategy, and a forecast of near-total supply control all describe the same narrowing pipeline of usable compute.
- Dylan Patel: OpenAI, Anthropic to control most usable compute by 2028. SemiAnalysis founder Dylan Patel says the two labs already take a third of new compute and could reach 70 to 80 percent of incremental supply within two years.
- OpenAI’s CFO makes the case for a nine-vendor compute portfolio. Sarah Friar argues OpenAI’s spread across chipmakers and clouds, not any single deal, is what lowers the cost of useful intelligence.
- Why cheap intelligence won’t kill the application layer. Aatish Nayak argues abundant models shift value toward firms that turn tokens into outcomes, not toward the labs alone.
- OpenAI’s data center chief exits, and his job gets cut into three. Chris Malone is out after roughly eighteen months running OpenAI’s infrastructure buildout. The company is already leaning toward a strategy he wasn’t hired to run.
- OpenAI’s first custom chip posts 1.5x to 4.1x gains in its own tests. Jalapeño, OpenAI’s first inference accelerator, beat comparison systems on power efficiency and latency across three public models, a step toward owning its own supply.
The Local Turn: Intelligence Off the Cloud
A local GPU agent, new desktop silicon, small open models, and a shrinking tokenizer all point the same direction: away from renting the cloud and toward running intelligence at the edge.
- Perplexity and Nvidia Ship an AI Agent That Never Leaves Your GPU. Portable Computer runs Perplexity’s agent stack on local Nvidia hardware, replacing token bills with a one-time GPU purchase.
- IBM Ships Granite 4.2, Betting Small Models on Agentic Skill. The 3B, 8B and 30B open weight models add a reinforcement learning pipeline built to act inside tools, not just answer questions.
- Claude’s tokenizer shrank while everyone else’s grew. Here’s why. Independent testing puts Anthropic’s vocabulary at roughly 16,000 tokens, a fraction of rivals like Qwen, and a possible workaround for a training bottleneck.
The Market-Size Math: IPOs and Geopolitics
One filing rests on a number bigger than an entire sector’s revenue, and one hosting deal is tangled in trade politics and a denied theft accusation.
- The real story in Anthropic’s IPO isn’t the filing. It’s the $30 trillion.. Anthropic’s pitch to investors rests on a total addressable market bigger than the entire S&P 1500 tech sector’s actual revenue.
- Moonshot AI in Early Talks to Put Kimi K3 on US Clouds. Reuters reports the Chinese startup wants up to 30 percent of revenue from Microsoft, Amazon, and Google, amid a US trade threat and a denied theft accusation.
Quick Hits
- JD.com Open Sources an Omnimodal World Model, Echo-WM. The Chinese retailer’s research arm released a generative world model tying video, sound, music, and speech to a single navigable scene, per its GitHub release.
- Vercel Connect exits beta, kills the long-lived API token for agents. The company says agents now request scoped, short-lived credentials at runtime instead of storing secrets that never expire, closing a common breach vector.
- Apple’s M6 and M5 Ultra chips target local AI on the desktop. The new Mac mini and Mac Studio silicon add cores, memory bandwidth, and a bigger Neural Engine aimed at running larger models on-device.
- Applied Compute launches AC2, a cloud for training and serving models. The new platform lets teams train, deploy, and refine their own open-weight models, entering an already crowded custom-model market with no independent benchmarks yet.
- Keenable exits stealth with $26M to build a search index for AI agents. The startup says its index already serves AI labs in production, but it will not name a single customer, leaving the claim unverified for now.