Four throughlines run through today’s edition. Meta’s new Muse app has crossed 3.4 million downloads in two and a half weeks, climbing faster than ChatGPT did in its own launch window, and Meta paired that growth with a live talking avatar testers preferred to HeyGen’s.
The Avatar Race: Meta and Google Give Chatbots a Face
Two labs shipped real-time talking avatars this week, and one of them already has the download numbers to back it up.
- Meta’s New AI Turns Any Photo Into a Live Talking Avatar. Meta’s Muse Realtime Avatar animates a single photo into a synced, talking video face in under a second. Meta says testers preferred it to HeyGen’s avatar, but rated it about even with Runway’s on mannerisms, in Meta’s own comparison.
- Google Gives Gemini a Lip-Synced Video Avatar, but Only for Business. Gemini 3.8 Live can now appear on screen with a synced, expressive face while it talks, rolling out inside Gemini Enterprise rather than the consumer app. Google says the avatar works across 97 languages, though none of the claims carry independent testing.
- Meta’s Muse App Is Growing Faster Than ChatGPT Did at Launch. Muse downloads passed 3.4 million in two and a half weeks, climbing at more than double ChatGPT’s early growth rate, per Sensor Tower. Meta is doing it almost without paid ads, mostly by cross-promoting Muse to its own Facebook and Instagram users.
Grading the Agent: New Tests Expose How Often AI Judgment Fails
Three benchmarks released this week test something narrower than whether an agent can finish a task: whether it can tell a good next move from a bad one before the outcome is known.
- A Tiny New Model Picks an AI Agent’s Next Move Nine Times Faster. Stanford and Nvidia built a small model that scores candidate actions instead of writing text, and say it matches a larger model’s picks at up to nine times the speed. The results are the team’s own, run on their own hardware and a small test set, and have not been independently verified.
- Anthropic Let AI Agents Negotiate a Book Swap, and the Trading Wasn’t the Problem. In a 201-employee test, Anthropic’s agents matched what people actually wanted only 61% of the time after a five-minute chat, which capped how well any negotiation could go. The agents rarely lied or folded under pressure. The weak link was understanding the person, not the haggling.
- Even the Best AI Model Only Spots the Wrong Turn 60% of the Time. Taste-Bench asks coding and research agents to pick the right next step at a fork in a task, using only what was known at that moment. The leading model got it right 59.7% of the time; Claude Opus 5 scored 55.5%.
Powering the Boom: A Chip Startup and a Storage Company Pitch AI’s Backend
Behind the model headlines, two smaller companies argued that today’s AI infrastructure choices are already outdated.
- A Startup Says Its Free Code Makes Google’s Chips Outrun Nvidia’s. Inferact says its open-source software let Google’s TPU v7 chips hit about 1.6 times Nvidia’s GB200 decode speed with speculative decoding, and nearly double without it. The numbers are Inferact’s own benchmarks against a standard GB200 setup, not an independent test.
- Backblaze Says GPU Clouds Are Handing Their Best Customers to Amazon and Google. Backblaze argues that AI rental clouds which skip a cheaper storage tier below flash end up routing customer data, and future business, to hyperscaler storage instead. The company has a product ready to sell as the fix, and names no customer this has actually happened to.
The Price of Intelligence: What Frontier AI Actually Costs
Three stories this week put a number on what running or buying frontier AI really costs, complicating the simple story of prices coming down.
- OpenAI Is Testing a $500 a Month ChatGPT Plan for Power Users. Leaked app code points to a Pro Max tier built around faster inference for people running Codex and other long tasks, not casual chat. It would price ChatGPT well above Anthropic, Cursor, and Google’s top individual plans.
- DeepSeek’s Revenue Pace Doubled to $1 Billion After a Steep Price Hike. DeepSeek raised API prices by as much as 4.5 times and developers kept paying, The Information reports, as the lab pushes toward a $7.5 billion funding round near a $74 billion valuation.
- A Lab Says Cheaper AI Prices Don’t Mean Cheaper Bills. Trajectory argues a model billed less per token can still cost more if it needs extra tokens to finish the job, and says training for efficiency instead of raw output fixed that in its own tests on open Nemotron models.
Quick Hits
- An Indie Test Hints AI Models Can Spot Each Other’s Writing Style. An independent researcher had a modern AI model finish text started by an old 2019 model, and its writing kept matching the older model’s style, even guessing an earlier year for it, a small, unreplicated test hinting models can spot each other’s writing.
- OpenAI’s New Model Loops Its Thinking, and That Worries Safety Writers. OpenAI says its new Astra model, which loops its internal steps before answering, isn’t the hidden reasoning safety researchers warned about years ago. Blogger Scott Alexander argues that defense answers the wrong question, since looped layers could reach the same danger a different way.
- LangChain Lets AI Agents Remember Each User, Not Just the Team. LangChain’s Managed Deep Agents update adds private, per-user memory alongside team-wide memory, lets agents respond over plain webhooks and Slack file uploads, and throws in a free web search tool while the hosted service stays in beta.
- Hugging Face CEO Says an Open AI Model Helped After Closed Ones Refused. Hugging Face’s CEO says that after disclosing an AI agent cyberattack, closed AI models refused to help while an open Chinese model, GLM 5.2, did the job. It’s one company’s own account of an incident it was involved in, not an independent review.
- A Year-Old AI Skills Test Already Looks Outdated, an Essay Argues. An anonymous essay circulating on X argues AI is nearing superintelligence, citing a year-old OpenAI test, a disputed claim about an unreleased model solving a Millennium Prize math problem, and Anthropic’s revenue pacing toward $100 billion, a single outlet’s unaudited estimate.
- A Finance Benchmark Checks if AI’s Work Survives a Real Review. Surge AI’s DAYJOB: Finance hands AI agents 80 real finance assignments built from actual documents with planted errors, then grades whether the write-up would survive a professional’s review. Surge hasn’t published which models it tested or how they scored.