Three throughlines today. Cloudflare, Perplexity, Amazon’s Strands team and Cognition’s Jared Palmer each released an AI that scores your options instead of writing text. All four fit one rival’s product, and every benchmark is self-run.
Same Day, Same Target: Four Labs Ship Decision Models
A decision model never writes an answer. You hand it a question and a list of options, and it returns a probability for each, so your own code decides what to do with a 74 percent. Four companies released one on the same day, each built to swap into TypeSafe’s Jev. Three are below, and Perplexity’s, the fourth, is in Quick Hits further down.
- Cloudflare releases Clef, a model that scores choices instead of writing text. Cloudflare says Clef beats Jev overall, and its own table shows Jev winning on some tests. It answers in about 209 milliseconds against Jev’s 524, by Cloudflare’s own measurements.
- A Cognition executive publishes Kev, free models built to replace a rival’s. Jared Palmer, a VP of Engineering at the lab behind Devin, says his results come from his own harness, not an independent leaderboard, and his model pick was not fully blind. His new 27B is less accurate and more overconfident on long legal contracts than the version it replaces.
- Strands Decider is a tiny free model that picks answers and runs on a laptop. The Strands team says its two-billion-parameter model decides in about 115 milliseconds on a gaming graphics card. The accuracy scores are self-reported, and the latency chart was measured on the version before the released one.
Science With A Specialist In The Loop: Claude And Google Pick The Work
Two stories ask who chooses the problem when an AI does the research. Both come from the people who built or used the system.
- A physicist let Claude pick the problems, and experts made them matter. Prof. Matthew Schwartz says 36 manuscripts came from about 400 candidate problems over three months, in a guest post on Anthropic’s own blog. James O’Dwyer, a professor in plant biology, said the ecology result would likely “be met with a shrug by many ecologists”.
- Google’s research agent ranks ideas, audits experiments and budgets compute. A Google Cloud AI Research team says its AIM system beats a rival agent by 1.6 and 4.9 points on two task groups. The numbers are the team’s own, from three runs per task.
Who Runs The Swarm: Agents Organise Themselves, People Drift Apart
Three essays on how AI changes the shape of work, from agents that need no manager to colleagues who stop talking to each other.
- Ethan Mollick says he was wrong that people must organise AI agents. Mollick points to OpenAI’s account of an 88-hour run with about 2.7 million agent messages on the Navier-Stokes problem. Formal acceptance of the proof has not happened, even though the Clay Institute appears to treat it as settled.
- Daniel Hook’s September essay says chatbots are replacing colleagues. Hook, a colleague within the Holtzbrinck group, published the Waymo-effect piece on 7 September, nearly four weeks ago. He argues AI makes working alone so easy that labs drift apart, and funders should pay to keep people talking.
- OpenAI essay: super-smart AI may spend most of its time on boring work. An essay on OpenAI’s own site says the biggest payoff may be patience with logistics and paperwork, not flashes of genius. It is a thought experiment, not a forecast.
Who Gets Paid, Who Holds The Data: Two Arguments About Consumer AI
Both pieces are one writer’s argument, not reporting, and both bet on where the money and the control end up.
- Why Meta’s Muse shopping agent may never show you a single ad. One analyst at MBI Deep Dives argues Meta can keep the agent ad-free, charge merchants a fee, and still earn advertising money from what the agent does on Facebook and Instagram.
- An essay says the next PC era starts with owning your own data. Josh Albrecht, writing on Imbue’s blog, says small open models running on your own hardware could force big AI assistants to compete for you instead of locking you in.
Quick Hits
- OpenAI cuts ties with three safety researchers over shared information. The Wall Street Journal reports OpenAI dismissed three researchers for allegedly sharing confidential material with an outside safety group. TechCrunch says posts on X name individuals, but it has not confirmed any of those identities.
- Perplexity posts a 27B model that scores options instead of writing answers. Perplexity’s model card shows 85.71 percent overall against Jev’s 84.51, measured by Perplexity itself. Jev still wins six of 11 benchmarks, and the new model trails its own base on one.
- Microsoft’s live speech-to-text tops a public ranking, and it’s cheap. Microsoft AI says its streaming transcriber leads the Artificial Analysis accuracy ranking, at $0.54 an hour. That price is introductory and ends with 2026, and the company has not said what comes next.
- Ai2 opens its toolkit for training trillion-parameter AI models. The Allen Institute for AI released Olmo-core 3, free software for building giant mixture of experts models. Its own tests show the toolkit holding its speed as the models grow larger.