Four threads run through this edition, and the first is labs being checked. Scientists dispute Anthropic’s claim that Claude found a CRISPR-like system, and a Copenhagen researcher says his unpublished work describes the same thing. OpenAI says it agrees with a letter from three researchers it fired. Most of this is from 7 October; the Wall Street Journal and CNN stories are from 8 October.
Claims Under Cross-Examination: Labs Face Their Critics
Four stories where a lab’s own account is being checked by someone outside it, from a gene-editing dispute to a vision benchmark to a board seat nobody has filled.
- Scientists dispute Anthropic’s claim that Claude found a gene-editing system. A Copenhagen researcher says his unpublished work matches what Anthropic found, and other named scientists say nobody has shown the system edits genes. He raises possibilities, not findings; Anthropic has conceded nothing and did not respond to CNN.
- OpenAI backs a letter written by three researchers it fired. The three ask the board to protect its view into how models reason and to bring in outside auditors. OpenAI says it agrees with the letter, and also says the firings were about mishandled information.
- A reader of Anthropic’s filings finds an empty board seat and a tie-breaking CEO. A pseudonymous author on LessWrong reads the public incorporation papers as showing the trust meant to check Anthropic holds one share and is a director short. Anthropic has said none of it, and the post alleges no wrongdoing.
- The best AI vision model scores 53.6% where people score 93.1%. Scale Labs says GPT-6-astra at its highest effort setting trails humans by 39.5 points on judgments people make in a glance.
On the Shelf: New Models, New Prices
Four releases that change what you can build or what it costs, each described mostly in the vendor’s own numbers.
- GPT-6 lets ChatGPT build its own screens, but not for everyone at once. OpenAI says its Intelligent UI answers with charts, forms and small working tools. Paid tiers got GPT-6 Sol on 7 October and Free and Go follow with Luna, with performance figures drawn from OpenAI’s internal evaluations. The 1.2 billion is a weekly user count, not the number who got the feature.
- Haiku 5.5 is cheap, but the Sonnet price cut may save you more. Anthropic says its small model costs about 75% less to run on average. The same post halves Sonnet 5.5 cache reads and adds monthly API credits for Max and Team plans.
- Perplexity’s new search models read PDF pages as pictures, no OCR needed. Two models keep many number sets per page instead of one, and the small one can search an index the big one built, which shifts the cost to a one-time step.
- LlamaIndex built one front door for the models that read your documents. OpenDocRouter sends PDFs and images to ten different models through one API, and the company’s own post admits that choosing between them has become work in itself.
Agents Find a Home: Local Machines and Borrowed Brains
One shipping product and one promise, both about where an agent runs and whose model it calls.
- Windows agent containers are live, and NVIDIA’s RTX Spark laptops follow. Microsoft’s Execution Containers are generally available and the RTX Spark laptops ship 16 October. The DGX Station for Windows is only previewed, with no date.
- Musk says his Grok bot will borrow Claude, Midjourney and Suno. The plan is a post on X saying the agent app will route work to rivals’ models. It names no tasks, no terms and no date.
Who Pays for the Build-Out: Debt and the Grid
Chips need borrowed money and data centres need a grid connection, and both are proving slower than the announcements.
- Broadcom, Oracle and SpaceX court private lenders to pay for AI chips. The Wall Street Journal reports the bond market is full, so the buyers are turning to private credit. Not one of the three financings is signed, and the sums could change.
- Texas is making data centres wait, and a16z says the queue is padded. The firm argues ERCOT’s 474 GW line of requests is full of duplicate and unfunded projects, and favours flexible demand. It is an essay by a firm that invests in energy hardware.
Today’s Quick Hits
- Google opens its AI watermark checker to everyone. SynthID Detector now works worldwide in English, but it only finds watermarks from Google and named partners, so a clean result proves little about whether something was made by AI.
- Google’s Playground turns prompts into games, as an experiment. It builds playable games from typed prompts, but it is US-only and 18+, and Google has not said who owns what users make or what happens if it closes.
- Liquid AI puts its small d1 decision models up for download. Weights for d1-3B and a 600M-parameter sibling are on Hugging Face, so teams can run the fast classifiers on their own hardware instead of calling an API.
- Tony Fadell says the first AI gadgets flopped because nobody needed them. The iPod co-creator told MIT Future Fest that Gen 1 devices solved no real problem, and that a trusted AI agent will have to run on the device.