Most of this issue is about the distance between what AI is claimed to do and what anyone has shown. Fourteen of the fifteen stories carry an 8 October source date and one repository page carries none, so nothing here happened today.
Show Your Receipts: Big Claims, Thin Evidence
Four companies described what their AI can do. In each case the proof is missing, undated or limited to people who already have a seat.
- OpenAI’s fastest GPT-6.1 Sol mode is for select plans, with no benchmark. Ultrafast works in Codex and ChatGPT Work only on the Pro 500 plan and on qualifying Enterprise and Edu accounts, and an Enterprise admin has to switch it on. OpenAI says “up to 8x faster” with “near-Astra intelligence” and shows no benchmark for either.
- Google calls Gemini a universal worker, but dates almost none of it. Google Cloud’s post gives a shipping status only to the industry editions, with Financial Services and Legal in preview. The governance and cost controls carry no dates, and the capability and adoption claims are Google’s own.
- Devin’s Security Swarm docs say nothing about how often it is wrong. Cognition’s documentation says it tested against published vulnerabilities and reports no results. It makes no precision claim at all, and no false-positive rate appears anywhere on the page.
- Hone raises $60 million for business-running agents nobody has seen. The five-month-old startup says the round values it at $285 million, a figure Hone supplied itself. Bloomberg shows no customer, deployment or demo, and an investor on the board talks about agents needing “immune systems”.
Graded by the Author: Tests With a Catch
Three groups measured agents and published the result. Each one says plainly how the measuring was done, which is what makes the numbers worth reading.
- Epoch AI gave six models its own jobs and says they are not ready. Eleven tasks, six models, one Epoch grader and one run per task. The best models handled well-defined work and fell down on open-ended research, and Epoch publishes no per-model rubric scores.
- A search company says search agents miss a third of the answers. Exa’s ATLAS benchmark has the best system at 0.66 row F1, after 18 minutes and $8.92 a task. Exa sells a search API, designed and ran the test, and its answer key was built and audited by models.
- Harvey’s legal agent went from 2.9 to 15.7 percent and still fails most tasks. A memory of lessons from graded work lifted full-pass rates fivefold without retraining the model. That still leaves most tasks failing at least one check, and the post never says what the rates are averaged over.
Whose Number Is It: OpenAI Under Two Spotlights
Two OpenAI stories where the account depends on who is telling it and what they could see.
- The $70 billion OpenAI revenue figure came from investors, not from OpenAI. Investors put it together from information OpenAI had shared with them, to compare it with Anthropic. OpenAI has since told them its run rate is approaching $50 billion, per the FT through TechCrunch.
- The fired OpenAI researchers publish their own side. Jasmine Wang, Tomek Korbak and Mikita Balesni deny leaking to The Information and explain a dispute over email access. OpenAI’s position reaches us through the Wall Street Journal, and we have no fresh comment from the company.
On the Builder’s Bench: Limits to Know Before You Ship
Four pieces for people running this stuff in production, each with a limit that is easy to miss on first read.
- A popular AI speed trick stops helping once your servers get busy. Engineer Joe Barrow argues speculative decoding does extra work rather than less. It only wins while the chip has idle capacity, and his five-slot example is a teaching device, not a measured crossover.
- Microsoft ships a throwaway computer for agents to wreck. Quicksand boots disposable virtual machines with no admin rights or Docker. Network access is off by default and most useful agents will turn it on, and a snapshot rollback cannot unsend a request.
- NVIDIA’s video lab now trains robot policies for YAM, Franka and Unitree G1. The LongLive repository grew to four projects, and Long-WAM trains policies that remember long stretches of video. Code is Apache 2.0, the README quotes no success rates, and the newest work is the least reviewed.
- Anthropic will send maintainers bug reports no human read. OSS Scanner is free, and its reports are model-generated and sent without human review, with a true-positive rate above 90 percent expected. The verifying lands on volunteers, and Anthropic admits finding bugs was never the hard part.
Today’s Quick Hits
- Midjourney tests a Thinking rerun button. Alpha users can rerun a finished image in Thinking mode. Midjourney says it helps with prompts and lettering, shows no measurements, and calls wider release a possibility.
- Voyager pitches an AI toolkit for artists. The site promises faster, cheaper, better creative work from top models, with no measurement, price, launch date or documentation behind any of it. Until someone tests it, creative teams have nothing to benchmark.