Epoch AI, the research group that tracks AI progress, handed six models eleven tasks from its own daily work and concluded they cannot yet replace its staff. The report, by Kelly Hong and Greg Burnham and dated 8 October 2026, is unusual because the organisation measured whether AI could do its own job and published the failure. A vendor benchmark never reads like that.
The suite covers five categories: graphic design, Data Insights (short articles on AI trends), Data Explorers, research on AI data centers, and a research-design task that asks a model to propose a project and run a pilot. The models were GPT-6 Astra in Codex, Claude Fable 5.1 in Claude Code, Grok 4.6 in Grok Build, Gemini 3.8 Flash in Antigravity, Kimi K3 in Kimi Code, and Qwen 3.8 Max in Qwen Code. Each ran at its highest reasoning setting with the context Epoch says it would give a new hire, including a Figma file and its own Google account.
The grading is the weak point, and Epoch says so. A single Epoch grader scored each output against the group’s own employee standards, using rubrics built with its staff. Each model attempted each task once. The authors call the scores noisy and subjective, and they lean on qualitative observations instead. The text publishes no per-model rubric scores, so the only number comparing models is the Capabilities Index figure below. They also warn that the results say nothing general about what AI can do across all work.
Within those limits, the split is clear. Fable 5.1 and GPT-6 Astra were broadly tied at the top, and their edge came from the well-defined parts: computer use, coding, and data analysis. Earlier this year, frontier models failed at porting an article to Substack. This time they succeeded. Epoch even dropped one design task, restyling a Matplotlib chart into its house look, after Fable 5.1 produced Epoch-quality work, and says it now largely automates that step.
The open-ended work is where it broke, and the pattern is a judgement gap. Models could name a promising research direction but could not design an experiment that measured what it claimed. GPT-6 Astra asked a sharp question: when an agent gets no better with practice, did it gather poor evidence or misread good evidence? Its pilot capped model outputs at 4,096 tokens, so 61 of 280 responses were truncated before naming an input. Astra noticed, raised the budget, and scores recovered. Its write-up still presented the swing as “sensitivity to the acquisition budget.” Epoch calls that framing misleading, since the models never had a fair chance to answer.
In practice, the model finds the right question and then reports its own setup errors as discoveries. That is a different problem from a model that is simply wrong. It produces work that looks finished.
Taste failed too, even with reference material to hand. Fable 5.1 was shown Epoch’s design files and still produced a far denser diagram than the human designer’s simple chart, copying “A” and “B” badges from a complex example without grasping why they earned their place. GPT-6 Astra, told to write on AI in physics research, chose a topic Epoch judged too niche for its general audience. When asked for a Data Insight from its polling data, three of the six models settled on the same topic.
The open-weight result is a separate finding, and Epoch attributes it to its own testing. Open-weight models trailed further and failed even on tasks the frontier models handled reliably. Kimi K3 built an entire Data Insight on a filtering error, counting AI mentions across all arXiv papers instead of physics papers and never checking its filter. Epoch reports that its Data Insights from every frontier closed-weight model were free of factual errors. Kimi K3 scores 158 on the Epoch Capabilities Index, roughly level with Grok 4.6, so the benchmark score hid the gap.
For a team selling or buying knowledge-work automation, Epoch’s account suggests a division: automate the steps that have a checkable answer, and keep a senior reviewer on everything that requires taste or a decision about what a result means. That reviewer is the cost most automation pitches leave out, including from the startups raising money this week on the promise of the opposite.
Based on “Can AI automate Epoch?” by Kelly Hong and Greg Burnham, published by Epoch AI on 8 October 2026.