Fig, an AI infrastructure startup, tested seven frontier models on the same 177 web tasks and found that no model beat all the others. In a technical report on its own site dated 29 September, the company says the model with the best average score still failed tasks that weaker rivals completed.

The top of the table was a near tie. OpenAI’s GPT-6 Astra solved 139 of the 177 tasks (78.5 percent) and Fable 5 solved 138 (78.0 percent), a gap of one task that sits inside the roughly three-point noise band Fig measured by rerunning Opus 4.7. Opus 5.5 solved 124, and Kimi K3 came last with 102. The tasks come from VisualWebArena, an open benchmark of browser jobs on shopping, classifieds and Reddit-style sites.

Fig’s sharper finding is about overlap. In all 21 pairings of the seven models, the lower scorer solved at least three tasks the higher scorer missed, and in the widest pairing it solved 21. Kimi K3 trailed Astra by 20.9 points yet still cracked six tasks Astra could not. All seven models solved 66 tasks and none solved 21, so the ranking was decided entirely by the 90 contested ones. Taken together, the seven models solved 156 tasks, 17 more than Astra alone.

The site mattered about as much as the model. Fig reports that Opus 4.7 scored 79 on classifieds, 73 on Reddit and 55 on shopping, a 23.9-point swing across one benchmark. Astra swung 18 points (77, 91 and 73), while Fable 5 was steadiest at 10.4. The benchmark’s own easy, medium and hard labels predicted model success only weakly.

Model updates behaved the same way. Going from Opus 4.7 to Opus 5, the average score fell just 1.1 points, but 36 individual tasks changed outcome: Opus 5 gained 17 and lost 19, and 11 of 21 task categories got worse. Going from GPT-5.6 to Astra, the average rose 7.9 points, and Fig calls that gain statistically distinguishable from chance. Even then, two categories slipped, and one of them, tasks that ask for the cheapest or largest item in a set, also regressed in the Opus step.

The physical-world tests told a similar story. Astra and Opus 5.5 both beat Google’s Gemini 3.5-Flash on the average of all four offline domains, a group that spans robot handling, building things, factory routines and road driving. Across the 185 task categories those domains define, Astra scored below Gemini on 16 and Opus 5.5 on 26.

Fig lists its own limits. Each model ran once per task, so variance was not measured directly. The Opus models used computer-use actions while the others used function calling, which muddies comparisons of absolute scores. Kimi K3 sometimes ran out of output budget before picking an action. The benchmark’s automatic checker also rejected correct answers such as “4200” written without a comma, though Fig says re-scoring those flipped only 4 to 9 episodes per model and left the ranking intact. The authors also state they did not study why the unevenness occurs. Fig has released its item-level results, called RIDGE, so others can check the numbers.

The report is a small, single-pass study published on the company’s own site, and it should be read that way. The takeaway for anyone choosing a model for a browsing agent is still practical: a leaderboard average hides which tasks a model drops, and an upgrade that lifts the average can break a workflow you rely on. The 156 versus 139 gap suggests that a system able to pick the right model per task would beat any single one, so re-run your own task list on each new release before you switch.

Reported by Fig in its technical report published on 29 September 2026.