Scale Labs, the research arm of Scale Inc., published a benchmark on 7 October showing that the strongest AI vision model it tested gets barely more than half of a set of everyday visual judgments right, while people get nearly all of them. Human participants reached 93.1 percent accuracy. GPT-6-astra, run at maximum reasoning effort, reached 53.6 percent. We are describing the abstract of the paper, Humanity’s Sixth Sense; we have not read the full text.
The gap works out to 39.5 points. Maximum reasoning effort means the model spends the most computation thinking before it answers, so this was its best shot. More thinking time did not buy the model the quick read that people bring to a scene without trying.
The authors argue that existing visual benchmarks test either slow expert analysis or low-level perception, and skip the snap inference people make without noticing. They give three examples. One look at a scene tells you what just happened and roughly what comes next. A quick glance at a gap tells you whether your car fits between two parked ones. A few seconds of video can show you who is really in charge. The benchmark covers images and video, with human-written prompts probing time, space, social relationships, and abstraction.
Scale Inc. sells data and model evaluation. A benchmark showing frontier models failing a new axis of testing serves a company in that business, which is worth remembering when you read the headline number. That is no reason to doubt the result, only to wait for other labs to run the test.
The authors also tried letting a model manipulate the image as it works, an agentic setup. According to the abstract, that shrinks the gap without erasing it. No figure is given.
The reason this matters outside a leaderboard is where these judgments get used. The paper says current models excel at tasks needing advanced perception and knowledge, and that profile wins deals: a system that reads a dense diagram looks ready for more. But a warehouse robot judging whether a pallet clears a doorway, an inspection camera sensing a stack about to shift, and a driver-assist system guessing what a pedestrian does next all lean on the fast inference this benchmark targets. That link is our inference, not something the abstract tests. Nobody ran these models on a loading dock.
Buyers of vision systems for physical work should read strong knowledge-test scores as silent on this skill. If your use case depends on spatial or social judgment from camera input, build a small test from your own footage, score it against your own staff, and keep a person in the loop until the model’s number comes close to theirs.
Based on the abstract of “Humanity’s Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models,” published by Scale Labs (Scale Inc.) on 7 October 2026.