RoboCurve gave OpenAI’s GPT-6 Astra control of a pair of bimanual YAM robot arms and ran it through the same two manipulation tasks it had already used to grade Anthropic’s Claude Fable models. The independent evaluator published the results on September 4. The headline number, a 95 percent completion rate on one task, is not the story. The 10 percent completion rate on the other one is.

Task one asked Astra to pick up a red block and drop it into a bowl. It succeeded in 19 of 20 trials, against 8 of 20 for Fable 5.1 and 1 of 20 for the original Fable 5, according to RoboCurve. Astra also finished faster, at roughly 2.5 minutes per run against Fable 5.1’s 6.8, and at an estimated $0.94 per run against Fable 5.1’s $2.12.

Task two put the same arms on a circular blue puzzle piece, which had to be lifted by the small knob at its middle and then seated into the matching round recess. Astra completed that insertion 2 times in 20, the same rate as Fable 5.1 and better than Fable 5’s zero. RoboCurve’s own scoring rubric, which runs from 0 (no purposeful approach) to 4 (placed), shows why the gap matters: on the puzzle task, Astra’s mean stage reached was actually lower than Fable 5.1’s, even though the two models tied on completions. RoboCurve describes Astra reaching the groove and stalling at the same final step where Fable stalls.

The two tasks look similar on paper: pick something up, put it somewhere. They are not the same problem. Dropping a block into a bowl only requires the gripper to close on a roughly graspable object and release it inside a wide target area, a task that tolerates a fair amount of positional slop. Seating a puzzle piece into a groove requires the gripper to find one specific feature, the knob, grasp it at a precise point, and then rotate the piece into a matching orientation before insertion succeeds. Fail the alignment step and the piece never goes in, no matter how close the arm gets. That is a fine-motor problem, not a reach-and-drop problem, and RoboCurve’s data suggests neither lab’s model has solved it yet.

This is worth stating plainly because most published robotics benchmarks skew toward tasks that resemble the bowl problem rather than the puzzle problem: move an object into a container, a bin, a general region. Coarse pick-and-place tasks are easier to build test rigs for and easier to score, which is exactly why headline manipulation numbers keep outrunning what these systems can do when a task demands precise grip and orientation control in a real workspace.

Each task ran 20 trials per model, a sample small enough that a single additional success or failure would move the completion rate by 5 points. RoboCurve’s cost figures are its own measurement at list price ($10 and $50 per million input and output tokens across all three models), not an audited or third-party benchmark, and the comparison has real gaps: the bowl trials for Astra and for the Fable models ran on different physical rigs, Astra’s runs happened two days after Fable’s rather than interleaved with them, and the human grader scoring each trial knew which model was running. RoboCurve also notes that OpenAI’s API automatically cached roughly a fifth of Astra’s input without a pricing discount applied in its calculation, meaning the $0.94-per-run figure is, if anything, an overstatement of what Astra actually cost to run.

Teams evaluating either model for warehouse or assembly-line pick tasks should treat the bowl result as evidence of competence on coarse placement and the puzzle result as the more relevant number for anything involving oriented parts, connectors, or small-feature grasping.

RoboCurve, an independent evaluation published September 4, 2026.