A new robot-arm comparison points to a familiar but commercially important pattern in embodied AI: a model can become highly effective at a bounded handling task without yet solving fine-contact assembly.
Robocurve tested OpenAI’s GPT-6 Astra against Claude Fable 5 and Fable 5.1 on the same YAM robot arms, using the same Inspect Robots agent policy and two 20-trial tasks. The reported results show Astra completing a basic “pick up a red block and put it in a bowl” task in 19 of 20 trials. But it completed a more demanding puzzle-piece insertion in only two of 20 attempts—the same completion count as Fable 5.1.
The result that changes the economics
On the block-to-bowl task, Astra’s 95% completion rate substantially exceeded Fable 5.1’s 40% and Fable 5’s 5%. It also used materially less output: an average 2,100 output tokens per run, versus 12,900 for Fable 5.1 and 19,200 for Fable 5.
The report estimates Astra’s per-run cost at $0.94 and its duration at 2.5 minutes, compared with $2.12 and 6.8 minutes for Fable 5.1. Those are benchmark estimates rather than a full deployment cost, but the direction matters. If a robot agent can consistently execute a simple pick-and-place operation with fewer model calls, it improves both throughput and the operating case for supervised automation.
For warehouse, lab and light-manufacturing operators, the immediate implication is narrow rather than universal: constrained, repeatable transport tasks may be nearer to practical pilot territory than headline demonstrations imply. A high completion rate on a task with clear objects, destinations and success conditions could reduce the amount of human intervention required—provided performance holds under production variation.
Precision is still the dividing line
The second task required the arm to grasp a round blue puzzle piece by its central knob and insert it into a matching circular groove. Astra reached the groove but, according to the report, commonly stalled at the final insertion step. Its completion rate was 10%, identical to Fable 5.1’s, though its estimated cost per run was lower at $1.36 versus $2.18.
That distinction is operationally significant. Moving an object to a broad target primarily tests perception, planning and gross manipulation. Insertion adds tight tolerances, contact dynamics and the ability to react reliably when the physical world departs slightly from the planned trajectory. Many valuable industrial workflows depend on that latter capability.
The benchmark’s scoring rubric captures partial progress—from purposeful approach through contact, lifting, positioning and final placement—but partial success does not necessarily translate into unattended work. In a production setting, a robot that reliably reaches an insertion point but cannot complete the fit may still need a human, a specialized fixture or conventional force-control logic.
How to read the benchmark
The test offers useful evidence, not a broad ranking of robotics intelligence. It covers two tasks, 20 trials per model per task, a particular hardware setup and a shared agent policy. Its cost figures are estimates, and its completion scoring was performed by a human grader. The source also makes individual transcripts, videos and rerun files available, which helps make the reported runs inspectable.
Executives evaluating model-directed robotics should therefore ask for task-specific evidence rather than extrapolate from the 95% figure. Key questions include how performance changes with lighting, object placement, surface friction, wear, sensor noise, cycle-time limits and recovery requirements.
What to watch next
The next meaningful test is not another broad pick-and-place challenge. It is whether model-directed systems can improve the last few centimeters of contact-rich work: insertion, fastening, cable routing and recovery from small misalignments.
Until then, the business opportunity is likely to be in carefully designed cells where fixtures and workflows turn a real operation into something closer to the bowl task. The model’s apparent efficiency gain could matter there; the puzzle result is the reminder that physical reliability remains the constraint.




