OpenAI's GPT-6 Astra scores 80% on Epoch AI's Furniture Assembly Benchmark (FAB). The benchmark photographs three IKEA pieces during assembly and leaves deliberate errors in the photos. A model compares the photos with the instructions, identifies the mistakes, and describes what went wrong.
The Decoder reports that the best score in November 2025 was Claude Opus 4.5, at 28%. Ten months later, Astra reaches 80% at about three minutes per photo. Claude Fable 5.1 scores 70%, and Claude Opus 5 scores 61%. Chinese open-weight models such as Kimi K3 trail the leaders by at least seven months.
[1]
It is still too slow for real-time help while someone is assembling the furniture. The researchers say the technique could eventually assist with car repairs or appliance fixes. The report also says Astra excels at visual robotic tasks, and it does not give a score or name those tasks.
The piece does not say which three pieces were photographed, how the errors were planted, or how many photos sit behind the 80%. Epoch AI's benchmark page did not load for this edition, so the figures are The Decoder's account of the benchmark.
[1]要点
- FAB photographs three IKEA pieces during assembly and plants deliberate errors.
- In November 2025 the best score was Claude Opus 4.5 at 28%.
- GPT-6 Astra scores 80% at about three minutes per photo; Fable 5.1 scores 70% and Opus 5 scores 61%.
- That pace is still too slow for real-time assembly help.