The Decoder reports that researchers at the University of Maryland and AWS built LEGO-Anything. A coding agent receives one photo, writes a Blender program, runs it, looks at the result, and revises until the scene is closer to the original. The output is a program that can be inspected and edited. It states objects, geometry, layout, and camera position. The article gives these results before a subscription wall. What sits behind the wall is not included here.

[1]
A camera on a tripod faces a small sofa built from wooden blocks, with a red brick lying outside it.
The camera faces a block-built seat, and a red brick sits outside it. That is a scene that was built, with the geometry still off. An illustration, not a news photograph., AI-generated illustration, not a news photograph

The accompanying LEGO-Bench has 208 images from 104 indoor and outdoor scenes and uses 443 registered assets. The images are rendered from simulator scenes that were already built, so geometry, depth, and object assignments can serve as the answer key, while the inputs still look fairly natural. The score has three parts: whether a usable scene was delivered, how accurate the visible geometry is, and how close a re-render looks to the original. All six tested GPT configurations delivered a working scene almost every time. The best, GPT-6 Astra, scored 53.4% on indoor scenes and 39.6% on outdoor scenes. Weaker configurations scored around 15%. When the researchers increased the reasoning budget, Astra’s score on an office subset rose from 32.3% to 61.8%.

[1]

The article says the common problems are a poor first attempt, revisions that undo earlier progress, and unreliable self-assessment. When models had to pick which of two versions better matched the original, their geometric judgments landed near or below chance. The researchers conclude that refinement should use concrete measurements, not the agent’s own judgment. They built LEGO-Plugin, which needs no extra training. It anchors the start in the reference image, replaces self-judgment with measurements, and blocks edits that make the scene worse. All six models improved. Weaker agents gained the most, up to 62.7 percentage points. The already strong top model gained only about two percentage points. Used directly for detection, segmentation, and depth, the scenes reached about half of the specialized model DINO on detection, with a larger gap on segmentation and depth. The authors say a working result and a faithful reconstruction are still far apart.

[1]

要点

  • An agent writes a Blender program from one photo and revises it. The benchmark has 208 images and 104 scenes.
  • GPT-6 Astra scores 53.4% indoors and 39.6% outdoors. Judging whether its own geometry improved lands near or below chance.
  • After measurements replace self-judgment, weaker models gain more. The top model gains about two points.