LEGO-Anything: How AI Turns a Photo Into Editable 3D Scenes

Wohn- und Essbereich mit grünem Sofa, Tisch und Küchenzeile
Photo by Darren Ahmed Arceo on Unsplash

One photo as the starting point, an editable 3D scene as the result: LEGO-Anything explores how AI agents can make that translation through Blender code. The research paper, involving researchers at the University of Maryland and AWS, was published on arXiv on September 28, 2026, and is now drawing public discussion. The interesting part is less another attractive preview than the ability to change objects, the camera, and the room layout afterward.

Key takeaways

  • LEGO-Anything has an agent write, execute, render, and iteratively revise Blender code.
  • LEGO-Bench evaluates 208 images from 104 simulated scenes. Working files and fidelity are assessed separately.
  • GPT-6-astra with Codex scores 53.4 percent overall indoors and 39.6 percent outdoors. These are not general accuracy rates for real photographs.
  • An additional control layer called LEGO-Plugin improves the evaluated agents on a limited office subset without retraining the models.
  • The project page still lists the code as coming soon. The research project is therefore not yet a complete downloadable tool for everyday use.

Why representing a scene as a program opens more possibilities

A computer-generated view can depict a room convincingly without providing a useful spatial description. That makes a major difference to subsequent work. Anyone who wants to move a table, reposition a camera, or replace individual objects needs more than the pixels in an image. An executable scene program records those elements explicitly and can render another image after a change.

LEGO-Anything uses a general-purpose coding agent for this task. It interprets the reference image, plans the scene, writes Blender code, and examines the rendered result. It can then revise the construction. This loop connects a natural-language task with an inspectable artifact: People can read the program, change parts of it, and check the effects again. The result remains editable in principle, rather than existing only as a finished view.

That is an interesting direction for design and visualization. As a possible workflow example, a rough room reconstruction could help people discuss different furniture arrangements. Whether its measurements are accurate would be a separate question. The paper does not establish survey-grade capture of homes or automatic production readiness for games or architecture. This distinction is precisely what makes the approach useful: It explicitly evaluates how far the executed scene still differs from the visible reference.

What the percentage scores actually measure

The associated evaluation suite, LEGO-Bench, uses images from fully specified simulator scenes. The evaluator therefore knows the correct geometry and object assignments, while the agent does not receive that information. The dataset includes 443 registered assets, or reusable scene components. Obtaining an equally precise reference for real photographs would be much harder. The simulated foundation enables robust measurements within the experiment, but limits their transfer to arbitrary phone photos.

The evaluation separates three questions. First, are the scene, export, and rendered image usable at all? Second, how closely do visible object surfaces match the reference? Third, how similar does a newly rendered image look? The overall score combines the latter two measures; unusable submissions receive zero. A scene can therefore be generated successfully while still getting objects, distances, or camera angles substantially wrong.

GPT-6-astra with Codex leads the evaluated configurations. The reported 53.4 percent indoors and 39.6 percent outdoors are averages over three runs. They mean neither that half of all rooms were reconstructed correctly nor that every distance is 53.4 percent accurate. The researchers measure a composite score under their experimental conditions. Outdoor environments and more complex scenes remain harder.

Any model ranking should also be limited to this task. The experiment measures the combination of model, agent environment, and 3D tool, rather than an isolated ability covering every form of spatial reasoning. A higher benchmark score also does not prove time savings in a subsequent design project. That would require knowing how much manual rework the result still needs there.

Why revisions sometimes make things worse

The authors examine not just the finished files but also how they were produced. Recurring difficulties include weak initialization, later changes that damage previously correct elements, and unreliable assessment of the agent’s own output. On geometry in particular, model judgments remain near or below chance. Asking the agent to inspect the scene carefully one more time is therefore no substitute for a reliable comparison.

LEGO-Plugin adds image-grounded initialization, concrete construction feedback, and version control to the working environment. A change is checked and can be repaired or rolled back. This is not a new generation of models. The control layer intervenes in the process and uses information from the reference image; it has no access to the benchmark’s hidden correct scene.

On the 42-case office subset, this addition improves all six evaluated models. The largest relative gain is 62.7 percent, while the already strongest model gains 2.1 percent. Both figures refer to the respective baseline and this particular experiment. They are neither additional percentage points nor a promise covering all indoor and outdoor scenes. Instead, the finding suggests that better workflows can especially help weaker agents.

Who can benefit from the approach now

For 3D developers and research groups, the inspectable intermediate representation is especially interesting. The team also derives object detection, segmentation, and depth estimation from reconstructed scenes. Results remain behind specialized vision models. Even so, the paper shows that the same spatial description can answer different questions, instead of generating a completely new result for each one.

Anyone interested in practical experiments with agent-assisted Blender should distinguish this research project from existing integrations. The separate community project MCP for Blender can manipulate objects, change materials, and execute Python code in Blender; it is not made by the Blender Foundation. Its availability does not mean that LEGO-Anything or LEGO-Plugin has already been fully reproduced through it. For local experiments, it is also useful to review agent access permissions on the Mac before giving a working environment access to personal projects.

For an introduction, the project page already provides comparison images and a metric viewer. That is currently the most concrete way to understand the approach’s strengths and errors. The next practical step depends on the announced code release and on whether editable rough scenes can become useful designs with a reasonable amount of rework. The real advance is the inspectable artifact: An image becomes the starting point of a construction whose individual parts people can continue to work on.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top