SEPTEMBER 2026

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

One image. A coding agent. An editable 3D world.

Xirui Li1,2,*Peng Shi2Mingwen Dong2Sheng Zhang2Zhuoyan Xu2
Dongkyu Lee2Shuaichen Chang2Yi Xiang2Lin Pan2Jiarong Jiang2

1 University of Maryland, College Park2 AWS

* Work done during an internship at AWS

Overview

From image to code: a coding agent writes, renders, and revises a Blender program to reconstruct an editable, queryable 3D scene.

One executable scene, with editable objects and deterministic vision readouts.

Write code. Render the scene. Inspect and refine.

LEGO-Anything

A coding agent turns the reference image into a Blender scene through planning, code execution, rendering, and iterative validation.

Agentic scene construction, abstracted from a GPT-6-astra trajectory. Four blocks define the process: the Agent Workspace manages artifacts, the Action Space exposes agent operations, the Blender Code Structure organizes editable scene state, and the Scene Construction Workflow iterates from reference interpretation to validated final artifacts. Together, they enable executable, feedback-driven scene reconstruction.

LEGO-Bench

1 / 2

Dataset

208 images from 104 scenes, spanning 8 environments, 17 themes, and 443 assets. Indoor and outdoor scenes use nested Easy ⊂ Medium ⊂ Hard object sets with fixed architecture, camera, and lighting. Agents receive one RGB image and public metadata; geometry, depth, and instance masks remain private.

Extensible benchmark-construction pipeline.

Evaluation

Validity (V)
Usable Blender scene, GLB export, and render; the scene must contain geometry and an active camera.
Reconstruction (R)
Object-averaged visible-surface F1 at a depth-scaled 5% tolerance, with no alignment or rescaling.
Appearance (A)
Fraction of pixels within 30/255 error in every RGB channel, using an evaluator re-render.

Overall = mean[V × (R + A) / 2]. All attempted cases count; invalid submissions and unresolved evaluation failures score zero.

LEGO-Bench

2 / 2

Results

Overall score (%) · Mean ± standard deviation over three runs.

Overall LEGO-Bench scores, mean and standard deviation over three runs. Higher is better.
Coding agentIndoor ↑Outdoor ↑
GPT-6-astra53.4 ± 0.839.6 ± 0.8
GPT-6-sol32.3 ± 1.024.2 ± 0.4
GPT-6-luna23.2 ± 0.717.3 ± 0.5
GPT-5.6-sol14.8 ± 1.015.3 ± 0.8
GPT-5.6-terra15.4 ± 0.611.5 ± 0.4
GPT-5.6-luna14.1 ± 1.312.1 ± 1.0

Findings

GPT-6-astra leads on both splits. High artifact validity still leaves substantial gaps in geometric and visual fidelity, and outdoor scenes are harder at every complexity tier.

LEGO-Plugin

1 / 2

Trajectory analysis

Scene quality over construction time. GPT-6-astra reaches an evaluable scene earlier and refines it steadily, whereas GPT-5.6-sol starts later and regresses more often. The gap shows that reliable construction depends on fast initialization and protection against regressive edits.
First evaluable scene
GPT-6-astra: ~10% of budget; GPT-5.6-sol: ~20%.
Regressive edits
29.6% of GPT-5.6-sol’s edits decrease the score.
Lost progress
GPT-5.6-sol finishes 3.2 points below its best scene.

LEGO-Plugin

2 / 2

Training-free initialization, grounded refinement, and version control. Improves all six evaluated models, with up to 62.7% relative gain on the Office subset.

With and without the plugin

GPT-5.6-luna
+62.7%
GPT-5.6-terra
+55.8%
GPT-5.6-sol
+55.3%
GPT-6-luna
+27.5%
GPT-6-sol
+12.1%
GPT-6-astra
+2.1%

Relative overall-score gain · 42-case Office subset.

LEGO-World

One reconstructed scene yields object detections, instance masks, and depth—through deterministic queries, without task-specific training.

LEGO-World results compared with specialized vision models.
Task / datasetMetricLEGO-AnythingSpecialized model
Detection COCOBox AP ↑30.1459.88 DINO
Segmentation LVISMask AP ↑14.7553.96 Segment Anything 3
Depth ETH3DAbsRel ↓0.15540.0783 Depth Anything 3

GPT-6-astra, without LEGO-Plugin · 100 images per task. AP uses equal-confidence predictions.

Scene readouts support all three tasks, with a clear gap to specialized vision models.

Explorer

Explore the reconstructions and the measurements behind them.

Examples

Six scenes from the appendix. Compare each reference image with reconstructions from six coding agents.

Metric viewer

Look inside two recorded evaluations. Inspect visible-surface geometry, object scores, and rendered appearance.

Citation

@misc{li2026legoanything,
  title   = {LEGO-Anything: Coding Agents for 3D Scene Reconstruction},
  author  = {Li, Xirui and Shi, Peng and Dong, Mingwen and
             Zhang, Sheng and Xu, Zhuoyan and Lee, Dongkyu and
             Chang, Shuaichen and Xiang, Yi and Pan, Lin and Jiang, Jiarong},
  year    = {2026},
  month   = sep,
  note    = {Preprint},
  url     = {https://xirui-li.github.io/lego-anything-website/}
}

Explorer

Loading explorer…

With and without the plugin

Reference image
Without plugin14.7
With plugin32.5

GPT-5.6-sol · Overall score (%) · Baseline: best of three runs.

LEGO-Bench findings

How well do coding agents reconstruct 3D scenes?

GPT-6-astra achieves the best overall score, surpassing task-specific and single-image scene-construction baselines. Coding agents deliver valid artifacts across indoor and outdoor scenes.

Mean ± standard deviation (%) over three runs. Higher is better.

LEGO-Bench indoor results. Mean and standard deviation over three runs.
Model / systemValidity ↑Reconstruction ↑Appearance ↑Overall ↑
GPT-6-astra + Codex100.0 ± 0.052.4 ± 1.154.4 ± 0.753.4 ± 0.8
GPT-6-sol + Codex99.4 ± 1.122.2 ± 2.442.5 ± 0.332.3 ± 1.0
GPT-6-luna + Codex99.7 ± 0.515.8 ± 1.230.5 ± 0.323.2 ± 0.7
GPT-5.6-sol + Codex97.8 ± 2.39.8 ± 2.119.8 ± 0.614.8 ± 1.0
GPT-5.6-terra + Codex99.7 ± 0.59.5 ± 1.021.3 ± 0.515.4 ± 0.6
GPT-5.6-luna + Codex99.1 ± 0.010.2 ± 2.118.0 ± 0.714.1 ± 1.3

Performance on LEGO-Bench. Values follow the manuscript’s main results table. Unsupported outdoor settings are marked “—”.

Evaluation metrics
Validity (V)
Whether the submission contains usable Blender, GLB, and rendered-image artifacts.
Reconstruction (R)
Object-averaged visible-surface F1 at a depth-scaled 5% tolerance, without alignment or rescaling.
Appearance (A)
The fraction of pixels whose evaluator-render error is at most 30 in every 8-bit sRGB channel.
Overall (S)
The mean of V × (R + A) / 2 over all attempted cases; invalid artifacts and unresolved evaluation failures receive zero.

Findings 1 of 3