Examples
Six scenes from the appendix. Compare each reference image with reconstructions from six coding agents.
From image to code: a coding agent writes, renders, and revises a Blender program to reconstruct an editable, queryable 3D scene.
Write code. Render the scene. Inspect and refine.
A coding agent turns the reference image into a Blender scene through planning, code execution, rendering, and iterative validation.
GPT-6-astra trajectory. Four blocks define the process: the Agent Workspace manages artifacts, the Action Space exposes agent operations, the Blender Code Structure organizes editable scene state, and the Scene Construction Workflow iterates from reference interpretation to validated final artifacts. Together, they enable executable, feedback-driven scene reconstruction.208 images from 104 scenes, spanning 8 environments, 17 themes, and 443 assets. Indoor and outdoor scenes use nested Easy ⊂ Medium ⊂ Hard object sets with fixed architecture, camera, and lighting. Agents receive one RGB image and public metadata; geometry, depth, and instance masks remain private.
Overall = mean[V × (R + A) / 2]. All attempted cases count; invalid submissions and unresolved evaluation failures score zero.
Overall score (%) · Mean ± standard deviation over three runs.
| Coding agent | Indoor ↑ | Outdoor ↑ |
|---|---|---|
| GPT-6-astra | 53.4 ± 0.8 | 39.6 ± 0.8 |
| GPT-6-sol | 32.3 ± 1.0 | 24.2 ± 0.4 |
| GPT-6-luna | 23.2 ± 0.7 | 17.3 ± 0.5 |
| GPT-5.6-sol | 14.8 ± 1.0 | 15.3 ± 0.8 |
| GPT-5.6-terra | 15.4 ± 0.6 | 11.5 ± 0.4 |
| GPT-5.6-luna | 14.1 ± 1.3 | 12.1 ± 1.0 |
GPT-6-astra leads on both splits. High artifact validity still leaves substantial gaps in geometric and visual fidelity, and outdoor scenes are harder at every complexity tier.
GPT-6-astra reaches an evaluable scene earlier and refines it steadily, whereas GPT-5.6-sol starts later and regresses more often. The gap shows that reliable construction depends on fast initialization and protection against regressive edits.Training-free initialization, grounded refinement, and version control. Improves all six evaluated models, with up to 62.7% relative gain on the Office subset.
Relative overall-score gain · 42-case Office subset.
One reconstructed scene yields object detections, instance masks, and depth—through deterministic queries, without task-specific training.
| Task / dataset | Metric | LEGO-Anything | Specialized model |
|---|---|---|---|
| Detection COCO | Box AP ↑ | 30.14 | 59.88 DINO |
| Segmentation LVIS | Mask AP ↑ | 14.75 | 53.96 Segment Anything 3 |
| Depth ETH3D | AbsRel ↓ | 0.1554 | 0.0783 Depth Anything 3 |
GPT-6-astra, without LEGO-Plugin · 100 images per task. AP uses equal-confidence predictions.
Scene readouts support all three tasks, with a clear gap to specialized vision models.
Explore the reconstructions and the measurements behind them.
Six scenes from the appendix. Compare each reference image with reconstructions from six coding agents.
Look inside two recorded evaluations. Inspect visible-surface geometry, object scores, and rendered appearance.
@misc{li2026legoanything,
title = {LEGO-Anything: Coding Agents for 3D Scene Reconstruction},
author = {Li, Xirui and Shi, Peng and Dong, Mingwen and
Zhang, Sheng and Xu, Zhuoyan and Lee, Dongkyu and
Chang, Shuaichen and Xiang, Yi and Pan, Lin and Jiang, Jiarong},
year = {2026},
month = sep,
note = {Preprint},
url = {https://xirui-li.github.io/lego-anything-website/}
}