VLM-Based Automatic Multi-Granularity Graph Representation of Building Layouts for Design Informatics

Abstract


Architectural floorplan images encode rich relational knowledge among functional spaces, a resource that underpins design retrieval, knowledge-based reasoning, and Building Information Modeling (BIM) enrichment throughout the building lifecycle. Yet, automatically constructing task-adaptive graph representations for public buildings remains a significant challenge. To address this gap, we first define a multi-granularity Level-of-Graphs (LoGs) taxonomy specifically for public building layouts. Methodologically, we introduce a Vision-Language Model (VLM)-based automatic LoG construction pipeline that encompasses node identification, edge inference, text parsing, and graph coarsening. The VLM-generated representations are systematically evaluated and tested in real-world tasks, using 147 academic library floorplans from around the world as a case study. Our experiments demonstrate that VLM-generated graphs are broadly consistent with human-labeled graphs, achieving a matched node ratio of 92% or higher, with a generation time of 509.3 seconds per floorplan across three LoG levels. Notably, meso-grained graphs yield the best node-level zone prediction performance (Macro F1 = 0.647 at 65% of fine-grained complexity), while coarse-grained graphs prove most effective for graph-level layout quality evaluation (Spearman's ρ = 0.610 at 16% of fine-grained complexity). By enabling scalable, annotation-free extraction of structured layout information from floorplan images, this study advances design informatics—converting plan images into knowledge representations and thereby enhancing the utilization of design information across the building lifecycle.

via ArXiv AI

Related