A Semantic Parsing Method for Indoor Scene Images Based on Prior Knowledge of Building Structure.
This paper proposes a semantic parsing method that leverages building-structure priors that uses a shifted-window hierarchical transformer encoder to extract multi-scale visual features and combines a gradient-direction-consistency line segment detection algorithm to construct a Manhattan 3D bounding box.