Skip to content
#explainable ai Open access

子空间稳定、轴语义易位与连续谱:跨城市路网形态主成分表征的测量条件审计 (Stable Subspaces, Variable Axes, and the Morphological Continuum: Auditing Measurement Conditions in Cross-City Street-Network PCA)

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)
Urban Design and Spatial Analysis

Abstract

【目的】 跨城市路网形态比较常依赖开放地理空间数据(OSM)与主成分分析(PCA)构建低维表征,但既有研究多预设离散分类范式,且普遍忽视数据质量异质性、空间尺度推移与特征共线性对主成分轴语义的构造性干扰。本文旨在检验跨城市路网形态主成分表征的前提假设是否成立,对数据适用性、空间尺度与特征共线性等关键测量条件实施系统审计,评估低维形态空间的代数稳定性与解释边界。 【方法】 选取42个城市(含31个中国内地直辖市与省会首府、香港及10个国际基准城市)的126个嵌套OSM窗口(5×8 km、2×与4×面积),提取5项道路级形态描述符(蜿蜒度、归一化曲率、方向一致性、平面交叉度、邻接类型熵)执行PCA降维。建立面向描述符的三级数据适用性分类(参考适用13城、支持受限21城、表征受限8城);采用旋转不变经典RV系数与轴锚定余弦分别度量子空间代数稳定性与轴语义方向;通过Procrustes对齐合成城市形态空间;实施2几何对2拓扑的对称特征消融以剥离共线特征造成的方差虚假集中;并结合SRTM 30 m DEM与OSM三维层差代理(p3d),检验宏观地形与立体交通对拓扑份额的线性解释力。 【结果】 ① 在全部126个窗口中,49%(62个)触发至少一项质量风险标记,证实开放数据跨城形态推断需要前置分层;在低维空间中,城市路网形态呈现为平滑过渡的几何–拓扑连续谱而非离散群落,PC2在39/42个城市中一致承载拓扑轴,初评标记的“反相位类型”在参考类复核下未能独立成簇(载荷幅值极低且与几何组共簇),现有数据不支持离散划分。② 载荷子空间跨尺度表现出高代数稳定性(参考类经典RV中位数0.997),但轴语义在特定尺度易受特征值排序扰动(重庆5×8↔4×轴锚定余弦仅0.228对经典RV 0.983);且拓扑极点读数(重庆0.675、贵阳0.657)高度依赖长尾极端弯路,剔除单条道路即骤降至0.130与0.255。③ 对称特征消融表明:剔除方向一致性后,92.9%的窗口(117/126)与92.9%的城市(39/42)中PC1转为拓扑主导(37座城市发生主导轴角色易位),城际拓扑排序完全解耦(Spearman ρ = -0.142,p = 0.370)。④ 机制假说检验表明,宏观地形与城市拓扑份额无统计关联(ρ = +0.063,p = 0.690),3D连接代理在参考类中虽呈名义弱相关(ρ = 0.569,p = 0.042)但未通过Benjamini–Hochberg校正,局部环境属性无法线性解释全局网络结构。 【结论】 基线分析中主成分PC1的“几何主导”受制于共线几何变量方差聚集的测量协议构造。对称消融表明其轴语义具有框架依赖性。在本文道路级表征框架与适用性复核下,现有数据不支持普适性离散形态划分假设,城市路网形态更接近低维连续谱分布;城市在轴上的相对位次仍由真实形态数据驱动。基于开放数据的跨城市形态比较,需将数据适用性分层、空间尺度外推、轴语义易位与尾部极端样本敏感性纳入前置边界条件,以防范将测量构造误读为城市空间规律。 ================================ Stable Subspaces, Variable Axes, and the Morphological Continuum: Auditing Measurement Conditions in Cross-City Street-Network PCAAuthor: Wenhao WANG*Affiliation: School of Geography and Urban-Rural Planning, Longdong University, Qingyang 745000, ChinaCorresponding Author: Wenhao WANG, E-mail: gew500781@gmail.com [Objectives] Cross-city comparative analyses of urban street networks frequently employ OpenStreetMap (OSM) data and principal component analysis (PCA) to extract low-dimensional morphological representations. However, existing studies commonly presuppose discrete typologies while overlooking how data quality heterogeneity, spatial scale shifts, and descriptor collinearity structurally distort low-dimensional representations. This study systematically audits the measurement conditions and foundational assumptions underlying PCA-derived street network spaces, evaluating the algebraic stability of subspaces, the semantic consistency of dominant axes, and their empirical interpretive boundaries across spatial scales.[Methods] Using 126 nested OSM windows (5×8 km, 2× area, and 4× area) across 42 global cities (31 Chinese provincial capitals and municipalities, Hong Kong, and 10 international benchmarks), we extracted five road-level descriptors (sinuosity, normalized curvature, directional consistency, planar degree, and adjacency-type entropy) to perform PCA. We instituted a three-tier descriptor-specific fitness-for-use classification (reference: 13 cities; support-limited: 21; representation-limited: 8) and decoupled cross-scale structural stability into two analytical tiers: the rotation-invariant classical RV coefficient for loading subspaces and the axis-anchored cosine for directional semantic fidelity. City-level morphology spaces were synthesized via Procrustes alignment. To diagnose measurement framework artifacts, we conducted a symmetric feature ablation (two geometric versus two topological descriptors) by removing collinear directional consistency. Finally, we evaluated macro-terrain (SRTM 30 m DEM) and vertical connectivity (OSM 3D layer-difference proxy p3d) hypotheses to assess whether local environmental attributes linearly explain city-level topology share.[Results] (1) Across all 126 windows, 49% (62 windows) trigger data quality risk flags, demonstrating that open-data cross-city morphological inference requires systematic upfront sample conditioning. In the low-dimensional space, urban street networks occupy a smooth geometric–topological continuum rather than forming discrete clusters; PC2 consistently serves as the primary topological axis across 39 of 42 cities, while a hypothesized "counter-phase" typology fails to form an independent cluster upon reference-class re-examination (manifesting low loading magnitude and co-clustering with the geometric group), rejecting discrete clustering. (2) Loading subspaces exhibit high cross-scale algebraic stability (median classical RV 0.997 among reference-class pairs), yet axis semantics remain vulnerable to eigenvalue reordering (e.g., Chongqing 5×8↔4× exhibits an axis-anchored cosine of only 0.228 alongside classical RV 0.983). Furthermore, topological pole scores (Chongqing 0.675, Guiyang 0.657) are governed by extreme descriptor tails, collapsing to 0.130 and 0.255 upon removing a single most sinuous road. (3) In symmetric feature ablation, removing directional consistency caused PC1 to flip to topological dominance in 92.9% of windows (117/126) and 92.9% of cities (39/42; 37 cities underwent dominant-axis role reversals), completely decoupling inter-city topological rankings (Spearman ρ = -0.142, p = 0.370). (4) Macro-terrain exhibits no statistical association with topology share (ρ = +0.063, p = 0.690), and the 3D proxy's nominal correlation in the reference class (ρ = 0.569, p = 0.042) fails Benjamini–Hochberg false-discovery-rate control, demonstrating that macro-terrain does not drive city-level topology share; furthermore, road-level environmental attributes fail to explain city-level morphology linearly (mean regression R² < 0.163).[Conclusions] The apparent geometric dominance of PC1 in baseline street-network representations is constrained by the measurement protocol through collinear variance concentration; symmetric ablation demonstrates that axis semantics are framework-dependent, while relative city positions along these axes remain driven by empirical morphological data. Under this road-level PCA framework and fitness-for-use re-examination, the data provide no empirical support for universal discrete typological classifications, with street networks instead following a low-dimensional continuum. Cross-city morphological comparisons using open data must treat fitness-for-use, spatial scale, axis instability, and descriptor-tail sensitivity as indispensable boundary conditions, to prevent observational artifacts from being mistaken for universal urban regularities. Key words: street network / OpenStreetMap / principal component analysis / subspace stability / axis semantics / cross-city comparison / spatial scale / measurement uncertainty基金项目: 无作者简介: 王文浩,陇东学院地理与城乡规划学院。E-mail: gew500781@gmail.com —— v1.0.1(2026-09-12):表3消融列三处数字经复核后修正并由落盘产物锚定:基线行 PC1 几何主导 123/126→120/126 (95.2%);剔除 E_entropy 行 123/126→126/126、城级 40/42→42/42、ρ +0.676→+0.675;§4.3 主导判据(123/126,例外为海口三尺度)与幅度阈值判据(124/126 即 98.4%)拆分陈述。—— v1.0.2(2026-09-16):二轮修改稿,新增三组经落盘产物锚定的内容:①固定队列对照(126 对 RV 中位 0.999、最低 0.881,9 对轴锚定余弦 <0.9,重庆 2×↔4× −0.9985);②重庆 5×8 窗口命名走廊实际跨度 7.3×11.4 km(约 2.08 倍名义面积);③道路分段制图粒度检验(ρ=+0.624/+0.535/参考类 +0.069)。—— v1.0.3(2026-09-17):三轮修改稿——全文语言精修、标题升版、表3 消融对照移入 §4.3;复现脚本口径与 AI 工具声明规范化;数值零改动。 —— v1.0.4(2026-09-20):版本升版发布;英文摘要 (4) 句新增量化上限——局部环境对几何形态的平均回归 R² < 0.163(正文 §4.5 同行印 0.163,产物锚定城级最大值 0.1625)。 The full standalone reproduction package (816 files, verified via metric contracts and ledger assertions) is bundled in limit-sense-repro-v1.0.4.tar.gz (SHA-256: 9d4efb0ca8641692d2f1fb9b5b2c905a78f5dd7e03a7253ae56bc762b53e2d25).

View source

Similar papers

#artificial intelligence Conference Open access Apr 2020

ECCOLA - a Method for Implementing Ethically Aligned AI Systems

The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.

Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson · 64 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Open access Mar 2024

LLM-based agents for automating the enhancement of user story quality: An early report

The use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams is explored, with a reference model for an Autonomous LLM-based Agent System developed and implemented at the company.

Zheying Zhang, M. Rayhan, Tomas Herda et al. · 48 citations · ⚡4
#computer vision Review Mar 2024

System for systematic literature review using multiple AI agents: Concept and an empirical evaluation

This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.

Abdul Malik Sami, Z. Rasheed, Kai-Kristian Kemell et al. · 44 citations · ⚡2
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 41 citations
#artificial intelligence Conference Open access Jun 2018

The Key Concepts of Ethics of Artificial Intelligence

It is suggested that the focus on finding keywords is the first step in guiding and providing direction for future research in the AI ethics field.

Ville Vakkuri, P. Abrahamsson · 39 citations · ⚡2

Related blog posts

Google DeepMind Blog Sep 30, 2026

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.