Systeembewuste Edge Intelligence: Adaptief en Gedistribueerd Deep Learning voor O-RAN
Abstract
Wireless networks are undergoing a paradigm shift with the advent of 6G, in which artificial intelligence (AI) is becoming a necessity at the physical layer for tasks like spectrum sensing, interference mitigation, and adaptive resource allocation. The Open Radio Access Network (O-RAN) architecture enables this transition by disaggregating network functions and pushing intelligence to the edge. However, deploying deep learning (DL) models on resource-constrained O-RAN Radio Units introduces a critical challenge: balancing the computational demands of AI with the stringent latency, fronthaul, and hardware constraints of real-time wireless systems. Current cloud-optimised AI architectures, designed for centralised data centres with abundant compute, fail to meet these edge requirements. Furthermore, the dynamic nature of wireless channels and the need for distributed coordination across multiple nodes complicate practical deployment. While theoretical advances in edge AI exist, a gap remains between isolated machine learning models and their system-level integration in real O-RAN environments. This dissertation bridges the gap by introducing System-Aware Edge Intelligence, a design paradigm that jointly optimises deep learning models alongside the physical and architectural constraints of O-RAN. Validated through synthetic datasets, software defined radio experimentation, over-the-air measurements, and hardware profiling on commercial off-the-shelf (COTS) platforms, this work provides empirical evidence for the advantages of a system-aware approach aligned with emerging industry standards. To systematically address these challenges, the thesis adopts a bottom-up approach, progressing from individual edge node optimisation to network-wide distributed coordination. First, focusing on the individual edge nodes, the dissertation investigates how physical-layer DL models can dynamically adapt their computational complexity to varying signal conditions. To achieve this, it introduces Width-Wise Early Exiting (WWEE), an adaptive inference framework that scales its active parameter count based on instantaneous signal complexity. By integrating selective classification, WWEE can reliably identify and reject uncertain, low-SNR signals. This capability to abstain from processing unreliable data reduces the average computational load while preserving overall classification quality, demonstrating how algorithmic adaptivity can successfully align with strict node-level hardware limitations. Second, at the system level, the thesis explores how unique O-RAN functional splits can be exploited to enable efficient inference. It introduces OSIRIS, a representation-aware split inference architecture aligned with O-RAN Split 7-2x. OSIRIS sequentially evaluates multi-domain representations (time, frequency, and Channel State Information) and dynamically routes computation based on intermediate confidence. This co-design of inference pipelines with system-level functional splits enables sub-millisecond physical-layer inference on standard COTS hardware. Third, the work leverages collaboration between different edge devices by distributing physical-layer deep learning across cell-free O-RAN topologies. By evaluating centralised, hybrid, and fully distributed inference paradigms, the research demonstrates that local feature extraction and soft decision fusion can maintain competitive accuracy relative to centralised baselines. Crucially, this collaborative approach mitigates limitations of single-node sensing (e.g., coverage gaps, interference blind spots) and enables the dynamic reallocation of compute resources across the network, all without overwhelming fronthaul capacity. Finally, to make this distributed collaboration practically viable, the dissertation addresses the critical need for low-overhead physical-layer coordination. It introduces localised machine learning frameworks to maintain temporal synchronisation and verify spatial event matching, ensuring that geographically separated nodes observe the exact same physical transmission. By shifting the coordination burden from continuous network signalling to predictive local computation, these solutions reduce over-the-air synchronisation overhead and enable fronthaul compression, thereby facilitating efficient ad hoc coordination. Overall, this dissertation explores structural methodologies for integrating AI into 6G networks more efficiently, suggesting a transition from treating machine learning as an isolated external tool to co-designing it with the wireless system itself. By addressing adaptive computation, representation-aware inference, distributed collaboration, and practical coordination in a unified framework, this work establishes a practical foundation for scalable, AI-native wireless systems tailored for the extreme edge.