Online Adaptive Learning for Fair Joint Hybrid Precoding and Task Scheduling in Dynamic Wireless - Computing Power Networks
Emerging 6G applications impose heterogeneous, time-varying requirements on both wireless spectral resources and distributed computing resources. Existing works address hybrid precoding (HP) design and computing power network (CPN) task scheduling in isolation, and predominantly rely on static, offline optimization that cannot adapt to dynamic channel fluctuations and computing load variations. More critically, no prior work enforces a unified fairness criterion that spans the communication rate region and the computing latency distribution simultaneously. In this paper, we formulate the joint HP and CPN scheduling problem as a fairness-constrained Markov decision process (MDP) and propose an Online Adaptive Actor-Critic $(\text{OA}^{2} \mathrm{C})$ algorithm to solve it in real time. ${OA}^{2} \mathrm{C}$ incorporates three novel components: (i) a dual-timescale update rule that decouples fast-varying channel adaptation from slow-varying load balancing, (ii) a heterogeneity-aware state encoder that jointly embeds channel state information, task-type profiles, and hardware capability tensors, and (iii) a Lyapunov-guided fairness regularizer that provably bounds the long-run deviation from the target $\alpha$-fair rate-latency trade-off surface. Simulation results on a spherical-wave massive MIMO channel with realistic CPN hardware heterogeneity show that OA ${ }^{2} \mathrm{C}$ achieves up to $2 2. 4 \%$ improvement in Jain's fairness index, 19.7% reduction in mean task completion latency, and throughput within 3.1% of the offline DPC-based nonlinear HP upper bound, while converging ${2. 3} \times$ faster than vanilla proximal policy optimization (PPO) baselines.