Skip to content
Preprint

Geometric Distributional Control: Learning Progress with Partial Structural Knowledge

Sep 2026 · 0 citations · 38 references
Computer Science Engineering

TL;DR

Geometric Distributional Control is validated on structured multilevel optimization and SUMO route-progress driving, where it improves over known-only solvers and learning baselines while preserving scaffold-enforced feasibility.

Abstract

Real-time control often sits between two limiting regimes. Predictive optimization and model-based control are powerful when dynamics, parameters, objectives, and online planning models are specified; reinforcement learning can relax this requirement, but must infer long-horizon value signals from sequential data and interaction, making training slow, high-variance, and hard to scale in large action spaces. This middle regime is common in systems including autonomous driving, warehouse robotics, traffic control, and delivery drones: partial geometry, physics, rules, or constraints are known, yet the local direction of task progress remains uncertain. Geometric Distributional Control (GDC) is designed for this partial-knowledge setting. It factorizes control into feasibility and progress: known geometry, rules, constraints, and response maps define an executable scaffold, while progress-weighted feasible data learns the missing directional signal on that scaffold. The learned score acts as a Bellman-like local value-gradient, selecting actions that make progress without requiring global Bellman recursion, a fully specified planner, or a black-box policy that absorbs both feasibility and preference. This knowledge can be lightweight and partial, such as simple dynamics, safety filters, local maps, constraint projectors, or lower-level response maps; it need not encode full dynamics or a long-horizon objective. Offline, GDC fits a progress-tilted distribution from short known-feasible snippets with weak signed progress certificates. Online, its score is projected through the scaffold and applied in receding-horizon feedback. We prove that this score descends a data-induced soft progress value and validate GDC on structured multilevel optimization and SUMO route-progress driving, where it improves over known-only solvers and learning baselines while preserving scaffold-enforced feasibility.

View source

Similar papers

#artificial intelligence Preprint Oct 2026

Optimal Control with Learned Critics under Unmodeled State Dependencies

Model Predictive Control (MPC) provides a structured and constraint-aware mechanism for decision-making, but its reliance on optimization-friendly analytical dynamics models limits its use in tasks with contacts and other hard-to-model state dependencies. Model-free reinforcement learning avoids explicit modeling assum...

P. Schöch, Markus Ryll · 0 citations
#reinforcement learning Preprint Aug 2026

Guided Riemannian Optimization (GuRO): Bridging Model Predictive Control and Decision Transformers

A novel framework is proposed that integrates MPC with RL in a sequence decision-making framework and leverages a curvature-aware optimization to efficiently tackle non-convex loss landscapes and achieves higher returns and faster convergence.

Hossein Abdi, S. Dash, Ming-Fei Sun · 0 citations
Preprint Sep 2026

Learning Robot Policies from Sparse Success Signals via STL-Guided Stein Variational Policy Gradient

Learning robot policies for tasks with sparse success signals is challenging when completion depends on coordinated actions, precise contact outcomes, or satisfying several conditions together. Intricate physical interactions with the world further complicate these requirements. Prior work using conventional reward sha...

Hong-Rui Zheng, C. Vasile, Antonio Loquercio et al. · 0 citations
Preprint Sep 2026

Tractable Reinforcement Learning for Full Class of Signal Temporal Logic Specifications Using Spatiotemporal Tube Reward

This paper addresses the control problem for robotic systems, including non-holonomic and underactuated platforms operating under unknown dynamics and strict actuator limits to satisfy complex high-level specifications. We denote these high-level specifications using Signal Temporal Logic (STL) and propose a novel time...

Vaishnavi Jagabathula, P. Sangeerth, Pushpak Jagtap · 0 citations
Preprint Oct 2026

Benchmarking Generative Trajectory Models for Active-Inference Control

Learning from trajectory demonstrations offers a route to active-inference control of complex systems whose dynamics are difficult to model explicitly. We introduce generative active-inference control (GenAIF), in which one generative trajectory model learns from demonstrations and measured action interventions to supp...

Yu-Lin Li, Mohsen A. Jafari, Andrea Matta · 0 citations
#machine learning Preprint Sep 2026

Runtime-Incremental Transformer for Reinforcement-Learning-Based Adaptive Control

A runtime mechanism that grows and prunes the heads of the attention block during reinforcement learning, governed by two signals: the effective rank of the on-policy context distribution, which triggers growth when representational capacity becomes insufficient, and the per-head output magnitude, which flags redundant...

G. Cirrincione, Adriano Fagiolini · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.