Reinforcement Learning for Long-Horizon Decision Policies in Autonomous Systems
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Autonomous physical systems must make sequential decisions whose effects unfold over long temporal horizons. In domains such as space autonomy and embodied robotic control, these decisions are shaped by nonlinear dynamics, physical constraints, delayed feedback, limited interaction, fixed datasets, and imperfect decision context. A policy that is locally effective may still accumulate error, violate feasibility, or select actions unsupported by prior experience. This dissertation studies how reinforcement learning can produce long-horizon decision policies that remain feasible, controllable, support-aware, and robust across extended physical decision processes. The central thesis is that reinforcement learning becomes more suitable for autonomous physical systems when its learning architecture reflects the temporal and physical structure of the task. The dissertation develops four structured approaches. First, cascaded reinforcement learning decomposes long-horizon precision control into coarse progress and terminal refinement. Second, attention-enhanced actor–critic learning improves feature relevance in complex physical dynamics. Third, support-preserving latent planning enables offline embodied agents to compose long-horizon behavior in a learned skill space while using conservative value estimation and deterministic generative execution to reduce unsupported extrapolation. Fourth, context-robust masked skill inference trains policies to recover missing latent decisions from partial temporal context, improving robustness when decision histories are degraded. These methods are evaluated in two complementary settings: trajectory-level space autonomy and offline embodied robotic control. Across these settings, the dissertation shows that decomposition, attention, temporal abstraction, conservative value learning, generative execution, and masked inference address different aspects of the long-horizon feasibility gap. The broader contribution is a unified design perspective for reinforcement learning in autonomous physical systems: reliable long-horizon decision policies require learning structures aligned with horizon length, physical feasibility, data support, controllability, and context reliability.