Reinforcement Learning for Long-Horizon Decision Policies in Autonomous Systems

Date

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Autonomous physical systems must make sequential decisions whose effects unfold over long temporal horizons. In domains such as space autonomy and embodied robotic control, these decisions are shaped by nonlinear dynamics, physical constraints, delayed feedback, limited interaction, fixed datasets, and imperfect decision context. A policy that is locally effective may still accumulate error, violate feasibility, or select actions unsupported by prior experience. This dissertation studies how reinforcement learning can produce long-horizon decision policies that remain feasible, controllable, support-aware, and robust across extended physical decision processes. The central thesis is that reinforcement learning becomes more suitable for autonomous physical systems when its learning architecture reflects the temporal and physical structure of the task. The dissertation develops four structured approaches. First, cascaded reinforcement learning decomposes long-horizon precision control into coarse progress and terminal refinement. Second, attention-enhanced actor–critic learning improves feature relevance in complex physical dynamics. Third, support-preserving latent planning enables offline embodied agents to compose long-horizon behavior in a learned skill space while using conservative value estimation and deterministic generative execution to reduce unsupported extrapolation. Fourth, context-robust masked skill inference trains policies to recover missing latent decisions from partial temporal context, improving robustness when decision histories are degraded. These methods are evaluated in two complementary settings: trajectory-level space autonomy and offline embodied robotic control. Across these settings, the dissertation shows that decomposition, attention, temporal abstraction, conservative value learning, generative execution, and masked inference address different aspects of the long-horizon feasibility gap. The broader contribution is a unified design perspective for reinforcement learning in autonomous physical systems: reliable long-horizon decision policies require learning structures aligned with horizon length, physical feasibility, data support, controllability, and context reliability.

Description

Keywords

Reinfrocement Learning, Robotics, Long Horizon Learning, Offline Reinforcment Learning, Autonomous Planning

Graduation Month

August

Degree

Doctor of Philosophy

Department

Department of Computer Science

Major Professor

Arslan Munir

Date

Type

Dissertation

Citation