dc.contributor.authorZaidi, Syed Muhammad Talha
dc.date.accessioned2026-08-11T21:45:36Z
dc.date.available2026-08-11T21:45:36Z
dc.date.graduationmonthAugust
dc.date.issued2026
dc.description.abstractAutonomous physical systems must make sequential decisions whose effects unfold over long temporal horizons. In domains such as space autonomy and embodied robotic control, these decisions are shaped by nonlinear dynamics, physical constraints, delayed feedback, limited interaction, fixed datasets, and imperfect decision context. A policy that is locally effective may still accumulate error, violate feasibility, or select actions unsupported by prior experience. This dissertation studies how reinforcement learning can produce long-horizon decision policies that remain feasible, controllable, support-aware, and robust across extended physical decision processes. The central thesis is that reinforcement learning becomes more suitable for autonomous physical systems when its learning architecture reflects the temporal and physical structure of the task. The dissertation develops four structured approaches. First, cascaded reinforcement learning decomposes long-horizon precision control into coarse progress and terminal refinement. Second, attention-enhanced actor–critic learning improves feature relevance in complex physical dynamics. Third, support-preserving latent planning enables offline embodied agents to compose long-horizon behavior in a learned skill space while using conservative value estimation and deterministic generative execution to reduce unsupported extrapolation. Fourth, context-robust masked skill inference trains policies to recover missing latent decisions from partial temporal context, improving robustness when decision histories are degraded. These methods are evaluated in two complementary settings: trajectory-level space autonomy and offline embodied robotic control. Across these settings, the dissertation shows that decomposition, attention, temporal abstraction, conservative value learning, generative execution, and masked inference address different aspects of the long-horizon feasibility gap. The broader contribution is a unified design perspective for reinforcement learning in autonomous physical systems: reliable long-horizon decision policies require learning structures aligned with horizon length, physical feasibility, data support, controllability, and context reliability.
dc.description.advisorArslan Munir
dc.description.degreeDoctor of Philosophy
dc.description.departmentDepartment of Computer Science
dc.description.levelDoctoral
dc.identifier.urihttps://hdl.handle.net/2097/47375
dc.language.isoen_US
dc.subjectReinfrocement Learning
dc.subjectRobotics
dc.subjectLong Horizon Learning
dc.subjectOffline Reinforcment Learning
dc.subjectAutonomous Planning
dc.titleReinforcement Learning for Long-Horizon Decision Policies in Autonomous Systems
dc.typeDissertation

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
SyedmuhammadtalhaZaidi2026.pdf
Size:
16 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.65 KB
Format:
Item-specific license agreed upon to submission
Description: