Scalable Graph Learning and Causal Inference for Resilience Analysis in Complex Infrastructure Networks
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Critical infrastructure systems such as power, water, stormwater, and transportation form tightly interconnected networks vital for essential services and public safety. Failure in one often cascades across others, impacting community lifelines. Traditional approaches analyze each system in isolation, overlooking interdependencies, network evolution, and the causal factors that drive failures. This research develops a scalable framework addressing four challenges: (1) representing complex infrastructure and identifying critical components; (2) capturing continuous network evolution; (3) understanding causal drivers of failure propagation; and (4) building explainable models that reveal why predictions are made and what interventions will restore stability. Our core questions are: which functionalities most influence community-level resilience, and how feasible interventions can minimize cascading failures and inequities? We propose an integrated framework combining Graph Neural Networks (GNNs) and Hetero-Functional Graphs (HFGs). HFGs model functional interdependencies across power, water, and transportation systems from operational data. We develop a Generalized Power GNN (GP-GNN) for critical node identification and a Cascading Outage GNN (CO-GNN) for critical link classification, both scalable across large networks. A dynamic GNN–BiLSTM model tracks temporal network evolution, while Wasserstein-based graph embeddings (WEGL) detect structural shifts — triggering reanalysis only when changes genuinely alter system behavior. Causal inference methods uncover root drivers of failure propagation, and a counterfactual explainability layer translates predictions into minimal-change operator recommendations. GP-GNN and CO-GNN achieve 97–99% accuracy for critical nodes and 98–99% for links, at twice the speed of simulation. WEGL-based shift detection reduces recomputation overhead by 65–95%. Causal and explainability methods identify root failure drivers and delivers specific, actionable guidance to operators. Together, these capabilities shift infrastructure management from reactive to anticipatory. By equipping planners and emergency managers with smarter, proactive decision-making tools, this research enables them to anticipate cascading failures, understand root causes, and implement targeted interventions that build more equitable and resilient communities.