Skip to main navigation menu Skip to main content Skip to site footer

Time-Varying Operational Graph Reasoning and Risk-Constrained Repair Planning for Tool-Augmented Large Language Model Agents

Abstract

Addressing the operational challenges of cloud-native distributed systems, such as frequent topology reconfiguration, heterogeneous coupling of multi-source telemetry, and complex fault propagation paths, this paper studies a distributed automatic diagnosis and repair framework for operations and maintenance based on a large language model and tool-invoking intelligent agents. This framework uses a time-varying operations and maintenance graph as a unified semantic carrier, aligning and expressing observations such as service nodes, call relationships, and log indicator tracing within the same structure. It also introduces time-aware representation to enhance the characterization of non-stationary disturbances and change drift. Building upon this, the framework constructs a global representation for root cause localization through multimodal evidence fusion and context aggregation. An iterative belief update mechanism is employed to perform closed-loop absorption and correction of verifiable evidence returned by the tool, thus forming a continuous reasoning process from evidence retrieval and hypothesis generation to root cause inference. For automatic repair requirements, the framework models action selection as a planning problem that considers utility, cost, and risk constraints. Risk budgeting and rollback conditions are used to constrain the repair sequence, ensuring the controllability and auditability of the output actions. Comparative experiments on public benchmarks validated the advantages of the proposed framework in root cause localization, ranking, and remediation effectiveness, demonstrating that the method can achieve more reliable automatic diagnosis and remediation decisions in complex operational data and toolchain environments.

pdf