Skip to main navigation menu Skip to main content Skip to site footer

Behavior Representation Learning for Reliable Monitoring of AI Inference Backends

Abstract

In multi-tenant and multi-model concurrent AI inference backends, system behavior exhibits high dynamics and complex structure. This property poses significant challenges to monitoring mechanisms based on rules or single metrics. From the perspective of system behavior modeling, this paper proposes a learning driven backend monitoring approach. Multi-source observations generated during inference execution are mapped into a unified latent representation space. This mapping captures the global evolution of system states. By jointly modeling temporal consistency and global consistency of behavior representations, the method forms stable and discriminative behavior structures in the latent space. These structures provide a reliable basis for runtime state assessment. On this foundation, a unified monitoring scoring mechanism is constructed. It quantifies the deviation between the current system state and the overall behavior pattern. This design enables effective identification of anomalous behavior. Comparative analysis with multiple monitoring methods shows that system behavior learning based monitoring achieves clear advantages in discrimination stability and false alarm control. The method adapts well to complex operating conditions in inference backends. This work presents a system-level modeling perspective for reliable AI inference backend management. It supports the practical adoption of learning based monitoring methods in real-world backend systems.

pdf