Workshop AIWILD @ ICLR

AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems

Zhaohui Wang

Workshop on Agents in the Wild: Safety, Security, and Beyond at the International Conference on Learning Representations (ICLR), 2026

type Workshop
year 2026
venue AIWILD @ ICLR
arXiv 2603.14688

// FIGURE_1

// ABSTRACT

As multi-agent AI systems are increasingly deployed in real-world settings—from automated customer support to DevOps remediation—failures become harder to diagnose due to cascading effects, hidden dependencies, and long execution traces. We present AgentTrace, a lightweight causal tracing framework for post-hoc failure diagnosis in deployed multi-agent workflows. AgentTrace reconstructs causal graphs from execution logs, traces backward from error manifestations, and ranks candidate root causes using interpretable structural and positional signals—without requiring LLM inference at debugging time. Across a diverse benchmark of multi-agent failure scenarios designed to reflect common deployment patterns, AgentTrace localizes root causes with high accuracy and sub-second latency, significantly outperforming both heuristic and LLM-based baselines. Our results suggest that causal tracing provides a practical foundation for improving the reliability and trustworthiness of agentic systems in the wild.

// BIBTEX

@inproceedings{wang2026agenttrace,
  title     = {AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems},
  author    = {Zhaohui Wang},
  booktitle = {Workshop on Agents in the Wild: Safety, Security, and Beyond at the International Conference on Learning Representations (ICLR)},
  year      = {2026},
  month     = {4},
  eprint    = {2603.14688},
  archivePrefix = {arXiv},
}