A multi-dataset benchmark for evaluating LLM agents in microservice failure diagnosis. This repository bundles the two datasets introduced in *"A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis"*: **503 expert-labeled failure cases**, **~15.3 GB** of multimodal observability data, validated through **6,093 teams** across two 2025 national-level competitions.
Both datasets evaluate agents along three shared pillars — **Localization** (where the fault is), **Identification** (what fault type it is), and **Reason** (whether the reasoning trace is grounded in the right diagnostic evidence) — using two complementary labelling forms: per-modality **key-evidence** (AIOps2025) and typed **causal chains** (RCA100).
@@ -9,6 +11,7 @@ Both datasets evaluate agents along three shared pillars — **Localization** (w