CyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain Provenance

Qi Chen1, Fushuo Huo1,*, Hangli Shen1, Jingcai Guo2, Shuhao Li3, Guang Cheng1
1School of Cyber Science and Engineering, Southeast University
2The Hong Kong Polytechnic University
3Zhongguancun Laboratory
* Corresponding author.
From security logs to complete attack chains. CyberClear evaluates whether agents can discover evidence, infer attack progression, and reconstruct APT campaigns without prior attack clues.

Abstract

Large language model (LLM) agents have demonstrated promising capabilities in cybersecurity tasks, yet their ability to reconstruct complete Advanced Persistent Threat (APT) attack campaigns from complex security logs remains largely unexplored. Existing cybersecurity benchmarks for agents mainly focus on vulnerability discovery, exploitation, and security analysis tasks, leaving the evaluation of attack chain provenance under realistic security logs insufficiently studied. To address this gap, we introduce CyberClear, a benchmark for evaluating LLM agents and advanced agent systems on APT attack chain provenance from long-context security logs. CyberClear covers both single-step attacks and multi-stage attack chains, requiring agents to identify attack evidence, infer attack progression, and generate provenance graphs containing entities, causal relationships, MITRE ATT&CK techniques, and forensic evidence. To enable comprehensive evaluation, we develop an evaluation method tailored to APT attack chain provenance. Unlike conventional text similarity metrics that focus on surface-level matching, our evaluation examines whether reconstructed graphs preserve the semantics of attack chains across single-step behavior correctness, multi-step behavior identification, temporal and causal consistency, entity and relationship fidelity, and overall attack narrative consistency. Advanced multi-agent systems powered by state-of-the-art LLMs still struggle on CyberClear, motivating us to propose CyberProvenance, an agent cyber harness designed for multi-agents that augments LLM agents with evidence accumulation, execution-based validation, and feedback-guided refinement mechanisms for reliable attack-chain provenance. Extensive evaluations on CyberClear demonstrate the effectiveness of CyberProvenance in improving evidence reasoning, execution-grounded validation, and complete APT attack chain reconstruction. The full benchmark and code will be released publicly upon publication.

The CyberClear Benchmark

450 benchmark instances · 89 attack behaviors · 531K characters per instance on average. CyberClear contains 318 single-step and 132 multi-step instances, organized into 318 easy, 122 medium, and 10 hard cases. Each instance pairs defender observations with an expert-verified attack provenance graph.

The construction pipeline combines isolated attack execution, defender-side evidence collection, ground-truth graph construction, and expert verification. The benchmark asks agents to recover entities, causal relations, ATT&CK techniques, and forensic evidence from the logs.

CyberProvenance

CyberProvenance connects three specialized agents through evidence accumulation, execution-based validation, and feedback-guided refinement.

Figure 4. The CyberProvenance harness grounds attack-chain reconstruction in log evidence and execution feedback.
  1. LogAnalysis Agent: searches defender logs, accumulates findings in EvidenceMemory, and drafts an attack provenance graph.
  2. AttackValidation Agent: checks predicted behaviors through controlled execution and compares resulting observations with the evidence.
  3. Check Agent: identifies inconsistencies and directs local refinement while preserving validated parts of the graph.

Experimental Results

The paper evaluates 11 agent configurations across DeepSeek-V4-Flash, Qwen3.8-Max, and GPT-5.6-Terra. Code-level metrics measure graph generation and surface similarity; semantic metrics assess behavior coverage, temporal and causal consistency, entity fidelity, and attack narrative consistency.

CyberProvenance improves attack-chain reconstruction across the evaluated settings. The results also show that valid graph code alone does not guarantee a faithful reconstruction of an attack.

Table 3. Main results on CyberClear. Within each method, the rows follow DeepSeek-V4-Flash, Qwen3.8-Max, and GPT-5.6-Terra. Blue cells mark the best result for each metric.

Performance Analysis

Evidence-Grounded Case Study

Figure 6. CyberProvenance removes unsupported causal relations by checking reconstructed edges against execution evidence.

In this example, python3 launches bash. Bash then invokes both stty and chmod, and chmod changes permissions on linpeas.sh. CyberProvenance recovers this branching structure and corrects unsupported dependencies introduced by AgentVerse.

Paper

Read the complete 40-page manuscript, including benchmark details, evaluation rubrics, ablations, and additional case studies.

Open full paper PDF ↗ · Download paper

Code & Data

The full benchmark and code will be released publicly upon publication.

Dataset on Hugging Face

Code on GitHub — Link forthcoming.

BibTeX

Citation for the manuscript. Download BibTeX

@misc{chen2026cyberclearbenchmarkllmagent,
      title={CyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain Provenance}, 
      author={Qi Chen and Fushuo Huo and Hangli Shen and Jingcai Guo and Shuhao Li and Guang Cheng},
      year={2026},
      eprint={2609.32424},
      archivePrefix={arXiv},
      primaryClass={cs.CR},
      url={https://arxiv.org/abs/2609.32424}, 
}