← 피드로
[Submitted on 4 Jun 2026 (v1), last revised 12 Jun 2026 (this version, v3)]
Abstract:Deep learning has enabled significant advances in time-series causal inference, yet progress remains constrained by the lack of realistic benchmarks with observable counterfactual outcomes. Existing datasets either rely on real-world observations without ground-truth counterfactuals or on simplified simulations that fail to capture complex causal dynamics. To address this gap, we develop a large-scale benchmark for counterfactual prediction in epidemic time series under dynamic interventions. Unlike existing benchmarks, it supports static and time-varying treatments, as well as both single-policy and multi-policy intervention settings, enabling evaluation of causal inference methods across a broad range of causal inference scenarios. Leveraging a calibrated agent-based model grounded in real-world demographic, mobility, epidemiological, and policy data, we generate realistic counterfactual trajectories across more than 150 U.S. counties. Using this benchmark, we evaluate widely used and state-of-the-art causal inference methods, revealing substantial performance differences and highlighting the challenges of realistic time-series causal reasoning.
Submission history
From: Wenhao Mu [view email]
[v1]
Thu, 4 Jun 2026 04:18:28 UTC (1,033 KB)
[v2]
Thu, 11 Jun 2026 01:52:53 UTC (1,032 KB)
[v3]
Fri, 12 Jun 2026 19:52:00 UTC (1,032 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.05692
답글 남기기