Evaluating Generative Time-Series Models on Data with Point Masses

작성자

카테고리:

← 피드로
arXiv cs.AI · Jian Xu · 2026-08-11 AI

[Submitted on 10 Aug 2026]

View PDF HTML (experimental)

Abstract:Many of the series that generative time-series models are benchmarked on place a large probability mass on a single value — it does not rain, no ride is requested, no part is ordered. We report what happens when such data is evaluated carefully. First, the standard rolling-origin protocol can score a model on a window whose atom structure bears no resemblance to the dataset: on one benchmark the dataset is $42\%$ zeros and the evaluation windows are $13\%$, on another $47\%$ against $5\%$. This is not a cosmetic problem — it reversed one of our own conclusions, turning the strongest occurrence model in our study into what looked like a cautionary tale. Second, we give a control in which CRPS is invariant \emph{by construction} while the temporal coupling is destroyed, which measures exactly how much that coupling contributes to a chosen statistic. Third, benchmarking seven models on a matched protocol over five seeds, an autoregressive hurdle beats a conditional flow on five of six datasets, by up to a factor of $153$, while the flow’s own occurrence statistics vary by up to $62\%$ across training seeds and every baseline is deterministic. Finally, the model ordering is not the same under five different occurrence statistics, and the two that do not share a construction agree with each other least.

Submission history

From: Xu Jian [view email]
[v1] Mon, 10 Aug 2026 14:57:36 UTC (96 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.09692

코멘트

답글 남기기