SUNTA: Hierarchical Video Prediction with Surprise-based Chunking

작성자

카테고리:

← 피드로
arXiv cs.AI · Tomoshi Iiyama, Masahiro Suzuki, Yutaka Matsuo · 2026-07-14 AI

[Submitted on 2 Jul 2026 (v1), last revised 10 Jul 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Hierarchical state-space models (HSSMs) offer a promising approach to long-horizon prediction by segmenting sequences into temporal chunks. However, their performance hinges on how chunk boundaries are determined. While prior HSSMs typically rely on fixed-length chunking or similarity-based boundary detection, these methods often misalign with the intrinsic temporal structure of the data. We argue that chunking should instead be driven by prediction errors, which more directly indicate when longer-range context becomes necessary. Nevertheless, integrating surprise-based chunking into HSSMs introduces critical challenges, including hierarchical collapse during end-to-end training and the absence of surprise signals during open-loop prediction. To address these issues, we propose Surprise-based Nested Temporal Abstraction (SUNTA), a method that employs a decoupled training strategy to preserve surprise signals and uses internal inconsistency as a top-down surprise metric to determine chunk boundaries within imagined rollouts. Experiments on video prediction tasks in 2D and 3D environments demonstrate that SUNTA outperforms baselines, uniquely maintaining accurate predictions over 250 timesteps, whereas all baselines degrade within the first 10 timesteps.

Submission history

From: Tomoshi Iiyama [view email]
[v1] Thu, 2 Jul 2026 12:27:54 UTC (8,629 KB)
[v2] Fri, 10 Jul 2026 20:58:05 UTC (8,608 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.02087

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다