Multi-Modal Time Series Prediction via Mixture of Modulated Experts

작성자

카테고리:

← 피드로
arXiv cs.AI · Lige Zhang, Ali Maatouk, Jialin Chen, Karthik Charan Konduri, Leandros Tassiulas, Rex Ying · 2026-09-14 AI

[Submitted on 29 Jan 2026 (v1), last revised 11 Sep 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Real-world time series exhibit complex and evolving dynamics, making accurate forecasting extremely challenging. Recent multi-modal forecasting methods leverage textual information such as news reports to improve prediction, but most rely on token-level fusion that mixes temporal patches with language tokens in a shared embedding space. However, such fusion can be ill-suited when high-quality time-text pairs are scarce and when time series exhibit substantial variation in characteristics, thus complicating cross-modal alignment. In parallel, mixture-of-experts (MoE) architectures have proven effective for both time series modeling and multi-modal learning, yet many existing MoE-based modality integration methods still depend on token-level fusion. To address this, we propose Expert Modulation, a new mechanism for multi-modal time series prediction that conditions both routing and expert computation on textual signals, enabling direct and efficient cross-modal control over expert behavior. Through theoretical analysis and experiments, our proposed method demonstrates strong improvements in multi-modal time series prediction. The current code implementation is available at this https URL

Submission history

From: Lige Zhang [view email]
[v1] Thu, 29 Jan 2026 11:03:09 UTC (26,683 KB)
[v2] Fri, 4 Sep 2026 04:08:10 UTC (12,319 KB)
[v3] Fri, 11 Sep 2026 13:04:59 UTC (12,319 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2601.21547