← 피드로
[Submitted on 2 Feb 2026 (v1), last revised 26 Jun 2026 (this version, v3)]
Abstract:Multimodal time series forecasting is crucial in real-world applications, where decisions depend on both numerical data and contextual signals. The core challenge is to effectively combine temporal numerical patterns with the context embedded in other modalities, such as text. While most existing methods align textual features with time-series patterns one step at a time, they neglect the multiscale temporal influences of contextual information such as time-series cycles and dynamic shifts. This mismatch between local alignment and global textual context can be addressed by spectral decomposition, which separates time series into frequency components capturing both short-term changes and long-term trends. In this paper, we propose SpecTF, a simple yet effective framework that integrates the effect of textual data on time series in the frequency domain. Our method extracts textual embeddings, projects them into the frequency domain, and fuses them with the time series’ spectral components using a lightweight cross-attention mechanism. This adaptively reweights frequency bands based on textual relevance before mapping the results back to the temporal domain for predictions. Experimental results demonstrate that SpecTF significantly outperforms state-of-the-art models across diverse multi-modal time series datasets while utilizing considerably fewer parameters. Code is available at this https URL.
Submission history
From: Huu Hiep Nguyen [view email]
[v1]
Mon, 2 Feb 2026 03:28:21 UTC (3,708 KB)
[v2]
Tue, 3 Feb 2026 05:00:55 UTC (3,706 KB)
[v3]
Fri, 26 Jun 2026 12:41:08 UTC (3,706 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2602.01588
답글 남기기