AAAI 2026 Time SeriesMultimodal LearningLLMs

T3Time: Tri-Modal Time Series Forecasting via Adaptive Multi-Head Alignment and Residual Fusion

Abdul Monaf Chowdhury1, Rabeya Akter1, Safaeid Hossain Arib1

1University of Dhaka, Bangladesh

Proceedings of the AAAI Conference on Artificial Intelligence, 40(25), 2026 · Main Technical Track

TL;DR

Short-horizon forecasts need local detail; long-horizon forecasts need periodic structure. T3Time adds a frequency branch to time-and-language models, lets a horizon-aware gate decide how to mix them, and aligns everything with multiple adaptively weighted cross-modal heads.

  • 7 / 8datasets with the lowest MSE
  • −3.28%average MSE vs. state-of-the-art baselines (−2.29% MAE)
  • −11.3%MSE on ILI, the largest single gain
  • −4.13%MSE in the 5% few-shot setting
Overview figure for T3Time: Tri-Modal Time Series Forecasting via Adaptive Multi-Head Alignment and Residual Fusion
Bimodal vs. trimodal. (a) Prior work fuses a time-series encoder with an LLM statically. (b) T3Time adds a frequency encoder, horizon-aware gating, and cross-modal alignment with adaptive head fusion.

Abstract

Multivariate time series forecasting (MTSF) seeks to model temporal dynamics among variables to predict future trends. Transformer-based models and large language models (LLMs) have shown promise due to their ability to capture long-range dependencies and patterns. However, current methods often rely on rigid inductive biases, ignore inter-variable interactions, or apply static fusion strategies that limit adaptability across forecast horizons. These limitations create bottlenecks in capturing nuanced, horizon-specific relationships in time-series data. To solve this problem, we propose T3Time, a novel trimodal framework consisting of time, spectral, and prompt branches, where the dedicated frequency encoding branch captures the periodic structures along with a gating mechanism that learns prioritization between temporal and spectral features based on the prediction horizon. We also proposed a mechanism which adaptively aggregates multiple cross-modal alignment heads by dynamically weighting the importance of each head based on the features. Extensive experiments on benchmark datasets demonstrate that our model consistently outperforms state-of-the-art baselines, achieving an average reduction of 3.28% in MSE and 2.29% in MAE. Furthermore, it shows strong generalization in few-shot learning settings: with 5% training data, we see a reduction in MSE and MAE by 4.13% and 1.91%, respectively; and with 10% data, by 3.62% and 1.98% on average.

Method

Encode three views, gate by horizon, align, keep a residual path

T3Time architecture with frequency, time-series, and LLM encoding branches, horizon-aware gating, multi-head cross-modal attention, adaptive head fusion, and a channel-wise residual connection
T3Time architecture. Frequency, time-series, and LLM-prompt branches feed horizon-aware gating, multi-head cross-modal attention with adaptive head fusion, and a channel-wise residual connection before a Transformer decoder.
  1. Tri-modal encodingA frequency branch encodes FFT magnitudes, a time branch encodes each variable's window, and a frozen GPT-2 encodes a per-variable text prompt describing the series.
  2. Horizon-aware gatingA small MLP looks at the temporal encoding and the prediction length and decides, per channel, how much to rely on temporal versus spectral features.
  3. Adaptive multi-head alignmentSeveral independent cross-modal attention heads align the series with the prompt embedding; a gating network weights the heads per variable.
  4. Channel-wise residual fusionLearned per-channel coefficients balance aligned features against the raw temporal–spectral path before decoding.

Results

Long-term forecasting · MSE averaged over four horizons · lower is better

DatasetT3TimeTimeCMATime-LLMUniTimeiTransformerPatchTSTDLinear
ETTm10.3720.3800.4100.3850.4070.3920.403
ETTm20.2790.2750.2960.2930.2880.2850.350
ETTh10.4180.4230.4480.4420.4540.4630.456
ETTh20.3480.3720.3810.3780.3830.3950.559
ECL0.1700.1740.1950.2160.1780.2070.212
Weather0.2440.2500.2750.2530.2580.2570.265
ILI1.7051.9222.4322.1082.4442.3882.616
Exchange0.3530.3950.3720.3640.3600.3900.354

96-step input windows (36 for ILI), averaged over three seeds. Full MAE results and more baselines are in the paper.

Few-shot forecasting

With only 10% of the training data, T3Time beats the strongest baseline on all five evaluated datasets (−3.62% MSE on average); with 5% of the data the margin grows to −4.13% MSE.

What matters

Ablations · mean MSE increase when a component is removed

  • Channel-wise residual connection: the most important part; removing it raises MSE on ILI from 1.705 to 2.176.
  • Frequency branch: removing it hurts most on periodic data (e.g. Exchange 0.353 → 0.374, ILI 1.705 → 1.786).
  • Multi-head alignment and horizon gating each contribute consistent smaller gains across datasets.
t-SNE of time-series, frequency, prompt, and forecasted embeddings coloured by dataset
What the three views learn. t-SNE of learned embeddings across six datasets. Time-series and frequency embeddings form clear dataset clusters, reflecting distinct temporal and periodic structure; prompt embeddings are more dispersed, reflecting the diversity of the prompts.

Citation

@inproceedings{chowdhury2026t3time,
  title     = {{T3Time}: Tri-Modal Time Series Forecasting via Adaptive
               Multi-Head Alignment and Residual Fusion},
  author    = {Chowdhury, Abdul Monaf and Akter, Rabeya and Arib, Safaeid Hossain},
  booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence},
  volume    = {40},
  number    = {25},
  pages     = {20597--20605},
  year      = {2026},
  doi       = {10.1609/aaai.v40i25.39196}
}