T3Time: Tri-Modal Time Series Forecasting via Adaptive Multi-Head Alignment and Residual Fusion
1University of Dhaka, Bangladesh
Proceedings of the AAAI Conference on Artificial Intelligence, 40(25), 2026 · Main Technical Track
TL;DR
Short-horizon forecasts need local detail; long-horizon forecasts need periodic structure. T3Time adds a frequency branch to time-and-language models, lets a horizon-aware gate decide how to mix them, and aligns everything with multiple adaptively weighted cross-modal heads.
- 7 / 8datasets with the lowest MSE
- −3.28%average MSE vs. state-of-the-art baselines (−2.29% MAE)
- −11.3%MSE on ILI, the largest single gain
- −4.13%MSE in the 5% few-shot setting

Abstract
Multivariate time series forecasting (MTSF) seeks to model temporal dynamics among variables to predict future trends. Transformer-based models and large language models (LLMs) have shown promise due to their ability to capture long-range dependencies and patterns. However, current methods often rely on rigid inductive biases, ignore inter-variable interactions, or apply static fusion strategies that limit adaptability across forecast horizons. These limitations create bottlenecks in capturing nuanced, horizon-specific relationships in time-series data. To solve this problem, we propose T3Time, a novel trimodal framework consisting of time, spectral, and prompt branches, where the dedicated frequency encoding branch captures the periodic structures along with a gating mechanism that learns prioritization between temporal and spectral features based on the prediction horizon. We also proposed a mechanism which adaptively aggregates multiple cross-modal alignment heads by dynamically weighting the importance of each head based on the features. Extensive experiments on benchmark datasets demonstrate that our model consistently outperforms state-of-the-art baselines, achieving an average reduction of 3.28% in MSE and 2.29% in MAE. Furthermore, it shows strong generalization in few-shot learning settings: with 5% training data, we see a reduction in MSE and MAE by 4.13% and 1.91%, respectively; and with 10% data, by 3.62% and 1.98% on average.
Method
Encode three views, gate by horizon, align, keep a residual path

- Tri-modal encodingA frequency branch encodes FFT magnitudes, a time branch encodes each variable's window, and a frozen GPT-2 encodes a per-variable text prompt describing the series.
- Horizon-aware gatingA small MLP looks at the temporal encoding and the prediction length and decides, per channel, how much to rely on temporal versus spectral features.
- Adaptive multi-head alignmentSeveral independent cross-modal attention heads align the series with the prompt embedding; a gating network weights the heads per variable.
- Channel-wise residual fusionLearned per-channel coefficients balance aligned features against the raw temporal–spectral path before decoding.
Results
Long-term forecasting · MSE averaged over four horizons · lower is better
| Dataset | T3Time | TimeCMA | Time-LLM | UniTime | iTransformer | PatchTST | DLinear |
|---|---|---|---|---|---|---|---|
| ETTm1 | 0.372 | 0.380 | 0.410 | 0.385 | 0.407 | 0.392 | 0.403 |
| ETTm2 | 0.279 | 0.275 | 0.296 | 0.293 | 0.288 | 0.285 | 0.350 |
| ETTh1 | 0.418 | 0.423 | 0.448 | 0.442 | 0.454 | 0.463 | 0.456 |
| ETTh2 | 0.348 | 0.372 | 0.381 | 0.378 | 0.383 | 0.395 | 0.559 |
| ECL | 0.170 | 0.174 | 0.195 | 0.216 | 0.178 | 0.207 | 0.212 |
| Weather | 0.244 | 0.250 | 0.275 | 0.253 | 0.258 | 0.257 | 0.265 |
| ILI | 1.705 | 1.922 | 2.432 | 2.108 | 2.444 | 2.388 | 2.616 |
| Exchange | 0.353 | 0.395 | 0.372 | 0.364 | 0.360 | 0.390 | 0.354 |
96-step input windows (36 for ILI), averaged over three seeds. Full MAE results and more baselines are in the paper.
Few-shot forecasting
With only 10% of the training data, T3Time beats the strongest baseline on all five evaluated datasets (−3.62% MSE on average); with 5% of the data the margin grows to −4.13% MSE.
What matters
Ablations · mean MSE increase when a component is removed
- Channel-wise residual connection: the most important part; removing it raises MSE on ILI from 1.705 to 2.176.
- Frequency branch: removing it hurts most on periodic data (e.g. Exchange 0.353 → 0.374, ILI 1.705 → 1.786).
- Multi-head alignment and horizon gating each contribute consistent smaller gains across datasets.

Citation
@inproceedings{chowdhury2026t3time,
title = {{T3Time}: Tri-Modal Time Series Forecasting via Adaptive
Multi-Head Alignment and Residual Fusion},
author = {Chowdhury, Abdul Monaf and Akter, Rabeya and Arib, Safaeid Hossain},
booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence},
volume = {40},
number = {25},
pages = {20597--20605},
year = {2026},
doi = {10.1609/aaai.v40i25.39196}
}