AAAI 2026 Time SeriesMultimodal LearningLLMs

T3Time: Tri-Modal Time Series Forecasting via Adaptive Multi-Head Alignment and Residual Fusion

Abdul Monaf Chowdhury1, Rabeya Akter1, Safaeid Hossain Arib1

1University of Dhaka, Bangladesh

Proceedings of the AAAI Conference on Artificial Intelligence, 40(25), 2026 · Main Technical Track

TL;DR

Short-horizon forecasts need local detail; long-horizon forecasts need periodic structure. T3Time adds a frequency branch to time-and-language models, lets a horizon-aware gate decide how to mix them, and aligns everything with multiple adaptively weighted cross-modal heads.

  • 7 / 8datasets with the lowest MSE
  • −3.28%average MSE vs. state-of-the-art baselines (−2.29% MAE)
  • −11.3%MSE on ILI, the largest single gain
  • −4.13%MSE with only 5% of the training data
Overview figure for T3Time: Tri-Modal Time Series Forecasting via Adaptive Multi-Head Alignment and Residual Fusion
Bimodal vs. trimodal. (a) Prior work fuses a time-series encoder with an LLM statically. (b) T3Time adds a frequency encoder, horizon-aware gating, and cross-modal alignment with adaptive head fusion.

Abstract

Multivariate time series forecasting (MTSF) seeks to model temporal dynamics among variables to predict future trends. Transformer-based models and large language models (LLMs) have shown promise due to their ability to capture long-range dependencies and patterns. However, current methods often rely on rigid inductive biases, ignore inter-variable interactions, or apply static fusion strategies that limit adaptability across forecast horizons. These limitations create bottlenecks in capturing nuanced, horizon-specific relationships in time-series data. To solve this problem, we propose T3Time, a novel trimodal framework consisting of time, spectral, and prompt branches, where the dedicated frequency encoding branch captures the periodic structures along with a gating mechanism that learns prioritization between temporal and spectral features based on the prediction horizon. We also proposed a mechanism which adaptively aggregates multiple cross-modal alignment heads by dynamically weighting the importance of each head based on the features. Extensive experiments on benchmark datasets demonstrate that our model consistently outperforms state-of-the-art baselines, achieving an average reduction of 3.28% in MSE and 2.29% in MAE. Furthermore, it shows strong generalization in few-shot learning settings: with 5% training data, we see a reduction in MSE and MAE by 4.13% and 1.91%, respectively; and with 10% data, by 3.62% and 1.98% on average.

Three ways to read one series

Motivation

Forecasting models usually commit to one view of the data. Time-domain Transformers such as PatchTST and iTransformer model local dynamics. Frequency models such as FEDformer capture periodicity. LLM-based models such as Time-LLM and TimeCMA describe the series in a text prompt and borrow a language model's priors. Each view sees something the others miss.

  • Isolated modalities. Prompt-based models pair time with language but ignore the spectral view that summarizes global periodicity.
  • Narrow alignment. A single cross-attention head limits how modalities interact.
  • Horizon rigidity. The same fusion is used for every forecast length, although short horizons lean on local detail and long horizons on periodic structure.

Method

Encode three views, gate by horizon, align, keep a residual path

T3Time architecture with frequency, time-series, and LLM encoding branches, horizon-aware gating, multi-head cross-modal attention, adaptive head fusion, and a channel-wise residual connection
T3Time architecture. Frequency, time-series and LLM-prompt branches feed horizon-aware gating, multi-head cross-modal attention with adaptive head fusion, and a channel-wise residual connection before a Transformer decoder.

1. Tri-modal encoding

Frequency. The input is transformed with a real FFT and each magnitude bin becomes a token. After a one-layer Transformer, an attention-weighted pooling summarizes the spectrum per variable.

\[ \widehat{\mathbf{X}} = \mathcal{F}_r(\mathbf{X}) \in \mathbb{C}^{B \times N \times L_f}, \quad L_f = \lfloor L/2 \rfloor + 1, \qquad \mathbf{F}_{\mathrm{pooled}} = \sum_{l=1}^{L_f} \alpha_l\, \widetilde{\mathbf{Z}}_{f,l} \]

Time. Each variable's window is linearly embedded and encoded by a Transformer.

Language. Each variable gets a prompt such as "From [t1] to [t2], the values were [v1, …, vn] every [hour]. The total trend value was [T]." A frozen GPT-2 encodes it, and the last-token embedding is kept.

2. Horizon-aware gating

A small MLP reads the pooled time encoding together with the normalized prediction length and outputs per-channel gates. Each forecast horizon gets its own mix of spectral and temporal evidence.

\[ \mathbf{g} = \sigma\big(\mathbf{W}_4\, \phi(\mathbf{W}_3\, \mathbf{g}_{\mathrm{in}})\big), \qquad \mathbf{Z}_g = \mathbf{g} \odot \widetilde{\mathbf{F}} + (1 - \mathbf{g}) \odot \widetilde{\mathbf{Z}}_t \]

3. Adaptive multi-head cross-modal alignment

Instead of one cross-attention head, \(H\) independent heads align the gated series features (queries) with the prompt embeddings (keys and values). A gating network then scores the heads separately for every variable, so different variables can rely on different alignments.

\[ \pi_{b,n}^{(h)} = \mathrm{softmax}_h\big(\mathbf{W}_6\, \phi(\mathrm{LN}(\mathbf{W}_5\, \mathbf{U}_{b,n}))\big), \qquad \mathbf{\Lambda}_{b,:,n} = \sum_{h=1}^{H} \pi_{b,n}^{(h)}\, \mathbf{H}^{(h)}_{b,:,n} \]

4. Channel-wise residual fusion

Learned per-channel coefficients decide how much each latent channel takes from the aligned features versus the raw temporal–spectral path, before a Transformer decoder and a linear forecasting head.

\[ \mathbf{\Theta}_{b,c,n} = \gamma_c\, \mathbf{\Lambda}_{b,c,n} + (1 - \gamma_c)\, \mathbf{Z}_{g,b,c,n}, \qquad \gamma_c \in [0, 1] \]

Results

Long-term forecasting · averaged over four horizons · lower is better

Following TimeCMA, inputs are 96 steps and horizons are {96, 192, 336, 720} (36 steps and {24, 36, 48, 60} for ILI), averaged over three seeds.

MSE

DatasetT3TimeTimeCMATime-LLMUniTimeTimesNetDLineariTransformerPatchTSTOFA
ETTm10.3720.3800.4100.3850.4000.4030.4070.3920.396
ETTm20.2790.2750.2960.2930.2910.3500.2880.2850.294
ETTh10.4180.4230.4480.4420.4580.4560.4540.4630.457
ETTh20.3480.3720.3810.3780.4140.5590.3830.3950.389
ECL0.1700.1740.1950.2160.1920.2120.1780.2070.217
Weather0.2440.2500.2750.2530.2590.2650.2580.2570.279
ILI1.7051.9222.4322.1082.1392.6162.4442.3882.623
Exchange0.3530.3950.3720.3640.4160.3540.3600.3900.519

MAE

DatasetT3TimeTimeCMATime-LLMUniTimeTimesNetDLineariTransformerPatchTSTOFA
ETTm10.3930.3920.4090.3990.4060.4070.4100.4020.401
ETTm20.3220.3230.3400.3340.3330.4010.3320.3280.339
ETTh10.4300.4310.4430.4480.4500.4520.4470.4490.450
ETTh20.3900.3970.4040.4030.4270.5150.4070.4140.414
ECL0.2660.2690.2880.3060.2950.3000.2700.2890.308
Weather0.2750.2760.2910.2760.2870.3170.2780.2800.297
ILI0.8350.9211.0120.9290.9311.0901.2031.0111.060
Exchange0.4010.4290.4160.4040.4430.4140.4030.4290.500

Few-shot forecasting

With only a tenth or a twentieth of the training data, T3Time beats the strongest baseline on all five datasets.

DatasetT3Time (10%)Best baseline (10%)T3Time (5%)Best baseline (5%)
ETTm10.3760.387 (TimeCMA)0.3840.396 (TimeCMA)
ETTm20.2660.277 (Time-LLM)0.2670.274 (Time-LLM)
ETTh10.4490.480 (TimeCMA)0.4420.472 (TimeCMA)
ETTh20.3570.370 (Time-LLM)0.3570.382 (Time-LLM)
Weather0.2260.229 (TimeCMA)0.2260.231 (TimeCMA)

MSE, lower is better. Average reductions are 3.62% (10%) and 4.13% (5%).

What matters

Ablations · MSE

DatasetT3Timew/o frequencyw/o multi-head CMAw/o residualw/o gating
ETTm10.3720.3810.3740.4040.373
ETTm20.2790.2790.2830.2880.280
ETTh10.4180.4330.4210.4330.425
ETTh20.3480.3640.3550.3840.363
Weather0.2440.2500.2490.2490.249
ILI1.7051.7861.8132.1761.724
Exchange0.3530.3740.3780.3960.373
  • The residual path matters most. Without it, MSE rises on every dataset, from 1.705 to 2.176 on ILI.
  • The spectral view pays off on periodic data, such as Exchange (0.353 → 0.374) and ILI (1.705 → 1.786).
  • Multi-head alignment and horizon gating each give consistent smaller gains.

What the three views learn

t-SNE of time-series, frequency, prompt, and forecasted embeddings coloured by dataset
t-SNE of learned embeddings across six datasets. Time-series and frequency embeddings form clear dataset clusters, reflecting distinct temporal and periodic structure. Prompt embeddings are more dispersed, reflecting the diversity of the prompts.

Limitations

Experiments use standard public benchmarks and a frozen GPT-2 for the language branch. Larger-scale pretraining and richer representations for each modality are left for future work.

Citation

@inproceedings{chowdhury2026t3time,
  title     = {{T3Time}: Tri-Modal Time Series Forecasting via Adaptive
               Multi-Head Alignment and Residual Fusion},
  author    = {Chowdhury, Abdul Monaf and Akter, Rabeya and Arib, Safaeid Hossain},
  booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence},
  volume    = {40},
  number    = {25},
  pages     = {20597--20605},
  year      = {2026},
  doi       = {10.1609/aaai.v40i25.39196}
}