NeurIPS 2026 Spotlight: Parallel-in-Time RNN Training Achieves >100x Speedup for Chaotic Dynamical Systems
Key point
Combining DEER with generalized teacher forcing enables stable, parallel training on time series exceeding 1 million steps, significantly outperforming Mamba.
Details
A NeurIPS 2026 spotlight paper introduces a method to train nonlinear Recurrent Neural Networks (RNNs) for Dynamical Systems Reconstruction (DSR) on chaotic systems with >100x speedup compared to standard approaches. The technique combines DEER (a Newton-type fixed-point solver) with generalized teacher forcing (GTF) to enable efficient GPU parallelization.
Methodology and Performance
Standard sequential training scales linearly with sequence length $O[T]$, creating a bottleneck for long time series. DEER reduces this complexity to $O[(\log T)^2]$ by solving the forward pass across the entire sequence simultaneously. However, DEER typically fails under chaotic dynamics, degrading to $O[T \log T]$ due to divergence. The proposed integration of GTF stabilizes DEER, preventing divergence and reducing exposure bias.
Key Results
- Speedup: Achieves more than 2 orders of magnitude (>100x) faster training than baseline methods.
- Scalability: Successfully trains on extremely long time series with $T > 10^6$ steps.
- Comparison: Outperforms Mamba and other state space models in the DSR setting for chaotic simulated and real-world systems.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.