DepthSSD: Rethinking Residual Connections via State Space Models on the Depth Axis
Abstract
Residual connections are fundamental to deep transformer training, yet their theoretical relationship to more expressive alternatives, such as attention-based depth mixing (AttnRes), remains poorly understood. We show that the depth-axis mixing matrix in transformers is semiseparable, and its rank precisely controls the expressiveness of inter-layer information flow. Standard residual connections correspond to rank-1 semiseparable matrices; attention residuals to full rank. We introduce DEPTHSSD (Depth State Space Duality), which applies the structured state space duality framework from Mamba-2 to the depth axis, yielding a depth mixing matrix of controllable rank N with O(L · N) parameters per sub-layer. We evaluate DEPTHSSD at three model scales (125M, 455M, and 1B parameters) against four depth-mixing baselines: standard residuals, BLOCKATTNRES, DENSEFORMER, and Hyper-Connections. At 1B, DEPTHSSD-N=8 achieves the best validation loss (3.124, perplexity 22.7), a 2.4-point perplexity reduction over standard residuals (PPL 25.1) and outperforming DENSEFORMER (PPL 24.1), while running 1.72× faster (16.5 vs. 9.6 TFLOPS). At 455M, DEPTHSSD-N=8 again leads (PPL 29.4 vs. 30.1 for DENSEFORMER and 31.2 for standard) at 1.88× the throughput of DENSEFORMER. At 125M, DEPTHSSD matches the validation quality of all baselines while maintaining 1.3–1.6× higher throughput than the other depth-mixing methods.