Conceptual

Temporal Reconstruction and Non-Aligned Residual for Spiking Neural Network Speech Classification

Two techniques that improve spiking neural networks for speech classification. Temporal Reconstruction rebuilds the temporal dimension of an audio spectrogram at multiple resolutions, mimicking the brain's hierarchical multi-time-scale processing so the network captures information at different time scales. Non-Aligned Residual enables residual (skip) connections between audio representations of different time lengths, which single-resolution SNN pipelines otherwise cannot use. Together they set state-of-the-art accuracy on the Spiking Speech Commands and Spiking Heidelberg Digits benchmarks while reducing processing time and energy.