Attention Kernels for Learning Maps Between Heavy-Tailed Measures
Abstract
Transformers can learn operators on probability measures, but softmax attention can become ill-defined when exponential weights are integrated against heavy-tailed distributions. We investigate alternatives with slower-growing attention kernels using two benchmarks with explicit target maps. On both heavy-tailed tasks, standard softmax exhibits ensemble collapse without preprocessing, while the alternative kernels avoid it. Symlog preprocessing helps softmax on one task but does not resolve the other. A Gaussian control gives comparable performance across kernels. These experiments support slower-growing kernels for learning maps between heavy-tailed measures.
Type
Publication
NeurIPS 2026 Workshop on AI for Stochastic Dynamics (accepted)