Research
I work on the mathematical theory of transformers — modeling tokens as interacting particles to understand how attention moves, clusters, and represents information. I work with Yury Polyanskiy and Philippe Rigollet at MIT.
Papers
Continuous First, Discrete Later: VQ-VAEs Without Dimensional Collapse
Trained VQ-VAE representations tend to collapse into a tiny subspace — 1–2% of full rank — and we show this dimensional collapse imposes a hard lower bound on the loss that codebook-improvement tricks cannot overcome. The fix is simple: warm up the model as an ordinary (unquantized) autoencoder before switching on vector quantization. On VQGAN and WavTokenizer, this warm-up restores representation dimension and improves reconstruction and perceptual quality at the same training budget, and our theory predicts the right moment to switch.
BibTeX
@misc{zhao2026continuousfirstdiscretelater,
title={Continuous First, Discrete Later: VQ-VAEs Without Dimensional Collapse},
author={Xinyu Zhao and Nikita Karagodin and Hamed Hassani and Sinan Hersek and Paul Pu Liang and Yury Polyanskiy},
year={2026},
eprint={2605.06870},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2605.06870},
} Normalization in Attention Dynamics
We model token representations in a deep transformer as interacting particles on the sphere and show that normalization acts as a form of speed regulation on their dynamics. This perspective gives a unified analysis of Post-LN, Pre-LN, Mix-LN, Peri-LN, and nGPT, explaining how each scheme shapes clustering and representation collapse across layers. The comparison identifies Peri-LN as a particularly effective choice.
BibTeX
@misc{karagodin2025normalizationattentiondynamics,
title={Normalization in Attention Dynamics},
author={Nikita Karagodin and Shu Ge and Yury Polyanskiy and Philippe Rigollet},
year={2025},
eprint={2510.22026},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2510.22026},
} Clustering in Causal Attention Masking
This is the first rigorous treatment of causally masked attention dynamics — the version transformers actually use for generation, which breaks the mean-field gradient-flow structure earlier analyses relied on. We prove that tokens converge to a single cluster for arbitrary key-query matrices with identity value matrix, going well beyond prior results that required all three matrices to be scaled identities. A connection to the classical Rényi parking problem takes first theoretical steps toward explaining the meta-stable states observed in practice.
BibTeX
@misc{karagodin2024clusteringcausalattentionmasking,
title={Clustering in Causal Attention Masking},
author={Nikita Karagodin and Yury Polyanskiy and Philippe Rigollet},
year={2024},
eprint={2411.04990},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2411.04990},
} Other fascinating things
Papers and ideas I didn't write but find remarkable.
The Unreasonable Effectiveness of Mathematics
The classic essay on why mathematical concepts developed in pure abstraction turn out to describe physical reality with uncanny precision.
Neural Ordinary Differential Equations
The paper that started the neural ODE revolution — replacing discrete residual layers with continuous dynamics defined by an ODE solver.