I am broadly interested in understanding machine learning systems through mathematics. For example, I think a lot about approximation theory, statistical learning theory, geometry, and dynamical systems. I am also tangentially interested in interactions between ML and more abstract fields (category theory, homotopy theory, algebraic topology).
Click on the for a short summary.not this one...
Preprints
- Training-Free Universal Approximation by Prompting Random Transformers. Alexander Hsu, Rongjie Lai. (2026) [arXiv]
We show that a single-layer softmax attention transformer with random Gaussian weights can approximate any Hölder function on a compact manifold when guided by an appropriate soft prompt, which we construct explicitly. The approach builds on the connection between attention and kernel regression to achieve universal approximation with minimax-optimal rates.
- Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer. Alexander Hsu, Zhaiming Shen, Wenjing Liao, Rongjie Lai. (2026) [arXiv]
We construct shallow transformers that use attention to exactly realize nonlinear features in-context, such as polynomial and spline bases, via explicit arithmetic computations in the forward pass. We derive generalization error bounds for the resulting nonlinear ICL pipeline.
- The layer number of grids. Gergely Ambrus, Alexander Hsu, Bo Peng, Shiyu Yan. (2020) [arXiv]
We study the number of convex layers for integer grids in higher dimensions. Undergrad summer research (unfortunately over Zoom). Never published; at the time there were sharper bounds via different techniques, and eventually our results were subsumed by newer work.
Publications
- Understanding In-Context Learning on Structured Manifolds: Bridging Attention to Kernel Methods. Zhaiming Shen, Alexander Hsu, Rongjie Lai, Wenjing Liao. International Conference on Learning Representations (ICLR), 2026.
We explore a connection between attention and kernel methods, which we extend to derive generalization error bounds for in-context kernel regression on manifolds. Along the way, we construct a transformer which realizes kernel regression in the ICL setting.