Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence
TP Nguyen, T Nguyen, MP Truong, T Nguyen, J Bailey, T Le
arXiv preprint arXiv:2605.13079

Kernel methods, optimization theory, and convergence analysis for modern deep learning.
To establish the theoretical foundations that explain how and why machine learning algorithms work, from kernel embeddings to optimizer dynamics.
This direction bridges classical machine learning theory with the training dynamics of deep networks. It spans kernel methods, large-scale optimization, Wasserstein distances, and principled analysis of learning rates and convergence.
Paper ... accepted at ...
TP Nguyen, T Nguyen, MP Truong, T Nguyen, J Bailey, T Le
T Le, T Nguyen, D Phung