HieRD: Hierarchical Relational Distillation for Vision-Language Embedding Models
V Le, N Hong Dang, T Vu, LN Van, DA Nguyen, T Le
International Conference on Machine Learning (ICML)

Knowledge distillation, Wasserstein transfer, and efficient student models from large teachers.
To compress the knowledge of large foundation models into smaller, faster students without sacrificing performance on downstream tasks.
This direction studies how to distill large language, vision-language, and generative models into efficient students. Research spans Wasserstein knowledge distillation, hierarchical relational distillation, and data-free black-box transfer.
Paper ... accepted at ...
V Le, N Hong Dang, T Vu, LN Van, DA Nguyen, T Le
HT Vuong, T Le, Q Tran, LN Van, T Le
TN Vo, D Nguyen, T Le, K Do, S Gupta