Udbhav Bamba

Applied Scientist II, Amazon Lab126 · Bangalore, India

prof_pic.jpg
ubamba98@gmail.com

I am an Applied Scientist II at Amazon Lab126, where I build efficient large language models for Alexa+. My work focuses on mixture-of-experts systems and model compression, including dense-to-MoE conversion and quantization-friendly training for low-latency on-device and cloud inference at 200B+ parameter scale. Previously, I productionized vision-language models at Amazon Ads, researched resource-efficient machine learning at Mila, and co-founded Transmute AI Research at IIT (ISM) Dhanbad.

I am also a Kaggle Competitions Master (top 0.02%, peak rank 45 worldwide). Outside work, I enjoy gaming, cycling, and badminton. For updates, see my Google Scholar or LinkedIn.

News

Jun 01, 2026 Received an ICML 2026 Silver Reviewer Award.
May 01, 2026 Three papers accepted at ICML 2026 :tada:: XRPO, DOT-MoE, and Reward Under Attack.
Feb 26, 2026 S2D accepted at CVPR 2026.
Jan 15, 2026 CRoPS accepted at TMLR.
Jul 01, 2025 Joined Amazon Lab126 as an Applied Scientist II, working on efficient GenAI for Alexa+.

Selected publications

* Equal contribution

  1. XRPO: Pushing the Limits of GRPO with Targeted Exploration and Exploitation
    Udbhav Bamba*, Minghao Fang*, Yifan Yu*, and 2 more authors
    In International Conference on Machine Learning (ICML), 2026
  2. DOT-MoE: Differentiable Optimal Transport for MoEfication
    Udbhav Bamba*, Arnav Chavan*, Aryamaan Thakur*, and 2 more authors
    In International Conference on Machine Learning (ICML), 2026
  3. S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations
    Arnav Chavan*, Nahush Lele*, Udbhav Bamba*, and 3 more authors
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  4. CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models
    Neeraj Anand, Samyak Jha, Udbhav Bamba, and 1 more author
    Transactions on Machine Learning Research, 2026
  5. Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
    Rishabh Tiwari*, Aditya Tomar*, Udbhav Bamba*, and 5 more authors
    In International Conference on Machine Learning (ICML), 2026