Udbhav Bamba

Applied Scientist II, Amazon Lab126 · Bangalore, India

prof_pic.jpg
ubamba98@gmail.com

I am an Applied Scientist II at Amazon Lab126, where I work at the intersection of large language models and low-latency on-device inference. My current work scales model compression to enable on-device and cloud-based GenAI deployment for Alexa+ — spanning dense-to-MoE conversion, quantization-friendly training, and compressing 200B+ parameter mixture-of-experts models. Previously at Amazon Ads, I built and productionized vision-language models for content understanding and moderation. Before Amazon, I focused on resource-efficient machine learning as a research intern at Mila – Quebec AI Institute (with David Rolnick and Gintare Karolina Dziugaite) and co-founded Transmute AI Research to foster a research culture among students at IIT (ISM) Dhanbad.

These days I am working on building better mixture-of-experts systems and helping LLMs reason efficiently under strict compute budgets. Beyond my research, I am an active competitor on Kaggle, where I hold the Competitions Master title (top 0.02%, peak rank 45th worldwide). Apart from work, I enjoy gaming, binge-watching shows, and playing badminton.

For updated details, please see my Google Scholar / LinkedIn pages.

news

Jun 01, 2026 Received the Silver Reviewer Award at ICML 2026, placing among the top reviewers.
May 01, 2026 Three papers accepted at ICML 2026 :tada:XRPO (targeted exploration and exploitation for GRPO), DOT-MoE (differentiable optimal transport for MoEfication), and Reward Under Attack (robustness and hackability of process reward models).
Feb 26, 2026 S2D, on quantization-friendly conditioning of neural activations, accepted at CVPR 2026.
Jan 15, 2026 CRoPS, a training-free hallucination mitigation framework for vision-language models, accepted at TMLR.
Jul 01, 2025 Joined Amazon Lab126 as an Applied Scientist II, working on model compression for on-device and cloud GenAI deployment for Alexa+.

Selected publications

  1. xrpo.png
    XRPO: Pushing the Limits of GRPO with Targeted Exploration and Exploitation
    Udbhav Bamba, Minghao Fang, Yifan Yu, and 2 more authors
    In International Conference on Machine Learning (ICML), 2026
  2. dotmoe.png
    DOT-MoE: Differentiable Optimal Transport for MoEfication
    Udbhav Bamba, Arnav Chavan, Aryamaan Thakur, and 2 more authors
    In International Conference on Machine Learning (ICML), 2026
  3. s2d.png
    S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations
    Arnav Chavan, Nahush Lele, Udbhav Bamba, and 3 more authors
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  4. crops.png
    CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models
    Neeraj Anand, Samyak Jha, Udbhav Bamba, and 1 more author
    Transactions on Machine Learning Research, 2026
  5. reward.png
    Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
    Rishabh Tiwari, Aditya Tomar, Udbhav Bamba, and 5 more authors
    In International Conference on Machine Learning (ICML), 2026