About

I am a Ph.D. candidate in Electrical & Computer Engineering at the University of Southern California, where I completed master's degrees in both Computer Science and Electrical Engineering in May 2025. My research, advised by Prof. Rahul Jain and Prof. Ashutosh Nayyar, centers on reinforcement learning (RL), with a particular focus on offline and robust imitation learning, behavior foundation models, and post-training LLMs. I also collaborate closely with Prof. Paria Rashidinejad on RL for LLMs research.

In May 2026, I joined Google as a Student Researcher. In summer 2025, I was an Applied Scientist intern at Amazon, working on reinforcement learning for agentic AI systems. Earlier, as a Research Engineer at Samsung Research, I applied deep RL to network resource problems in the 6G Lab. As an undergraduate I worked with Prof. Sriparna Bandopadhyay at IIT Guwahati on data augmentation with r-cyclic matrices.

Before that, I spent summer 2018 with Prof. Richard James at the University of Minnesota modeling light-induced phase transitions, and summer 2017 with Prof. Frank Chung-Hoon Rhee at Hanyang University, estimating the fuzzifier parameter for alpha-planes of general type-2 fuzzy sets.

Reinforcement Learning Imitation Learning Behavior Foundation Models Agentic AI Large Language Models Post-training
Selected Research
AdviSD: an advisor model learns to advise a frontier LLM via targeted multi-turn self-distillation

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

Rishabh Agrawal, Hejie Cui, Shasha Li, Shanchan Wu, Sercan Ö. Arık

arXiv preprint 2026

Abstract

A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2–6.4 percentage points on BFCL-v3 and by 3.9–5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.

DistIL: distributional DAgger with future-aware credit assignment

Reinforcement Learning from Rich Feedback with Distributional DAgger

Rishabh Agrawal, Jacob Fein-Ashley, Paria Rashidinejad

Advances in Neural Information Processing Systems (NeurIPS) 2026

Also accepted at
  • ICML 2026 Workshop on RL from World Feedback
  • 3rd AI for Math Workshop: Toward Self-Evolving Scientific Agents at ICML 2026
  • Continual Reinforcement Learning Workshop at RLC 2026
  • Second Workshop on the Foundations of Post-training at COLT 2026
Abstract

The dominant RL-from-verifiable-rewards recipe rewards each response with a single correctness bit, yet many settings provide far richer feedback — execution traces, tool outputs, expert corrections, self-evaluations. We study how to use such feedback through DistIL, a distributional variant of DAgger that optimizes a forward cross-entropy objective. Unlike reverse-KL or Jensen–Shannon self-distillation, DistIL guarantees monotonic policy improvement and sublinear regret, performs future-aware credit assignment, and improves Pass@N across scientific reasoning, coding, and hard mathematics.

RBFM: robust task inference under dynamics shift

When Dynamics Shift, Robust Task Inference Wins: Offline Imitation Learning with Behavior Foundation Models Revisited

Rishabh Agrawal, Rahul Jain, Ashutosh Nayyar

Advances in Neural Information Processing Systems (NeurIPS) 2026

Abstract

Behavior Foundation Models (BFMs) enable scalable imitation learning but assume fixed dynamics, leaving them brittle to real-world shifts in friction, actuation, or sensor noise. We recast BFM task inference as a robust minimax problem and introduce RBFM-Light and RBFM-Heavy — two variants that add robustness only at inference, with no change to pretraining and using offline data from a single nominal environment. Both substantially outperform standard BFM and robust offline IL baselines under dynamics shifts.

BE-DROIL method overview

Balance Equation-based Distributionally Robust Offline Imitation Learning

Rishabh Agrawal, Yusuf Alvi, Rahul Jain, Ashutosh Nayyar

8th Annual Conference on Learning for Dynamics and Control (L4DC) 2026

Also accepted at
  • Workshop on Embodied and Safe-Assured Robotic Systems at NeurIPS 2025 (Oral presentation)
Abstract

Standard imitation learning implicitly assumes the environment stays fixed between training and deployment — an assumption that rarely holds. We learn robust policies from expert demonstrations alone by solving a distributionally robust optimization over an uncertainty set of transition models, and show the worst-case objective can be rewritten entirely in terms of the nominal data distribution, enabling tractable offline learning with stronger robustness under shifted dynamics.

Markov Balance Satisfaction Improves Performance in Strictly Batch Offline Imitation Learning

Rishabh Agrawal, Nathan Dahlin, Rahul Jain, Ashutosh Nayyar

Proceedings of the AAAI Conference on Artificial Intelligence 2025

Abstract

We study imitation in a strictly offline setting — no environment interaction, no auxiliary data, no transition model. Our method uses the Markov balance equation with a conditional density estimation framework, employing conditional normalizing flows for dynamics, and consistently outperforms many state-of-the-art IL algorithms across Classic Control and MuJoCo.

View all publications, patents & projects
News
  1. Sep 2026
    Our paper “AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation”, from my Student Researcher work at Google , is out on arXiv.
  2. Sep 2026
    Our paper “Reinforcement Learning from Rich Feedback with Distributional DAgger” was accepted to NeurIPS 2026.
  3. Sep 2026
  4. May 2026
    Started as a Student Researcher at Google .
  5. Jan 2026
  6. Nov 2025
    Passed my Ph.D. Qualifying Exam, officially a Ph.D. Candidate!
  7. Nov 2025
    Selected for the AAAI 2026 Doctoral Consortium in Singapore.
  8. Nov 2025
    Our paper “Balance Equation-based Distributionally Robust Offline Imitation Learning” was accepted to the NeurIPS 2025 E-SARS Workshop for an oral presentation.
  9. May 2025
    Started as an Applied Scientist Intern at Amazon .
  10. May 2025
    Awarded two M.S. degrees in Computer Science and Electrical Engineering at USC.
  11. Apr 2025
    Gave a talk on offline imitation learning at the 45th SoCal Control Workshop, UC San Diego.
  12. Feb 2025
  13. Dec 2024
  14. Nov 2024
    Received the Outstanding Poster Award at USC's 14th Annual Research Festival.
  15. Sep 2024
    Our paper “Policy Optimization for Strictly Batch Imitation Learning” was accepted to the Workshop on Optimization for Machine Learning at NeurIPS 2024.
  16. Dec 2022
  17. Jan 2022
    Started my Ph.D. at USC with a broad focus on reinforcement learning.
  18. Aug 2020
    Our paper “A Reinforcement Learning Framework for QoS-Driven Radio Resource Scheduler” was accepted to IEEE Globecom 2020.
  19. Jun 2019
    Joined Samsung Research as a Research Engineer in the 6G Lab.
  20. May 2019
    Graduated from IIT Guwahati with a B.Tech in Mathematics & Computing.
  21. May 2018
    Started a summer research internship at the University of Minnesota, Twin Cities.
  22. Mar 2018
  23. May 2017
    Started a summer research internship at Hanyang University, South Korea.
  24. Jul 2015
    Started undergraduate studies at IIT Guwahati (Mathematics, CS & Financial Engineering).
Contact

I'm always glad to talk research and open to new collaborations. The quickest way to reach me is email, feel free to say hello.

Email
Office
335 Hughes Aircraft Electrical Engineering Center
3740 McClintock Ave, Los Angeles, CA 90089