Shresth Verma

me.jpeg

I am a fourth-year PhD student at Harvard University advised by Prof. Milind Tambe, working on LLM post-training and alignment, robust learning from synthetic data, and LLMs for public health. My recent work focuses on

  • Synthetic data selection for multi-turn LLM fine-tuning
  • Preference-robust post-training of LLMs
  • Balancing tradeoffs in multi-objective preference alignment
  • Using LLMs to design resource allocation policies in public health

In Summer 2026, I was an Applied Scientist Intern at Amazon Science, where I worked on identifying LLM-agent and traditional bot traffic streams on retail websites.

Previously, I spent two wonderful years at Google DeepMind (formerly Google Research India), working in the AI for Social Good lab where I was grateful to be advised by Dr. Aparna Taneja. I developed and deployed robust bandit algorithms to plan targeted mobile health interventions for more than 100K beneficiaries from underserved communities in India.

Before that, I was a Data Scientist at United Health Group where I worked in the Chief Medical Officer’s team for modelling readmission risks for millions of beneficiaries. I also worked with data from the world’s largest healthcare graph database, designing graph-based analytics and tools to model and interpret patients’ longitudinal wellness journeys.

Amazon Science Harvard University Google DeepMind UnitedHealth Group
Summer 2026 2023 - Present 2021 - 2023 2020 - 2021

News

Sep 25, 2026 Our work on synthetic data for multi-turn LLM fine-tuning is accepted at NeurIPS 2026!
May 18, 2026 Starting as an Applied Scientist Intern at Amazon Science!
Apr 15, 2026 Gave a talk on preference-robust fine-tuning of LLMs for public health at the NSF AI Institute for Societal Decision Making seminar.
Nov 10, 2025 Our work on Preference-Robust DPO is accepted at AAAI 2026!
May 12, 2025 Our work on Portfolios for Multi-objective RL has been accepted at ICML 2025. See you in Vancouver!
Nov 3, 2024 I will be presenting a demo on LLMs for RL Code Generation at AAAI 2025. See you in Philly!
Oct 19, 2024 Presenting my work on Social Choice Language Model at NeurIPS 2024 GenAI for Health Workshop. See you in Vancouver!

Selected Publications

2026

  1. NeurIPS’26
    Bilevel Optimization of Synthetic Trajectories for Multi-Turn LLM Fine-Tuning
    Shresth Verma, Mauricio Tec, Cheol Woo Kim, Kai Wang, and Milind Tambe
    In Advances in Neural Information Processing Systems, 2026
  2. AAAI’26
    DPO-PRO: Direct Preference Optimization with Preference Robustness
    Cheol Woo Kim*, Shresth Verma*, Mauricio Tec, and Milind Tambe
    In AAAI Conference on Artificial Intelligence, 2026

2025

  1. AAAI’25
    PRIORITY2REWARD: Incorporating Healthworker Preferences for Resource Allocation Planning
    Shresth Verma, Alayna Nguyen, Niclas Boehmer, Lingkai Kong, and Milind Tambe
    In Proceedings of the AAAI Conference on Artificial Intelligence, 2025
  2. ICML’25
    Navigating the Social Welfare Frontier: Portfolios for Multi-objective Reinforcement Learning
    Cheol Woo Kim, Jai Moondra, Shresth Verma, Madeleine Pollack, Lingkai Kong, and 2 more authors
    In Proceedings of the 42nd International Conference on Machine Learning, 13–19 jul 2025

2023

  1. IJCAI’23
    Limited Resource Allocation in a Non-Markovian World: The Case of Maternal and Child Healthcare
    Panayiotis Danassis, Shresth Verma, Jackson A. Killian, Aparna Taneja, and Milind Tambe
    In International Joint Conference on Artificial Intelligence, 13–19 jul 2023
  2. AAAI’23
    Scalable decision-focused learning in restless multi-armed bandits with application to maternal and child health
    Kai Wang*, Shresth Verma*, Aditya Mate, Sanket Shah, Aparna Taneja, and 3 more authors
    In AAAI Conference on Artificial Intelligence, 13–19 jul 2023
  3. AAMAS’23
    Restless Multi-Armed Bandits for Maternal and Child Health: Results from Decision-Focused Learning.
    Shresth Verma, Aditya Mate, Kai Wang, Neha Madhiwalla, Aparna Hegde, and 2 more authors
    In International Conference on Autonomous Agents and Multi Agent Systems, 13–19 jul 2023