Shresth Verma
I am a fourth-year PhD student at Harvard University advised by Prof. Milind Tambe, working on LLM post-training and alignment, robust learning from synthetic data, and LLMs for public health. My recent work focuses on
- Synthetic data selection for multi-turn LLM fine-tuning
- Preference-robust post-training of LLMs
- Balancing tradeoffs in multi-objective preference alignment
- Using LLMs to design resource allocation policies in public health
In Summer 2026, I was an Applied Scientist Intern at Amazon Science, where I worked on identifying LLM-agent and traditional bot traffic streams on retail websites.
Previously, I spent two wonderful years at Google DeepMind (formerly Google Research India), working in the AI for Social Good lab where I was grateful to be advised by Dr. Aparna Taneja. I developed and deployed robust bandit algorithms to plan targeted mobile health interventions for more than 100K beneficiaries from underserved communities in India.
Before that, I was a Data Scientist at United Health Group where I worked in the Chief Medical Officer’s team for modelling readmission risks for millions of beneficiaries. I also worked with data from the world’s largest healthcare graph database, designing graph-based analytics and tools to model and interpret patients’ longitudinal wellness journeys.
| | | | |
| Summer 2026 | 2023 - Present | 2021 - 2023 | 2020 - 2021 |
News
| Sep 25, 2026 | Our work on synthetic data for multi-turn LLM fine-tuning is accepted at NeurIPS 2026! |
|---|---|
| May 18, 2026 | Starting as an Applied Scientist Intern at Amazon Science! |
| Apr 15, 2026 | Gave a talk on preference-robust fine-tuning of LLMs for public health at the NSF AI Institute for Societal Decision Making seminar. |
| Nov 10, 2025 | Our work on Preference-Robust DPO is accepted at AAAI 2026! |
| May 12, 2025 | Our work on Portfolios for Multi-objective RL has been accepted at ICML 2025. See you in Vancouver! |
| Nov 3, 2024 | I will be presenting a demo on LLMs for RL Code Generation at AAAI 2025. See you in Philly! |
| Oct 19, 2024 | Presenting my work on Social Choice Language Model at NeurIPS 2024 GenAI for Health Workshop. See you in Vancouver! |
Selected Publications
2026
- NeurIPS’26Bilevel Optimization of Synthetic Trajectories for Multi-Turn LLM Fine-TuningIn Advances in Neural Information Processing Systems, 2026
- AAAI’26DPO-PRO: Direct Preference Optimization with Preference RobustnessIn AAAI Conference on Artificial Intelligence, 2026
2025
- AAAI’25PRIORITY2REWARD: Incorporating Healthworker Preferences for Resource Allocation PlanningIn Proceedings of the AAAI Conference on Artificial Intelligence, 2025
- ICML’25Navigating the Social Welfare Frontier: Portfolios for Multi-objective Reinforcement LearningIn Proceedings of the 42nd International Conference on Machine Learning, 13–19 jul 2025
2023
- IJCAI’23Limited Resource Allocation in a Non-Markovian World: The Case of Maternal and Child HealthcareIn International Joint Conference on Artificial Intelligence, 13–19 jul 2023
- AAAI’23Scalable decision-focused learning in restless multi-armed bandits with application to maternal and child healthIn AAAI Conference on Artificial Intelligence, 13–19 jul 2023
- AAMAS’23Restless Multi-Armed Bandits for Maternal and Child Health: Results from Decision-Focused Learning.In International Conference on Autonomous Agents and Multi Agent Systems, 13–19 jul 2023