Welcome! I am Hui Yuan [xweɪ ɥæn], and my name in Chinese is 袁慧. I am a Research Scientist at Utopai Studios, where I work on long-horizon agents for interactive video creation. Previously, I was a Research Scientist at Meta.
My research spans reinforcement learning, long-horizon agents, and generative-model post-training. Application-wise, I am particularly interested in developing creative intelligence through agentic media generation systems.
I received my Ph.D. in Electrical and Computer Engineering from Princeton University, advised by Mengdi Wang, and my B.S. in Statistics from the University of Science and Technology of China.
I have also been fortunate to work with Yinyu Ye, Csaba Szepesvári, and Yingyu Liang.
Research
Selected Publications
FMIP: Joint Continuous-Integer Flow for Mixed-Integer Linear Programming
MATH-Perturb: Benchmarking LLMs’ Math Reasoning Abilities against Hard Perturbations
A Common Pitfall of Margin-Based Language Model Alignment: Gradient Entanglement
Training-Free Guidance Beyond Differentiability: Scalable Path Steering with Tree Search in Diffusion and Flow Models
MaxMin-RLHF: Towards Equitable Alignment of Large Language Models with Diverse Human Preferences
Gradient Guidance for Diffusion Models: An Optimization Perspective
Reward-Directed Conditional Diffusion: Provable Distribution Estimation and Reward Improvement
Additional Publications
- A First-Order Generative Bilevel Optimization Framework for Diffusion Models. Q Xiao, Hui Yuan, A F M Saif, G Liu, R R Kompella, M Wang, T Chen. ICML 2025. Paper Code
- Conversational Dueling Bandits in Generalized Linear Models. S Yang, Hui Yuan, X Zhang, M Wang, H Zhang, H Wang. KDD 2024. Paper Code
- Tree Search-Based Evolutionary Bandits for Protein Sequence Optimization. J Qiu*, Hui Yuan*, J Zhang*, W Chen, H Wang, M Wang. AAAI 2024. Paper
- Unified Off-Policy Learning to Rank: A Reinforcement Learning Perspective. Zeyu Zhang, Yi Su, Hui Yuan, Yiran Wu, Rishab Balasubramanian, Qingyun Wu, Huazheng Wang, Mengdi Wang. NeurIPS 2023. Paper Code
- Bandit Theory and Thompson Sampling-Guided Directed Evolution for Sequence Optimization. Hui Yuan, Huazheng Wang, Chengzhuo Ni, Xuezhou Zhang, Le Cong, Csaba Szepesvári, Mengdi Wang. NeurIPS 2022. Paper
- Learning Entangled Single-Sample Gaussians in the Subset-of-Signals Model. Yingyu Liang, Hui Yuan. COLT 2020. Paper
- Learning Entangled Single-Sample Distributions via Iterative Trimming. Hui Yuan, Yingyu Liang. AISTATS 2020. Paper
- Uniform Joint Screening for Ultra-High Dimensional Graphical Models. Z Zheng, H Shi, Y Li, Hui Yuan. Journal of Multivariate Analysis. Paper