π€ About Me
I am Weixin Wang, a fourth year Ph.D. candidate in Prof. Pan Xuβs lab at Duke University.
My research centers on Machine Learning, with a focus on developing computationally and data-efficient algorithms that feature both strong empirical performance and rigorous theoretical guarantees. My core interests span Reinforcement Learning (RL) Theory, particularly Thompson Sampling, Ensemble Sampling, and other Randomized Exploration methods in RL. I also maintain broad interests in Diffusion Models, Large Language Models (LLMs), Robust RL, Artificial Intelligence, and High-Dimensional Statistics.
I am fortunate to collaborate with Yu Yang, Zhishuai Liu (lab members), Wei Deng, Andrew Bennett, Ruoxi Cheng and many other excellent researchers.
π₯ News
- 2026.6: Β I will be joining Morgan Stanley in New York this summer as a Machine Learning Research Associate mentored by Andrew Bennett!
- 2026.1: Β ππ Two papers are accepted to ICLR 2026!
- 2025.12: Β I attended NeurIPS 2025 at San Diego!
- 2025.8: Β I attended Princeton 2025 Machine Learning Theory Summer School!
- 2025.5: Β ππ Sample Complexity of Distributionally Robust Off-Dynamics Reinforcement Learning with Online Interaction is accepted as Poster to ICML 2025!
- 2024.12: Β I attended NeurIPS 2024 at Vancouver!
- 2024.9: Β ππ Randomized Exploration in Cooperative Multi-agent Reinforcement Learning is accepted as Poster to NeurIPS 2024!
π Publications
-
Decoupled Marginal Sharpening for Training-Free Inference-Time Scaling
Weixin Wang, Wei Deng, Anderson Schneider, Yuriy Nevmyvaka, Pan Xu, Andrew Bennett
Under review.
-
Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo [Paper]
Weixin Wang*, Yu Yang*, Wei Deng, Pan Xu
Under review.
-
Rethinking Langevin Thompson Sampling from A Stochastic Approximation Perspective [Paper]
Weixin Wang*, Haoyang Zheng*, Guang Lin, Wei Deng, Pan Xu
NeurIPS 2025 Workshop: Dynamics at the Frontiers of Optimization, Sampling, and Games (DynaFront).
-
Diffusion Posterior Sampling for Fast Adaptation in Multi-task Nonlinear Contextual Bandits
Weixin Wang*, Yu Yang*, Pan Xu
Under review.
-
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment [Paper]
Ruoxi Cheng*, Haoxuan Ma*, Weixin Wang*, Ranjie Duan, Jiexi Liu, Xiaoshuang Jia, Simeng Qin, Xiaochun Cao, Yang Liu, Xiaojun Jia
In Proc. of the 14th International Conference on Learning Representations (ICLR), Rio de Janeiro, Brazil, 2026.
-
Provable Anytime Ensemble Sampling Algorithms in Nonlinear Contextual Bandits [Paper]
Jiazheng Sun*, Weixin Wang*, Pan Xu
Under review.
-
Breaking the Total Variance Barrier: Sharp Sample Complexity for Linear Heteroscedastic Bandits with Fixed Action Set
Heyang Zhao*, Tianyuan Jin*, Weixin Wang, Vincent Y. F. Tan, Pan Xu, Quanquan Gu
In Proc. of the 14th International Conference on Learning Representations (ICLR), Rio de Janeiro, Brazil, 2026.
-
Near-Optimal Reinforcement Learning for Linear Distributionally Robust Markov Decision Processes [Paper]
Zhishuai Liu*, Weixin Wang*, Pan Xu
Reinforcement Learning Journal (RLJ), vol. 7, 2026. Presented at the Third Reinforcement Learning Conference (RLC 2026), MontrΓ©al, Canada.
-
Sample Complexity of Distributionally Robust Off-Dynamics Reinforcement Learning with Online Interaction [Paper]
Yiting He*, Zhishuai Liu*, Weixin Wang, Pan Xu
In Proc. of the 42nd International Conference on Machine Learning (ICML), Vancouver, Canada, 2025.
-
Randomized Exploration in Cooperative Multi-Agent Reinforcement Learning [Paper] [Code]
Hao-Lun Hsu*, Weixin Wang*, Miroslav Pajic, Pan Xu
In Proc. of the 38th Conference on Advances in Neural Information Processing Systems (NeurIPS), Vancouver, Canada, 2024.
π» Internships
- 2026.06 - 2026.08, Machine Learning Research Associate, Morgan Stanley, New York.
π Educations
- 2023.08 - now, Ph.D., Department of Electrical and Computer Engineering, Duke University.
- 2019.09 - 2023.06, B.S., School of the Gifted Young, University of Science and Technology of China.