Pu Wang 王普

Pu Wang 

Pu Wang
Undergraduate student
Turing Class, College of Computer Science and Technology,
Zhejiang University
Email: puwang0508@gmail.com
GitHub / Notebook

About me

I am a fourth-year undergraduate student (Sept. 2023  —  Present) in Turing Class, Chu Kochen Honors College, Zhejiang University, pursuing a B.E. in Artificial Intelligence with an honors degree. Since March 2025 I have been a research intern at the State Key Lab of CAD&CG, advised by Prof. Yao-Xiang Ding.

My research interests lie in the theory and algorithms of sequential decision making, with a growing focus on large language models and language agents. I am interested in how agents acquire, through interaction, the knowledge they need to act well, and how that knowledge can be learned efficiently and reused across tasks and environments. More broadly, while algorithms and theory remain at the core of my work, I aim to bring principled formulations from bandits and reinforcement learning to problems related to decision making in language models and language agents.

Research directions

  • Algorithms and theory for decision making. Bandits and reinforcement learning with provable guarantees, e.g., a unified framework for best arm identification and regret minimization in dueling bandits.

  • Learning from indirect and incomplete information. What agents can learn from comparisons, imperfect teachers, or partial observations, and how to learn it efficiently.

  • Language models and language agents. Bringing principled formulations from decision making to language models and agents, such as how they acquire and reuse knowledge through interaction.

News

  • September 2026  —  Our paper on a unified framework for dueling bandits (TG-ITE) is accepted to NeurIPS 2026.

  • June 2026  —  Our work on a unified framework for dueling bandits (TG-ITE) is available as a preprint on arXiv.

Selected Papers

Tree-Guided Identify-Then-Exploit: A Unified Framework of Best Arm Identification and Regret Minimization for Dueling Bandits
Pu Wang, Yao-Xiang Ding
NeurIPS 2026: The Fortieth Annual Conference on Neural Information Processing Systems
tl;dr: The first unified framework for N-armed dueling bandits that jointly handles best arm identification (BAI), weak regret, and strong regret under only the Condorcet-winner assumption. It achieves O(N) sample complexity for BAI and O(N) weak regret, and removes the O(log N) suboptimality gap of the existing approach when optimizing both jointly.

For the full list, see Publications.