Hello, I’m Wendi.

I am currently a first-year PhD student at University of Wisconsin–Madison, where I am fortunate to be advised by Prof. Sharon Li. I received my B.E. and M.E. in Computer Science and Technology from Huazhong University of Science and Technology.

My current research interests lie in reinforcement learning algorithms for large-scale models and their downstream applications, such as agentic systems.

I am always happy to chat and discuss about potential collaborations or my past research projects. Feel free to contact me via email wli679 AT wisc.edu

Research Interests

01

Reinforcement Learning

Learning reliable policies and reward signals for large-scale foundation models.

02

Reasoning Models

Improving exploration, process supervision, and test-time reasoning behavior.

03

Agentic Systems

Building capable agents that can plan, learn from feedback, and act robustly.

News

  • Progress Advantage won the best paper award at workshop RLxF@ICML 2026
  • I started my summer internship as a Research Intern in Microsoft, Redmond, WA.
  • Two papers were accepted at ACL 2026.
  • GEB was accepted by ICLR 2026.
  • GEB was selected for an oral presentation at ResponsibleFM@NeurIPS 2026.
  • I started my PhD journey at UW–Madison.
  • Free Process Rewards without Process Labels was accepted by ICML 2025.
  • PQM was accepted by ICLR 2025.
  • One paper was accepted by Findings of NAACL 2024.
  • I received the National Scholarship (top 3% nationwide).
  • One paper was accepted by the ACL 2023 main conference.

Publications

Google Scholar

14 papers

Papers can appear in more than one topic. * denotes equal contribution.

  • Illustration from Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning

    arXiv · 2026

    Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning

    Wendi Li, Shawn Im, Sharon Li

  • Illustration from Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

    arXiv · 2026

    RLxF @ ICML 2026

    Best Paper

    Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

    Changdae Oh, Wendi Li, Seongheon Park, Samuel Yeh, Tanwi Mallick, Sharon Li

  • Illustration from Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring

    arXiv · 2026

    Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring

    Seongheon Park, Wendi Li, Changdae Oh, Samuel Yeh, Zsolt Kira, Michael Hagenow, Sharon Li

  • Illustration from When Vision Speaks for Sound

    arXiv · 2026

    When Vision Speaks for Sound

    Xiaofei Wen, Wenjie Jacky Mo, Xingyu Fu, Rui Cai, Tinghui Zhu, Wendi Li, Yanan Xie, Muhao Chen, Peng Qi

  • Illustration from LAD: Learning Advantage Distribution for Reasoning

    ACL · 2026

    LAD: Learning Advantage Distribution for Reasoning

    Wendi Li, Sharon Li

  • Illustration from Towards Reducible Uncertainty Modeling for Reliable Large Language Model Agents

    ACL · 2026

    Towards Reducible Uncertainty Modeling for Reliable Large Language Model Agents

    Changdae Oh, Seongheon Park, To Eun Kim, Jiatong Li, Wendi Li, Samuel Yeh, Xuefeng Du, Hamed Hassani, Paul Bogdan, Dawn Song, Sharon Li

  • Illustration from Process Reinforcement through Implicit Rewards

    TMLR · 2026

    Process Reinforcement through Implicit Rewards

    Ganqu Cui*, Lifan Yuan*, Zefan Wang*, Hanbin Wang*, Wendi Li*, Bingxiang He*, Yuchen Fan*, Tianyu Yu*, Qixin Xu*, et al.

  • Illustration from General Exploratory Bonus for Optimistic Exploration in RLHF

    ICLR · 2026

    ResponsibleFM @ NeurIPS 2026

    Oral

    General Exploratory Bonus for Optimistic Exploration in RLHF

    Wendi Li, Changdae Oh, Yixuan Li

  • Illustration from Free Process Rewards without Process Labels

    ICML · 2025

    Free Process Rewards without Process Labels

    Lifan Yuan*, Wendi Li*, Huayu Chen, Ganqu Cui, Ning Ding, Kaiyan Zhang, Bowen Zhou, Zhiyuan Liu, Hao Peng

  • Illustration from Process Reward Model with Q-value Rankings

    ICLR · 2025

    Process Reward Model with Q-value Rankings

    Wendi Li, Yixuan Li

  • Illustration from Reinforcement Learning with Token-level Feedback for Controllable Text Generation

    NAACL · 2024

    Reinforcement Learning with Token-level Feedback for Controllable Text Generation

    Wendi Li, Wei Wei, Kaihe Xu, Wenfeng Xie, Dangyang Chen, Yu Cheng

  • Illustration from Position Debiasing Fine-Tuning for Causal Perception in Long-Term Dialogue

    IJCAI · 2024

    Position Debiasing Fine-Tuning for Causal Perception in Long-Term Dialogue

    Shixuan Fan, Wei Wei, Wendi Li, Xian-Ling Mao, Wenfeng Xie, Dangyang Chen

  • Illustration from TREA: Tree-Structure Reasoning Schema for Conversational Recommendation

    ACL · 2023

    TREA: Tree-Structure Reasoning Schema for Conversational Recommendation

    Wendi Li, Wei Wei, Xiaoye Qu, Xian-Ling Mao, Ye Yuan, Wenfeng Xie, Dangyang Chen

  • Illustration from Towards Hierarchical Policy Learning for Conversational Recommendation with Hypergraph-based Reinforcement Learning

    IJCAI · 2023

    Towards Hierarchical Policy Learning for Conversational Recommendation with Hypergraph-based Reinforcement Learning

    Sen Zhao, Wei Wei, Yifan Liu, Ziyang Wang, Wendi Li, Xian-Ling Mao, Shuai Zhu, Minghui Yang, Zujie Wen

Outside Research

I enjoy literature, movies, and music.

Recent favorite books: Life Ceremony by Sayaka Murata, All the Lovers in the Night by Mieko Kawakami, Satantango by László Krasznahorkai, and The Hunter by Shuang Xuetao.

Recent favorite films: Happy as Lazzaro and La Chimera by Alice Rohrwacher, If I Had Legs, I'd Kick You by Mary Bronstein, Kaili Blues by Gan Bi, The Florida Project by Sean Baker, Desert of Namibia by Yoko Yamanaka

Music: Billie Eilish, Lana Del Rey, Jude Chiu, and Qing-Feng Wu.