Yikun Ban 班义琨
I am a professor in the School of Computer Science and Engineering at Beihang University and a member of the State Key Laboratory of Software Development Environment. Previously, I was a postdoc and obtained my Ph.D. degree at Computer Science, University of Illinois at Urbana-Champaign. Prior to this, I obtained my Master's degree from EECS, Peking University and bachelor's degree from Wuhan University.
I am interested in principled algorithms in the space of reinforcement learning and deep learning, to solve real-world sequential decision-making problems. Current research topics:
- Reinforcement Learning Foundation
- Multi-Agent Reinforcement Learning
- Ensemble Learning of LLMs
News
- [2026.9] Four papers were accepted by NeurIPS'26
- [2026.5] Two papers were accepted by KDD'26.
- [2026.4] Four papers were accepted by ICML'26.
- [2026.2] Check our very interesting Learning finding for LLM! Weak-Driven Learning: How Weak Agents make Strong Agents Stronger.
- [2026.1] Check our very interesting theoretical finding for RLVR! Your Group-Relative Advantage Is Biased.
- [2025.11] Welcome to check our survey! A Survey on LLM Ensemble.
Selected Preprint (* Equal Contribution, # Corresponding)
Agent Exploration Toward Artificial General Intelligence: A Survey [Paper] [Paper-SSRN] [Github] [Website]
Weak-Driven Learning: How Weak Agents make Strong Agents Stronger [Paper] [Github] [Hugging Face] [PaperWeekly] [小红书]
Selected Publications — 2026
* Equal Contribution, # Corresponding
Your Group-Relative Advantage Is Biased[Paper] [Hugging Face] [机器之心] [小红书]
Conference on Neural Information Processing Systems (NeurIPS'26)
Heterogeneous Agent Collaborative Reinforcement Learning [Paper] [Project Page] [Hugging Face] [机器之心] [小红书]
Conference on Neural Information Processing Systems (NeurIPS'26)
Frequency-Aware Flow Matching for Continuous and Consistent Robotic Action Generation
Conference on Neural Information Processing Systems (NeurIPS'26)
Adaptive Robust Estimator for Policy Optimization in Reinforcement Learning
Conference on Neural Information Processing Systems (NeurIPS'26)
Alignment-Free Multi-Modality Large-Small Model Bidirectional Collaboration with Missing Modality [DBLP]
ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD'26), pp. 1451–1461
CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling [paper]
ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD'26)
Does Your Reasoning Model Implicitly Know When to Stop Thinking? [paper] [Code] [Hugging Face] [小红书]
International Conference on Machine Learning (ICML'26)
Real-Time Aligned Reward Model beyond Semantics [paper]
International Conference on Machine Learning (ICML'26)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards [paper]
International Conference on Machine Learning (ICML'26)
T-POP: Test-Time Personalization with Online Preference Feedback [paper]
International Conference on Machine Learning (ICML'26)
Bi-level Hierarchical Neural Contextual Bandits for Online Recommendation [DBLP]
Transactions on Machine Learning Research (TMLR, 2026)
Harmonizing Gradient Matching For Fairness [DBLP]
Transactions on Machine Learning Research (TMLR, 2026)
International Joint Conferences on Artificial Intelligence (IJCAI'26)
Neural Exploitation and Exploration of Contextual Bandits [paper]
Journal of Machine Learning Research (JMLR, 2026)
GCL-OT: Graph Contrastive Learning with Optimal Transport for Heterophilic Text-Attributed Graphs [DBLP]
AAAI Conference on Artificial Intelligence (AAAI'26), pp. 25142–25150
Selected Publications — 2025 & Earlier
Thirty-ninth Conference on Neural Information Processing Systems (NeurIPS'25, Spotlight)
Thirty-ninth Conference on Neural Information Processing Systems (NeurIPS'25)
LLM-Forest: Ensemble Learning of LLMs with Graph-Augmented Prompts for Data Imputation [paper]
The 63rd Annual Meeting of the Association for Computational Linguistics, Findings (ACL'25)
Adaptive Sampling-based Dynamic Graph Learning for Information Diffusion Prediction [paper]
ACM Transactions on Information Systems (TOIS, 2025)
Can Graph Neural Networks Learn Language with Extremely Weak Text Supervision? [paper]
The 63rd Annual Meeting of the Association for Computational Linguistics, Main (ACL'25)
Robust Neural Contextual Bandit against Adversarial Corruptions [paper]
Thirty-eighth Conference on Neural Information Processing Systems (NeurIPS'24)
PageRank Bandits for Link Prediction [paper]
Thirty-eighth Conference on Neural Information Processing Systems (NeurIPS'24)
ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD'24)
The Web Conference, Tutorial (WWW'24)
International Conference on Learning Representations (ICLR'24)
Contextual Bandits with Online Neural Regression [paper]
International Conference on Learning Representations (ICLR'24)
Meta-Learning with Neural Bandit Scheduler [paper]
Thirty-seventh Conference on Neural Information Processing Systems (NeurIPS'23)
ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD'23)
Thirty-sixth Conference on Neural Information Processing Systems (NeurIPS'22)
DISCO: Comprehensive and Explainable Disinformation Detection [paper]
ACM International Conference on Information and Knowledge Management (CIKM'22, Demo Track)
Neural Bandit with Arm Group Graph [paper]
ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD'22)
International Conference on Learning Representations (ICLR'22, Spotlight)
Preprint: ArXiv:2107.07438
ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD'21)
The Web Conference (WWW'21)
Dynamic Knowledge Graph Alignment [paper]
AAAI Conference on Artificial Intelligence (AAAI'21)
ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD'20)
Education
Aug 2019 – Jun 2024
Ph.D., Computer Science, Advised by Jingrui He and Hanghang Tong
University of Illinois at Urbana-Champaign, Illinois, US
Aug 2016 – Jul 2019
M.S., Computer Science
Peking University, Beijing, China
Aug 2012 – Jul 2016
B.S., School of Software Engineering
Wuhan University, Wuhan, China