Anhao Zhao.
Ph.D. Student at PolyU
I am Anhao Zhao, a joint Ph.D. student at the NLP Group of The Hong Kong Polytechnic University & EIT NLP of Eastern Institute of Technology, Ningbo, fortunately supervised by Dr. Xiaoyu Shen and Prof. Wenjie Li.

Research
My research focuses on LLM post-training and LLM efficiency, with the long-term goal of building general-purpose models that advance reasoning, interaction, and inference efficiency.
LLM post-training
I study how supervised fine-tuning (SFT), reinforcement learning (RL), and on-policy distillation (OPD) can develop stronger reasoning and agentic capability. My work explores SFT on self-generated data for efficient reasoning [On-Policy SFT], RL for adaptive streaming reasoning [AdaSR], and more stable and effective OPD [PowerOPD, KL Agreement Trap].
LLM efficiency
I explore efficiency at both the token and architecture levels. At the token level, I study alternatives to explicit reasoning tokens [Latent CoT Survey] and reasoning while reading to reduce response latency [StreamingThinker]. At the architecture level, I investigate token-adaptive computation [SkipGPT], the practical inference speedups of model pruning [Beyond FLOPs], and visual cache reuse and sparse interactions for efficient multimodal inference [miniReranker].
📬I am open to collaborations and discussions. Please feel free to reach out to me if you are interested in my research or any relevant topics.
News
Five papers accepted to EMNLP 2026: 3 Main Conference papers and 2 Findings papers 🎉!
Got one paper accepted by ICML 2026🎉!
Got one paper accepted by CVPR 2026 Spotlight🎉!
Got one paper accepted by ICLR 2026🎉!
Earlier news
Attended EMNLP 2025 in person for the first time — a truly exciting experience 🎉
Started my Ph.D. study at the NLP Group @ PolyU & EIT NLP, supervised by Dr. Xiaoyu Shen and Prof. Wenjie Li.
Got one paper accepted by EMNLP 2025🎉!
Released our new survey on Latent Chain-of-Thought Reasoning.
Got one paper accepted by ACL 2025🎉!
Got one paper accepted by ICML 2025🎉!
Got one paper accepted by EMNLP 2024🎉!
Publications
Most recent publications on Google Scholar.
* indicates equal contribution
Conference & Journal Papers
Escaping the KL Agreement Trap in On-Policy Distillation
Haoran Xin, Anhao Zhao, Ying Sun, Jin Li, Xiaoyu Shen, Hui Xiong
Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration
Shuhao Li, Guodong Du, Anhao Zhao, Wanyu Lin, Tianyu Yuan, Xiaoyu Shen
ArXiv Preprints
AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimization
Junlong Tong, Wenqi Xu, Yingqi Fan, Anhao Zhao, Xuan Lu, Yang Tan, Xiaoyu Shen
miniReranker: Efficient Multimodal Reranking through Visual Cache Reuse and Interaction Sparsity
Yingqi Fan, Xuan Lu, Anhao Zhao, Junlong Tong, Ping Nie, Kai Zou, Yunpu Ma, Wei Zhang, Xiaoyu Shen
ProactiveLLM: Learning Active Interaction for Streaming Large Language Models
Junlong Tong, Yao Zhang, Anhao Zhao, Yingqi Fan, Yunpu Ma, Xiaoyu Shen
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
Anhao Zhao, H Xin, Yingqi Fan, Junlong Tong, Wenjie Li, Xiaoyu Shen
SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
Zicheng He, Anhao Zhao, Xiaoyu Shen, Chen Wu, Lei He
Service
Reviewer/Program Committee Member:
ICLR27, ICLR26, CVPR26, ECCV26, ICML26, NeurIPS26
Teaching Assistant:
COMP1010_26271_C: Computational Thinking and Problem Solving
COMP 5311: Internet Infrastructure and Protocols, Fall 2025, PolyU
COMP 5532: DIGITAL TWINS & VIRTUAL HUMAN, Spring 2026, PolyU