Anhao Zhao

Anhao Zhao.

Ph.D. Student at PolyU

Research interests:
  • OPD
  • Agentic RL
  • Post-training

I am Anhao Zhao, a joint Ph.D. student at the NLP Group of The Hong Kong Polytechnic University & EIT NLP of Eastern Institute of Technology, Ningbo, fortunately supervised by Dr. Xiaoyu Shen and Prof. Wenjie Li.

Anhao Zhao
Hung Hom, Hong Kong, China

Research

My research focuses on LLM post-training and LLM efficiency, with the long-term goal of building general-purpose models that advance reasoning, interaction, and inference efficiency.

LLM post-training

I study how supervised fine-tuning (SFT), reinforcement learning (RL), and on-policy distillation (OPD) can develop stronger reasoning and agentic capability. My work explores SFT on self-generated data for efficient reasoning [On-Policy SFT], RL for adaptive streaming reasoning [AdaSR], and more stable and effective OPD [PowerOPD, KL Agreement Trap].

LLM efficiency

I explore efficiency at both the token and architecture levels. At the token level, I study alternatives to explicit reasoning tokens [Latent CoT Survey] and reasoning while reading to reduce response latency [StreamingThinker]. At the architecture level, I investigate token-adaptive computation [SkipGPT], the practical inference speedups of model pruning [Beyond FLOPs], and visual cache reuse and sparse interactions for efficient multimodal inference [miniReranker].

📬I am open to collaborations and discussions. Please feel free to reach out to me if you are interested in my research or any relevant topics.

News

Five papers accepted to EMNLP 2026: 3 Main Conference papers and 2 Findings papers 🎉!

Got one paper accepted by ICML 2026🎉!

Got one paper accepted by CVPR 2026 Spotlight🎉!

Got one paper accepted by ICLR 2026🎉!

Earlier news

Attended EMNLP 2025 in person for the first time — a truly exciting experience 🎉

Started my Ph.D. study at the NLP Group @ PolyU & EIT NLP, supervised by Dr. Xiaoyu Shen and Prof. Wenjie Li.

Got one paper accepted by EMNLP 2025🎉!

Released our new survey on Latent Chain-of-Thought Reasoning.

Got one paper accepted by ACL 2025🎉!

Got one paper accepted by ICML 2025🎉!

Got one paper accepted by EMNLP 2024🎉!

Publications

Most recent publications on Google Scholar.
* indicates equal contribution

Conference & Journal Papers

EMNLP2026 · Main

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy

Haozhe Hu, Hao Wu, Anhao Zhao, Longwei Ding, Peiran Yin, Yunpu Ma, Xiaoyu Shen

EMNLP2026 · Main

PowerOPD: Stabilizing On-Policy Distillation with Bounded Power Transformation

Anhao Zhao, Junlong Tong, Yingqi Fan, Ping Nie, Wenjie Li, Xiaoyu Shen

EMNLP2026 · Main

Escaping the KL Agreement Trap in On-Policy Distillation

Haoran Xin, Anhao Zhao, Ying Sun, Jin Li, Xiaoyu Shen, Hui Xiong

EMNLP2026 · Findings

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration

Shuhao Li, Guodong Du, Anhao Zhao, Wanyu Lin, Tianyu Yuan, Xiaoyu Shen

EMNLP2026 · Findings

Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning

Xinghao Chen*, Anhao Zhao*, Heming Xia, Xuan Lu, Hanlin Wang, Yanjun Chen, Wei Zhang, Jian Wang†, Wenjie Li, Xiaoyu Shen†

CVPR2026

What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models

Yingqi Fan, Junlong Tong, Anhao Zhao, Xiaoyu Shen†

ICLR2026

StreamingThinker: Large Language Models Can Think While Reading

Junlong Tong, Yingqi Fan, Anhao Zhao, Yunpu Ma, Xiaoyu Shen†

EMNLP2025 · Main

VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs

Yingqi Fan, Anhao Zhao, Jinlan Fu, Junlong Tong, Hui Su, Yijie Pan, Wei Zhang, Xiaoyu Shen†

ACL2025 · Findings

LLM as Effective Streaming Processor: Bridging Streaming-Batch Mismatches with Group Position Encoding

Junlong Tong, Jinlan Fu, Zixuan Lin, Yingqi Fan, Anhao Zhao, Hui Su, Xiaoyu Shen†

ICML2025

SkipGPT: Each Token is One of a Kind

Anhao Zhao, Fanghua Ye†, Yingqi Fan, Junlong Tong, Jing Xiong, Zhiwei Fei, Hui Su, Xiaoyu Shen†

EMNLP2024 · Main

Unveiling In-Context Learning: A Coordinate System to Understand Its Working Mechanism

Anhao Zhao, Fanghua Ye, Jinlan Fu, Xiaoyu Shen†

KBS2024

A dynamic multi-modal deep reinforcement learning framework for 3D bin packing problem

Anhao Zhao, Tianrui Li, Andrew Lim

ArXiv Preprints

arXiv2026

AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimization

Junlong Tong, Wenqi Xu, Yingqi Fan, Anhao Zhao, Xuan Lu, Yang Tan, Xiaoyu Shen

arXiv2026

miniReranker: Efficient Multimodal Reranking through Visual Cache Reuse and Interaction Sparsity

Yingqi Fan, Xuan Lu, Anhao Zhao, Junlong Tong, Ping Nie, Kai Zou, Yunpu Ma, Wei Zhang, Xiaoyu Shen

arXiv2026

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models

Junlong Tong, Yao Zhang, Anhao Zhao, Yingqi Fan, Yunpu Ma, Xiaoyu Shen

arXiv2026

Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation

Anhao Zhao, H Xin, Yingqi Fan, Junlong Tong, Wenjie Li, Xiaoyu Shen

arXiv2026

SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation

Zicheng He, Anhao Zhao, Xiaoyu Shen, Chen Wu, Lei He

arXiv2026

On-Policy Supervised Fine-Tuning for Efficient Reasoning

Anhao Zhao, Ziyang Chen, Junlong Tong, Yingqi Fan, Fanghua Ye, Shuhao Li, Yunpu Ma, Wenjie Li, Xiaoyu Shen†

arXiv2026

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention

Wenjie Liu, Hao Wu, Xin Qiu, Yingqi Fan, Yihan Zhang, Anhao Zhao, Yunpu Ma, Xiaoyu Shen†

arXiv2026

From LLMs to LRMs: Rethinking Pruning for Reasoning-Centric Models

Longwei Ding, Anhao Zhao, Fanghua Ye, Ziyang Chen, Xiaoyu Shen†

Service

Reviewer/Program Committee Member:
ICLR27, ICLR26, CVPR26, ECCV26, ICML26, NeurIPS26

Teaching Assistant:
COMP1010_26271_C: Computational Thinking and Problem Solving
COMP 5311: Internet Infrastructure and Protocols, Fall 2025, PolyU
COMP 5532: DIGITAL TWINS & VIRTUAL HUMAN, Spring 2026, PolyU