|
Qining Zhang (张启宁)
About me
Hi! I am a final-year Ph.D. candidate in Electrical Engineering and Computer Science at the University of Michigan, advised by Prof.
Lei Ying, and a recipient of the
Rackham Predoctoral Fellowship.
I received an M.S. in Mathematics from the University of Michigan and a B.E. degree in Electrical Engineering from Tsinghua University in 2021, where I was mentored by Prof.
Jintao Wang
and Prof.
Haoyue Tang.
In summer 2024, I was an Applied Scientist Intern with the Amazon Search Experience Science team, mentored by Dr.
Yi Liu
and Dr.
Tanner Fiez.
I am currently on the academic job market for tenure-track Assistant Professor or Postdoc positions in IE/OR, ECE, and CS, starting Fall 2027.
News
-
Oct. - Nov. 2026: I will be visiting several universities this fall to give talks on stochastic zeroth-order policy optimization for reinforcement learning.
-
Oct. 9, 2026: Computer Science Department at Carnegie Mellon University. Many thanks to Prof. Weina Wang for hosting!
-
Oct. 12, 2026: ECE Department at University of Virginia. Many thanks to Prof. Cong Shen for hosting!
-
Oct. 15, 2026: CS Department at Virginia Tech. Many thanks to Prof. Bo Ji for hosting!
-
Oct. 20, 2026: School of EECS at Penn State University. Many thanks to Prof. Bin Li for hosting!
-
Oct. 22, 2026: ECE Department at Johns Hopkins University. Many thanks to Prof. Laixi Shi for hosting!
-
Nov. 5, 2026: ECE Department at University of California Davis. Many thanks to Prof. Junshan Zhang for hosting!
-
Nov. 1, 2026: I will attend the INFORMS Annual Meeting and present my job-market talk in the SA26 - Queues, Learning, and Applied Probability session.
Research
My research develops foundations and algorithms for reinforcement learning with human feedback, resource constraints, and representation bottlenecks, using stochastic zeroth-order policy optimization (SZPO) as a unified approach.
My broader interests include:
Reinforcement learning and optimization for human alignment;
Stochastic bandits, best-arm identification, and adaptive experimentation;
Stochastic systems and control, including queueing, networked systems, and platform operations;
Applications to language model post-training, robotics, and autonomous systems.
Publications
Preprints
Actor-Predictor Architecture for Reinforcement Learning with Formal Specification Objectives
Josef Brozovich, Qining Zhang, Lei Ying.
Under Review, 2026.
Efficient Federated Reinforcement Learning from Human Feedback via Zeroth-Order Policy Optimization
Deyi Wang*, Qining Zhang*, Lei Ying.
Under Review, Equal contribution, 2026.
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling
Evan Assmus, Qining Zhang, Lei Ying.
Under Review, 2026.
Journals and Journal-Quality Conferences
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function
Qining Zhang, Lei Ying.
Reinforcement Learning Conference (RLC) and Reinforcement Learning Journal (RLJ), 2026.
(Acceptance Rate: 33.9%)
A short version was accepted at the MLxOR Workshop of NeurIPS, 2025.
Multi-Metric Adaptive Experimental Design Under a Fixed Budget with Validation
Qining Zhang, Tanner Fiez, Yi Liu, Wenyang Liu.
International Conference on Artificial Intelligence and Statistics (AISTATS), 2026.
(Acceptance Rate: 28.1%)
Fingerprinting and Quantification of Procyanidins via LC-MS/MS and ESI In-Source Fragmentation
Yanxin Lin, Helene Hopfer, Qining Zhang, Misha T. Kwasniewski.
Journal of Agricultural and Food Chemistry, 2025.
Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
Qining Zhang, Lei Ying.
International Conference on Learning Representations (ICLR), 2025.
(Acceptance Rate: 32.1%)
Reinforcement Learning from Human Feedback without Reward Inference: Model-Free Algorithm and Instance-Dependent Analysis
Qining Zhang, Honghao Wei, Lei Ying.
Reinforcement Learning Conference (RLC) and Reinforcement Learning Journal (RLJ), 2024.
(Acceptance Rate: ~40%)
Cost Aware Best Arm Identification
Kellen Kanarios, Qining Zhang, Lei Ying.
Reinforcement Learning Conference (RLC) and Reinforcement Learning Journal (RLJ), 2024.
(Acceptance Rate: ~40%)
Deep Reinforcement Learning for Early Diagnosis of Lung Cancer
Yifan Wang, Qining Zhang, Lei Ying, Chuan Zhou.
AAAI Conference on Artificial Intelligence (AAAI), 2024.
(Acceptance Rate: 24.2%)
Fast and Regret Optimal Best Arm Identification: Fundamental Limits and Low-Complexity Algorithms
Qining Zhang, Lei Ying.
Advances in Neural Information Processing Systems (NeurIPS), 2023.
(Acceptance Rate: 26.1%)
On Low-Complexity Quickest Intervention of Mutated Diffusion Processes Through Local Approximation
Qining Zhang, Honghao Wei, Weina Wang, Lei Ying.
ACM International Symposium on Theory, Algorithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing (MobiHoc), 2022.
(Acceptance Rate: 19.8%)
Online Utility Optimization in Multi-User Interference Networks Under a Long-Term Budget Constraint
Yuchao Chen, Jintao Wang, Qining Zhang, Feifei Gao, Jian Song.
IEEE Transactions on Vehicular Technology, 2022.
Other Publications
Early Lung Cancer Diagnosis from Virtual Follow-up LDCT Generation via Correlational Autoencoder and Latent Flow Matching
Yutong Wu, Yifan Wang, Qining Zhang, Chuan Zhou, Lei Ying.
W3PHIAI-26 Workshop at AAAI, 2026.
Online Optimizing Multi-user Interference Network Utility with Unknown CSI under Budget Constraint
Yuchao Chen, Jintao Wang, Qining Zhang, Feifei Gao, Jian Song.
IEEE Wireless Communications and Networking Conference (WCNC), 2022.
Minimizing the Age of Synchronization in Power-Constrained Wireless Networks with Unreliable Time-Varying Channels
Qining Zhang, Haoyue Tang, Jintao Wang.
Age of Information Workshop, IEEE INFOCOM Workshops, 2020.
Professional Services
Conference Reviewer: ICLR / NeurIPS / AISTATS / INFOCOM / L4DC / RLC / EWRL
Journal Reviewer: IEEE TPAMI / IEEE ToN / Automatica / Performance Evaluation / TMLR
|