Qining Zhang (张启宁)

Qining Zhang 

Ph.D. Candidate
Department of Electrical Engineering and Computer Science
University of Michigan, Ann Arbor
Address: 1301 Beal Ave, Ann Arbor, MI, USA, 48105
E-mail: qiningz AT umich Dot edu

About me

Hi! I am a final-year Ph.D. candidate in Electrical Engineering and Computer Science at the University of Michigan, advised by Prof. Lei Ying, and a recipient of the Rackham Predoctoral Fellowship. I received an M.S. in Mathematics from the University of Michigan and a B.E. degree in Electrical Engineering from Tsinghua University in 2021, where I was mentored by Prof. Jintao Wang and Prof. Haoyue Tang. In summer 2024, I was an Applied Scientist Intern with the Amazon Search Experience Science team, mentored by Dr. Yi Liu and Dr. Tanner Fiez.

I am currently on the academic job market for tenure-track Assistant Professor or Postdoc positions in IE/OR, ECE, and CS, starting Fall 2027.

News

  • Oct. - Nov. 2026: I will be visiting several universities this fall to give talks on stochastic zeroth-order policy optimization for reinforcement learning.

    • Oct. 9, 2026: Computer Science Department at Carnegie Mellon University. Many thanks to Prof. Weina Wang for hosting!

    • Oct. 12, 2026: ECE Department at University of Virginia. Many thanks to Prof. Cong Shen for hosting!

    • Oct. 15, 2026: CS Department at Virginia Tech. Many thanks to Prof. Bo Ji for hosting!

    • Oct. 20, 2026: School of EECS at Penn State University. Many thanks to Prof. Bin Li for hosting!

    • Oct. 22, 2026: ECE Department at Johns Hopkins University. Many thanks to Prof. Laixi Shi for hosting!

    • Nov. 5, 2026: ECE Department at University of California Davis. Many thanks to Prof. Junshan Zhang for hosting!

  • Nov. 1, 2026: I will attend the INFORMS Annual Meeting and present my job-market talk in the SA26 - Queues, Learning, and Applied Probability session.

Research

My research develops foundations and algorithms for reinforcement learning with human feedback, resource constraints, and representation bottlenecks, using stochastic zeroth-order policy optimization (SZPO) as a unified approach. My broader interests include:

  • Reinforcement learning and optimization for human alignment;

  • Stochastic bandits, best-arm identification, and adaptive experimentation;

  • Stochastic systems and control, including queueing, networked systems, and platform operations;

  • Applications to language model post-training, robotics, and autonomous systems.

Publications

Preprints

  1. Actor-Predictor Architecture for Reinforcement Learning with Formal Specification Objectives
    Josef Brozovich, Qining Zhang, Lei Ying.
    Under Review, 2026.

  2. Efficient Federated Reinforcement Learning from Human Feedback via Zeroth-Order Policy Optimization
    Deyi Wang*, Qining Zhang*, Lei Ying.
    Under Review, Equal contribution, 2026.

  3. SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling
    Evan Assmus, Qining Zhang, Lei Ying.
    Under Review, 2026.

Journals and Journal-Quality Conferences

  1. Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function
    Qining Zhang, Lei Ying.
    Reinforcement Learning Conference (RLC) and Reinforcement Learning Journal (RLJ), 2026. (Acceptance Rate: 33.9%)
    A short version was accepted at the MLxOR Workshop of NeurIPS, 2025.

  2. Multi-Metric Adaptive Experimental Design Under a Fixed Budget with Validation
    Qining Zhang, Tanner Fiez, Yi Liu, Wenyang Liu.
    International Conference on Artificial Intelligence and Statistics (AISTATS), 2026. (Acceptance Rate: 28.1%)

  3. Fingerprinting and Quantification of Procyanidins via LC-MS/MS and ESI In-Source Fragmentation
    Yanxin Lin, Helene Hopfer, Qining Zhang, Misha T. Kwasniewski.
    Journal of Agricultural and Food Chemistry, 2025.

  4. Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
    Qining Zhang, Lei Ying.
    International Conference on Learning Representations (ICLR), 2025. (Acceptance Rate: 32.1%)

  5. Reinforcement Learning from Human Feedback without Reward Inference: Model-Free Algorithm and Instance-Dependent Analysis
    Qining Zhang, Honghao Wei, Lei Ying.
    Reinforcement Learning Conference (RLC) and Reinforcement Learning Journal (RLJ), 2024. (Acceptance Rate: ~40%)

  6. Cost Aware Best Arm Identification
    Kellen Kanarios, Qining Zhang, Lei Ying.
    Reinforcement Learning Conference (RLC) and Reinforcement Learning Journal (RLJ), 2024. (Acceptance Rate: ~40%)

  7. Deep Reinforcement Learning for Early Diagnosis of Lung Cancer
    Yifan Wang, Qining Zhang, Lei Ying, Chuan Zhou.
    AAAI Conference on Artificial Intelligence (AAAI), 2024. (Acceptance Rate: 24.2%)

  8. Fast and Regret Optimal Best Arm Identification: Fundamental Limits and Low-Complexity Algorithms
    Qining Zhang, Lei Ying.
    Advances in Neural Information Processing Systems (NeurIPS), 2023. (Acceptance Rate: 26.1%)

  9. On Low-Complexity Quickest Intervention of Mutated Diffusion Processes Through Local Approximation
    Qining Zhang, Honghao Wei, Weina Wang, Lei Ying.
    ACM International Symposium on Theory, Algorithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing (MobiHoc), 2022. (Acceptance Rate: 19.8%)

  10. Online Utility Optimization in Multi-User Interference Networks Under a Long-Term Budget Constraint
    Yuchao Chen, Jintao Wang, Qining Zhang, Feifei Gao, Jian Song.
    IEEE Transactions on Vehicular Technology, 2022.

Other Publications

  1. Early Lung Cancer Diagnosis from Virtual Follow-up LDCT Generation via Correlational Autoencoder and Latent Flow Matching
    Yutong Wu, Yifan Wang, Qining Zhang, Chuan Zhou, Lei Ying.
    W3PHIAI-26 Workshop at AAAI, 2026.

  2. Online Optimizing Multi-user Interference Network Utility with Unknown CSI under Budget Constraint
    Yuchao Chen, Jintao Wang, Qining Zhang, Feifei Gao, Jian Song.
    IEEE Wireless Communications and Networking Conference (WCNC), 2022.

  3. Minimizing the Age of Synchronization in Power-Constrained Wireless Networks with Unreliable Time-Varying Channels
    Qining Zhang, Haoyue Tang, Jintao Wang.
    Age of Information Workshop, IEEE INFOCOM Workshops, 2020.

Professional Services

Conference Reviewer: ICLR / NeurIPS / AISTATS / INFOCOM / L4DC / RLC / EWRL

Journal Reviewer: IEEE TPAMI / IEEE ToN / Automatica / Performance Evaluation / TMLR