Shuang Qiu received his Ph.D. degree in Computer Science and Engineering from the University of Michigan, Ann Arbor. He is now an assistant professor at Department of Systems Engineering and Department of Computer Sciences, City University of Hong Kong. His primary research interest include agentic AI, embodied AI, reinforcement learning, optimization, and AI for engineering / science / industry. For more details, please refer to his personal homepage.
Service in CityUHK
Teaching Service
- 2026 - Now, Master's Core Course (MSc in AI-Driven Innovation), SYE6601 - Introduction to Artificial Intelligence: Concepts and Applications.
- 2025 - Now, PhD Core Course, SYE8203 - Applied Probability and Statistics.
Selected Publications
- For more details and full publications, please refer to his personal homepage and google scholar.
- Zhongjian Qiao, Jiafei Lyu, Chenjia Bai, Peisong Wang, Siyang Gao, Shuang Qiu. Unifying Value Alignment and Assignment in Cross-Domain Offline Reinforcement Learning with Heterogeneous Datasets. International Conference on Machine Learning (ICML), 2026
- Zhongjian Qiao, Rui Yang, Jiafei Lyu, Xiu Li, Zhongxiang Dai, Zhuoran Yang, Siyang Gao, Shuang Qiu. Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics Shifts. International Conference on Learning Representations (ICLR), 2026
- Zhongjian Qiao, Jiafei Lyu, Boxiang Lyu, Yao Shu, Siyang Gao, Shuang Qiu. Model-based Offline RL via Robust Value-Aware Model Learning with Implicitly Differentiable Adaptive Weighting. International Conference on Learning Representations (ICLR), 2026
- Yiran Guo, Lijie Xu, Jie Liu, Dan Ye, Shuang Qiu. Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models. Advances in Neural Information Processing Systems (NeurIPS), 2025
- Xize Liang*, Chao Chen*, Shuang Qiu*, Jie Wang, Yue Wu, Zhihang Fu, Zhihao Shi, Feng Wu, Jieping Ye. ROPO: Robust Preference Optimization for Large Language Models. International Conference on Machine Learning (ICML), 2025
- Chenjia Bai, Yang Zhang, Shuang Qiu, Qiaosheng Zhang, Kang Xu, Xuelong Li. Online Preference Alignment for Language Models via Count-based Exploration. International Conference on Learning Representations (ICLR Spotlight), 2025
- Shuang Qiu*, Boxiang Lyu*, Qinglin Meng*, Zhaoran Wang, Zhuoran Yang, Michael I. Jordan. Learning Dynamic Mechanisms in Unknown Environments: A Reinforcement Learning Approach. Journal of Machine Learning Research (JMLR), 2024
- Rui Yang*, Xiaoman Pan*, Feng Luo*, Shuang Qiu*, Han Zhong, Dong Yu, Jianshu Chen. Rewards-in-Context: Multi-Objective Alignment of Foundation Models with Dynamic Preference Adjustment. International Conference on Machine Learning (ICML), 2024
- Dake Zhang, Boxiang Lyu, Shuang Qiu#, Mladen Kolar, Tong Zhang. Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning. International Conference on Machine Learning (ICML Spotlight), 2024
- Shuang Qiu*, Ziyu Dai*, Han Zhong, Zhaoran Wang, Zhuoran Yang, Tong Zhang. Posterior Sampling for Competitive RL: Function Approximation and Partial Observation. Advances in Neural Information Processing Systems (NeurIPS), 2023
- Shuang Qiu, Lingxiao Wang, Chenjia Bai, Zhuoran Yang, Zhaoran Wang. Contrastive UCB: Provably Efficient Contrastive Self-Supervised Learning in Online Reinforcement Learning. International Conference on Machine Learning (ICML), 2022
- Shuang Qiu, Xiaohan Wei, Jieping Ye, Zhaoran Wang, Zhuoran Yang. Provably Efficient Fictitious Play Policy Optimization for Zero-Sum Markov Games with Structured Transitions. International Conference on Machine Learning (ICML), 2021
- Shuang Qiu, Xiaohan Wei, Zhuoran Yang, Jieping Ye, Zhaoran Wang. Upper Confidence Primal-Dual Reinforcement Learning for CMDP with Adversarial Loss. Advances in Neural Information Processing Systems (NeurIPS), 2020
Postdoc & RA & PhD Openings
- I am actively seeking self-motivated students with strong mathematical or programming skills for the following positions:
-
Postdoc & Research Assistant: Multiple Postdoc and long-term / short-term / remote RA positions available now
-
Agentic AI
-
World Model & Embodied AI
-
Large Language Model
-
Reinforcement Learning
-
AI for Engineering / Science / Industry
Please do not hesitate to contact me via my email with your CV and transcript if you are interested in the above research topics.
Last update date :
15 Aug 2026