I am a final-year Ph.D. candidate in the joint Ph.D. program between Shanghai AI Laboratory and Beihang University, supervised by Prof. Dahua Lin. At Shanghai AI Laboratory, I work closely with Jiaqi Wang and Yuhang Zang on multimodal LLMs. Before starting my Ph.D., I received my Bachelor’s degree from Beihang University in 2021 and completed two years of master’s study there under the supervision of Prof. Bowen Du, working on representation learning and forecasting for multivariate time series.
I am currently a Research Intern with Alibaba's Qwen Omni team and a core contributor to Video-to-Skill and Qwen-Music. My work spans agentic MLLMs, audio-visual understanding, and audio generation.
I am on the job market and seeking Research Scientist opportunities that build on my experience in agentic MLLMs, audio-visual understanding, and audio generation. Beyond these core areas, I am also excited about embodied intelligence, interactive multimodal systems, and video generation, and am eager to explore these emerging directions.
I am always open to research discussions and collaborations.
Email: liuzihan@buaa.edu.cn WeChat: ZinniaL19
Current Focus
🔥 News
- 2026.07: Qwen-Music Technical Report is released on arXiv.
- 2026: STAR-Bench is accepted by ICLR 2026.
- 2025.05: SongGen is accepted by ICML 2025.
- 2025.05: SongComposer is accepted by ACL 2025 main conference.
💼 Internship
Alibaba Cloud Computing Co., Ltd. | Qwen Omni Team, Tongyi Lab
Research Intern, 2026.01 - Present
- Video-to-Skill (Core Contributor): I distill reusable and transferable multimodal skills from human demonstrations and instructional videos to guide Computer Use Agents in completing complex customized or out-of-distribution (OOD) tasks. I build the Omni Skill Creator Plugin, spanning 17 categories of office and professional software, and generate high-quality agentic trajectory data for Omni model training and CUA capability enhancement.
- Qwen-Music (Core Contributor) [Technical Report]: I contribute to the development of Qwen-Music across model training, evaluation, and post-training, with a focus on improving overall generation quality.
📝 Publications 
* equal contribution.
Audio Understanding and Reasoning

STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence [ICLR 2026]
Zihan Liu*, Zhikang Niu*, Qiuyang Xiao, Zhisheng Zheng, Ruoqi Yuan, Yuhang Zang, Yuhang Cao, Xiaoyi Dong, Jianze Liang, Xie Chen, Leilei Sun, Dahua Lin, Jiaqi Wang
We formalize audio 4D intelligence as reasoning over sound dynamics across time and 3D space, and introduce STAR-Bench to evaluate fine-grained perceptual and spatio-temporal reasoning beyond caption-level semantics.
Audio and Music Generation

Qwen-Music Technical Report [arXiv 2026]
Core Contributors (alphabetical): Jin Xu, Kangdi Wang, Ruibin Yuan, Shun Lei, Xiong Wang, Xize Cheng, Xueyao Zhang, Yang Zhang, Yiheng Chen, Yongqi Wang, Yue Wang, Zhifang Guo, Zhiyong Wu, Zihan Liu, Zijian Lin
A large-scale music generation system supporting text-to-music and cover-song generation, with tokenizer/model/render components and post-training for musicality and instruction following.

SongGen-X: Unifying Versatile Editing for Autoregressive Song Generation via Mixture-of-Adapters [Submitted to TASLP]
Zihan Liu, Ruixing Zhang, Jiaqi Wang, Leilei Sun, Dahua Lin, Yuhang Zang
A unified song editing framework based on Mixture-of-Adapters, supporting fixed- and adaptive-duration inpainting, track-conditioned refinement, style transfer, and lyric editing with a frozen autoregressive backbone.

SongGen: A Single-Stage Auto-regressive Transformer for Text-to-Song Generation [ICML 2025]
Zihan Liu, Shuangrui Ding, Zhixiong Zhang, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang
A fully open-source single-stage autoregressive transformer for controllable song generation, supporting lyric/text control, optional reference voice, and mixed or dual-track output modes.

SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition [ACL Main 2025]
Shuangrui Ding*, Zihan Liu*, Xiaoyi Dong, Pan Zhang, Rui Qian, Junhao Huang, Conghui He, Dahua Lin, Jiaqi Wang
A language model that unifies lyric and melody generation in symbolic song representation, enabling multi-task song composition within a single framework.
Time Series Modeling

An NCDE-based Framework for Universal Representation Learning of Time Series [IJCAI 2024]
Zihan Liu, Bowen Du, Junchen Ye, Xianqing Wen, Leilei Sun
An NCDE-based framework for learning universal time-series representations through reconstruction and contrastive self-supervision across downstream tasks.
Github |

Learning the Evolutionary and Multi-scale Graph Structure for Multivariate Time Series Forecasting [KDD 2022]
Junchen Ye*, Zihan Liu*, Bowen Du, Leilei Sun, Weimiao Li, Yanjie Fu, Hui Xiong
An evolutionary and multi-scale graph learning framework that models dynamic dependencies among multivariate time series.
Github |

Adaptive Spatio-Temporal Graph Neural Network for Traffic Forecasting [Knowledge-Based Systems 2022]
Xuxiang Ta, Zihan Liu, Xiao Hu, Le Yu, Leilei Sun, Bowen Du
A dynamic traffic graph learning framework with macro-level self-learning and micro-level self-adaptation for traffic forecasting.
Github |
🎖 Honors and Awards
- 2021-2025, 1st Prize, Academic Outstanding Scholarship.
- 2022.10, National Scholarship, Ministry of Education of PRC.
- 2022.12, Outstanding Graduate Student.
- 2021.09, Graduate Entrance Scholarship.
- 2021.06, Excellent Bachelor’s Thesis; Outstanding Undergraduate Graduate.
📖 Education
- 2023.09 - 2027.03, Ph.D. Candidate, Joint Ph.D. Program between Shanghai AI Laboratory and Beihang University.
- 2021.09 - 2023.06, M.Sc. in Computer Science and Technology, Beihang University.
- 2017.09 - 2021.06, B.Sc. in Computer Science and Technology, Beihang University.
🖥️ Services
- Selected conference reviewing service: ICLR, NeurIPS, ICML, and others.