Hi👋 nice to meet you!
I am a final-year Ph.D. candidate in the joint Ph.D. program between Shanghai AI Laboratory and Beihang University, supervised by Prof. Dahua Lin. At Shanghai AI Laboratory, I work closely with Jiaqi Wang and Yuhang Zang on multimodal LLMs. Before starting my Ph.D., I received my Bachelor’s degree from Beihang University in 2021 and completed two years of master’s study there under the supervision of Prof. Bowen Du, working on representation learning and forecasting for multivariate time series.

I am currently a Research Intern with Alibaba's Qwen Omni team and a core contributor to Video-to-Skill and Qwen-Music. My work spans agentic MLLMs, audio-visual understanding, and audio generation.

I am on the job market and seeking Research Scientist opportunities that build on my experience in agentic MLLMs, audio-visual understanding, and audio generation. Beyond these core areas, I am also excited about embodied intelligence, interactive multimodal systems, and video generation, and am eager to explore these emerging directions.
I am always open to research discussions and collaborations.
Email:  liuzihan@buaa.edu.cn   WeChat:  ZinniaL19

Current Focus

Agentic MLLMs & Computer-Use Agents Developing multimodal agents that acquire reusable skills from human demonstrations and instructional videos, together with scalable rollout pipelines for agentic training data.
Audio-Visual Understanding Advancing audio-visual instruction following and reasoning across temporal, spatial, and cross-modal evidence, with an emphasis on reliable evaluation.
Audio Generation Systems Developing music generation systems through tokenizer and model training, controllable song generation and editing, and post-training for Qwen-Music.

🔥 News

  • 2026.07:  Qwen-Music Technical Report is released on arXiv.
  • 2026:  STAR-Bench is accepted by ICLR 2026.
  • 2025.05:  SongGen is accepted by ICML 2025.
  • 2025.05:  SongComposer is accepted by ACL 2025 main conference.

💼 Internship

Alibaba Cloud Computing Co., Ltd. | Qwen Omni Team, Tongyi Lab
Research Intern, 2026.01 - Present

  • Video-to-Skill (Core Contributor): I distill reusable and transferable multimodal skills from human demonstrations and instructional videos to guide Computer Use Agents in completing complex customized or out-of-distribution (OOD) tasks. I build the Omni Skill Creator Plugin, spanning 17 categories of office and professional software, and generate high-quality agentic trajectory data for Omni model training and CUA capability enhancement.
  • Qwen-Music (Core Contributor) [Technical Report]: I contribute to the development of Qwen-Music across model training, evaluation, and post-training, with a focus on improving overall generation quality.

📝 Publications

* equal contribution.

Audio Understanding and Reasoning

ICLR 2026
STAR-Bench

STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence [ICLR 2026]

Zihan Liu*, Zhikang Niu*, Qiuyang Xiao, Zhisheng Zheng, Ruoqi Yuan, Yuhang Zang, Yuhang Cao, Xiaoyi Dong, Jianze Liang, Xie Chen, Leilei Sun, Dahua Lin, Jiaqi Wang

We formalize audio 4D intelligence as reasoning over sound dynamics across time and 3D space, and introduce STAR-Bench to evaluate fine-grained perceptual and spatio-temporal reasoning beyond caption-level semantics.

Homepage| Github | arXiv|

Audio and Music Generation

arXiv 2026
Qwen-Music inference framework

Qwen-Music Technical Report [arXiv 2026]

Core Contributors (alphabetical): Jin Xu, Kangdi Wang, Ruibin Yuan, Shun Lei, Xiong Wang, Xize Cheng, Xueyao Zhang, Yang Zhang, Yiheng Chen, Yongqi Wang, Yue Wang, Zhifang Guo, Zhiyong Wu, Zihan Liu, Zijian Lin

A large-scale music generation system supporting text-to-music and cover-song generation, with tokenizer/model/render components and post-training for musicality and instruction following.

arXiv|

Submitted
SongGen-X Mixture-of-Adapters framework

SongGen-X: Unifying Versatile Editing for Autoregressive Song Generation via Mixture-of-Adapters [Submitted to TASLP]

Zihan Liu, Ruixing Zhang, Jiaqi Wang, Leilei Sun, Dahua Lin, Yuhang Zang

A unified song editing framework based on Mixture-of-Adapters, supporting fixed- and adaptive-duration inpainting, track-conditioned refinement, style transfer, and lyric editing with a frozen autoregressive backbone.

Homepage| Github

ICML 2025
SongGen

SongGen: A Single-Stage Auto-regressive Transformer for Text-to-Song Generation [ICML 2025]

Zihan Liu, Shuangrui Ding, Zhixiong Zhang, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang

A fully open-source single-stage autoregressive transformer for controllable song generation, supporting lyric/text control, optional reference voice, and mixed or dual-track output modes.

Homepage| Github | arXiv|

ACL 2025
SongComposer

SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition [ACL Main 2025]

Shuangrui Ding*, Zihan Liu*, Xiaoyi Dong, Pan Zhang, Rui Qian, Junhao Huang, Conghui He, Dahua Lin, Jiaqi Wang

A language model that unifies lyric and melody generation in symbolic song representation, enabling multi-task song composition within a single framework.

Homepage| Github | arXiv|

Time Series Modeling

IJCAI 2024
CTRL

An NCDE-based Framework for Universal Representation Learning of Time Series [IJCAI 2024]

Zihan Liu, Bowen Du, Junchen Ye, Xianqing Wen, Leilei Sun

An NCDE-based framework for learning universal time-series representations through reconstruction and contrastive self-supervision across downstream tasks.

Github |

KDD 2022
ESG

Learning the Evolutionary and Multi-scale Graph Structure for Multivariate Time Series Forecasting [KDD 2022]

Junchen Ye*, Zihan Liu*, Bowen Du, Leilei Sun, Weimiao Li, Yanjie Fu, Hui Xiong

An evolutionary and multi-scale graph learning framework that models dynamic dependencies among multivariate time series.

Github |

KBS 2022
Ada-STNet

Adaptive Spatio-Temporal Graph Neural Network for Traffic Forecasting [Knowledge-Based Systems 2022]

Xuxiang Ta, Zihan Liu, Xiao Hu, Le Yu, Leilei Sun, Bowen Du

A dynamic traffic graph learning framework with macro-level self-learning and micro-level self-adaptation for traffic forecasting.

Github |

🎖 Honors and Awards

  • 2021-2025, 1st Prize, Academic Outstanding Scholarship.
  • 2022.10, National Scholarship, Ministry of Education of PRC.
  • 2022.12, Outstanding Graduate Student.
  • 2021.09, Graduate Entrance Scholarship.
  • 2021.06, Excellent Bachelor’s Thesis; Outstanding Undergraduate Graduate.

📖 Education

  • 2023.09 - 2027.03, Ph.D. Candidate, Joint Ph.D. Program between Shanghai AI Laboratory and Beihang University.
  • 2021.09 - 2023.06, M.Sc. in Computer Science and Technology, Beihang University.
  • 2017.09 - 2021.06, B.Sc. in Computer Science and Technology, Beihang University.

🖥️ Services

  • Selected conference reviewing service: ICLR, NeurIPS, ICML, and others.