Hi! I’m Yifei Xia, a second-year Ph.D. student in the School of Computer Science at Peking University and a member of DAIR Lab, advised by Prof. Bin Cui. I received my B.Sc. degree in Computer Science and Technology from the Turing Experimental Class at Renmin University of China in June 2024, where I was advised by Prof. Feng Zhang.
I am the creator of HetuDiT, an efficient and dynamic serving system for image and video generation models, and a developer of Hetu, a high-performance distributed deep learning system for large-scale and automated parallel training, including LLM training workloads.
During my internship at ByteDance, I have been deeply involved in post-training and inference acceleration for the Seedance 1.0 and Seedance 2.0 series. Several of my designs, including AdaSpa, have been deployed in production and widely used across business scenarios.
My research interests lie in AI infrastructure, especially DiT post-train, inference and serving, LLM inference and serving, and multi-modal training/inference systems.
I am open to collaborations and discussions on efficient inference, serving systems, and AI infrastructure. Feel free to reach out!
📖 Educations
B.S. in Computer Science and Technology, Turing Experimental Class, Renmin University of China. 2020 - 2024. Adviser: Feng Zhang
Ph.D. in Computer Science and Technology, Peking University. 2024 - present. Adviser: Bin Cui
💻 Internships
- [2024.09 - present] Research Intern, Seed CV-AI Platform, ByteDance.
- Seedance 1.0 post-training inference acceleration: Designed and implemented the sparse attention solution for Seedance 1.0, enabling efficient long-video generation inference in production.
- Seedance 2.0 post-training inference acceleration: Contributed to post-training and inference acceleration for Seedance 2.0, including sparse attention and quantization-based optimizations.
📝 Publications
2026
[arXiv] Yifei Xia, Hao Yuan, Suhan Ling, Haoran Sun, Hanke Zhang, Xupeng Miao, Fangcheng Fu, Bin Cui. “Adaptive Resource Management and Quality Control for Streaming Video Generation.”
[ICML] Yifei Xia, Fangcheng Fu, Hao Yuan, Suhan Ling, Xupeng Miao, Huixia Li, Yuxi Ren, Xin Xia, Xuefeng Xiao, Bin Cui. “EchoAttention: Exploiting Token-Pair Redundancy and Frame-Block Similarity for Efficient Long Video Generation.”
2025
[arXiv] Yifei Xia, Fangcheng Fu, Hao Yuan, Hanke Zhang, Xupeng Miao, Yijun Liu, Suhan Ling, Jie Jiang, Bin Cui. “TridentServe: A Stage-level Serving System for Diffusion Pipelines.”
[ICCV] Yifei Xia, Suhan Ling, Fangcheng Fu, Yujie Wang, Huixia Li, Xuefeng Xiao, Bin Cui. “Training-free and Adaptive Sparse Attention for Efficient Long Video Generation.”
[KDD] Zheng Chen, Feng Zhang, Yifei Xia, Wentao Zhang, Xiaowei Zhu, Wenguang Chen, Xiaoyong Du. “CompressGNN: Accelerating Graph Neural Network Training via Hierarchical Compression.”
2024
[NeurIPS] Yifei Xia, Fangcheng Fu, Wentao Zhang, Jiawei Jiang, Bin Cui. “Efficient Multi-task LLM Quantization and Serving for Multiple LoRA Adapters.”
[VLDBJ] Yifei Xia, Feng Zhang, Qingyu Xu, Mingde Zhang, Zhiming Yao, Lv Lu, Xiaoyong Du, Dong Deng, Bingsheng He, Siqi Ma. “GPU-Based Butterfly Counting.”
2023
- [TPDS] Yihua Hu, Feng Zhang, Yifei Xia, Zhiming Yao, Letian Zeng, Haipeng Ding, Zhewei Wei, Xiao Zhang, Jidong Zhai, Xiaoyong Du, Siqi Ma. “Enabling Efficient Random Access to Hierarchically Compressed Text Data on Diverse GPU Platforms.”
🚀 Systems

HetuDiT - A Dynamic Serving System for Diffusion Transformers
HetuDiT is a stage-level, dynamically parallel serving system for Diffusion Transformers (DiTs). It targets real-world image and video generation serving workloads where requests differ in prompt length, resolution, frame count, and stage-level compute demand.
- Stage-level scheduling: Profiles condition encoding, denoising, VAE decoding, and other stages separately, then rewrites execution plans on the fly.
- Dynamic parallelism: Assigns stage-specific batch sizes and parallelism strategies, including SP, CP, TP, and PP, for each request stage.
- Model support: Provides serving scripts for Stable Diffusion 3, CogVideoX 1.5-5B, Flux, and HunyuanVideo.
- Serving performance: README benchmarks report higher SLO attainment and lower mean/P95 latency than static pipeline-level serving on a 128x NVIDIA L20 cluster.

Hetu - A Distributed Deep Learning System
Developer | GitHub | Documentation
Hetu is a high-performance distributed deep learning system developed by DAIR Lab at Peking University. It targets large-scale and automated distributed training, including trillion-parameter model training and modern LLM-related workloads.
- Parallel training: Supports multiple training protocols and communication architectures, including data/model/pipeline parallelism, parameter server, and AllReduce.
- Large-scale workloads: Designed for giant models and large-scale training workloads across distributed GPU clusters.
- Research coverage: Serves as a platform for system research across Transformer/LLM, MoE, embedding, diffusion, GNN, heterogeneous-resource, kernel, and memory-management workloads.
- Open ecosystem: Provides documentation, examples, and an Apache-2.0 open-source codebase maintained under PKU-DAIR.
🎖 Honors and Awards
Peking University Presidential Scholarship, 2025 (top 1.5%, highest honor for Ph.D. students at Peking University)
Peking University Presidential Scholarship, 2024 (top 1.5%, highest honor for Ph.D. students at Peking University)
China Institute of Electronics - Tencent Ph.D. Research Incentive Program, Hunyuan Foundation Model Special Track, 2024 (100,000 RMB, only 17 recipients nationwide)
Best Paper Award at CCF National Database Conference NDBC, 2024
Outstanding Graduate, Renmin University of China, 2024
Outstanding Undergraduate Thesis, Renmin University of China, 2024
PetroChina Scholarship, 2023
National Scholarship, 2022 (top 1%)
RUC-Huawei “Intelligent Base” Student Scholarship, 2022 and 2023
