Hello! I am currently a Ph.D. student in Computer Science at Xi’an Jiaotong University (XJTU), working under the supervision of Prof. Weizhan Zhang in Academician Qinghua Zheng’s lab. My research focuses on:
My current research interests lie in:
- Multimodal Large Models – Cross-Modal Representation, Vision-Language Alignment
- Computer Vision: Object Detection, Large Vision Model, AI-Generated Content
- Machine Learning: Multi-modal Learning, Zero-shot Learning
🔥 News
- 2026.03: Our paper VLDUS: Vision-Language Distillated Unseen Synthesizer for Zero-Shot Object Detection has been accepted by Neural Networks (NN).
📝 Researches & Projects

VLDUS: Vision-Language Distillated Unseen Synthesizer for Zero-Shot Object Detection
-
Muyan Jiao+, Caixia Yan+, Nuohan Xue, et al. (+ means equal contribution)
-
Neural Networks

Content-Aware Player Identification with Multi-Task Fusion for Soccer Matches
- Joint Laboratory of 5G Future Media and Artificial Intelligence, XJTU & MiGu
- Deploy during the live stream of EURO 2024
🎖 Honors and Awards
- 2023.10, The First Prize Scholarship of Xi’an Jiaotong University.
- 2023.08, Third Place in Task 2: Zero-Shot Object Detection in Image Challenge at ICCV 2023 Workshop on “Vision Meets Drones: A Challenge”
- 2022.10, Outstanding Freshman Scholarship (Grade 1).
- 2021.10, First Prize in 32nd Innovation Track of XJTU Tengfei Cup (<5%).
📖 Educations
- 2025.09 - now, Ph.D. in Computer Science and Technology, XJTU
- 2022.09 - 2025.06, M.S. in Computer Science and Technology, XJTU
- 2018.09 - 2022.06, B.E. in Computer Science and Technology, XJTU.
- 2016.09 - 2018.06, XJTU Special Class for Gifted Young.
🎙 Skills
- Programming: Python, C/C++, PyTorch, Git, Linux
- Language: English (CET-6, IELTS)
- Foundation Models: LLaMA, Qwen, DeepSeek; post-training methods including SFT, RLHF, DPO, and GRPO
- Fine-tuning: LoRA, LLaMA Factory, LLaVA
- Music: Guitar, Piano