Hello! I am currently a Ph.D. student in Computer Science at Xi’an Jiaotong University (XJTU), working under the supervision of Prof. Weizhan Zhang in Academician Qinghua Zheng’s lab. My research focuses on:

My current research interests lie in:

  • Multimodal Large Models – Cross-Modal Representation, Vision-Language Alignment
  • Computer Vision: Object Detection, Large Vision Model, AI-Generated Content
  • Machine Learning: Multi-modal Learning, Zero-shot Learning

🔥 News

  • 2026.03: Our paper VLDUS: Vision-Language Distillated Unseen Synthesizer for Zero-Shot Object Detection has been accepted by Neural Networks (NN).

📝 Researches & Projects

sym

VLDUS: Vision-Language Distillated Unseen Synthesizer for Zero-Shot Object Detection

  • Muyan Jiao+, Caixia Yan+, Nuohan Xue, et al. (+ means equal contribution)

  • Neural Networks

sym

Content-Aware Player Identification with Multi-Task Fusion for Soccer Matches

  • Joint Laboratory of 5G Future Media and Artificial Intelligence, XJTU & MiGu
  • Deploy during the live stream of EURO 2024

🎖 Honors and Awards

  • 2023.10, The First Prize Scholarship of Xi’an Jiaotong University.
  • 2023.08, Third Place in Task 2: Zero-Shot Object Detection in Image Challenge at ICCV 2023 Workshop on “Vision Meets Drones: A Challenge”
  • 2022.10, Outstanding Freshman Scholarship (Grade 1).
  • 2021.10, First Prize in 32nd Innovation Track of XJTU Tengfei Cup (<5%).

📖 Educations

  • 2025.09 - now, Ph.D. in Computer Science and Technology, XJTU
  • 2022.09 - 2025.06, M.S. in Computer Science and Technology, XJTU
  • 2018.09 - 2022.06, B.E. in Computer Science and Technology, XJTU.
  • 2016.09 - 2018.06, XJTU Special Class for Gifted Young.

🎙 Skills

  • Programming: Python, C/C++, PyTorch, Git, Linux
  • Language: English (CET-6, IELTS)
  • Foundation Models: LLaMA, Qwen, DeepSeek; post-training methods including SFT, RLHF, DPO, and GRPO
  • Fine-tuning: LoRA, LLaMA Factory, LLaVA
  • Music: Guitar, Piano