Research + engineering

I am an algorithm engineer at TikTok LIVE, working on VLM, LLM, and speech models for real-time understanding of large-scale live content.

My current interests include long-context modeling, efficient AI through quantization, distillation and hardware-software co-design, and multilingual and multimodal learning. I care about research that remains effective under real-world latency, throughput, and reliability constraints.

Before TikTok, I worked on real-time edge AI and large-scale content understanding at YY Live, and on computer vision algorithms and platforms at Lenovo. I received my Ph.D. in Computer Science from Hong Kong Baptist University and was a visiting Ph.D. student at Michigan State University.

Current Focus

My work connects fundamental model research with production-scale AI for live content.

Multimodal live understanding

VLM, LLM, and speech models for real-time understanding of multilingual live content.

Long-context modeling

Modeling long, evolving streams where temporal context and cross-modal evidence both matter.

Efficient AI

Quantization, distillation, and hardware-software co-design for efficient large-scale inference.

Experience Snapshot

Roles spanning livestream AI, audio-video understanding, and applied computer vision.

Oct 2024 – Present

Algorithm Engineer, TikTok

VLM, LLM, and speech models for real-time understanding of large-scale live content.

Aug 2020 – Oct 2024

Senior Computer Vision Algorithm Engineer, YY Live (Baidu Group)

Computer vision and AI algorithm and architecture R&D for live video, including real-time edge AI and large-scale content understanding.

Dec 2018 – Aug 2020

Staff Researcher, Lenovo Machine Intelligence Center

Computer vision algorithms and R&D platforms for industry and scientific research.

2013 – 2018

Ph.D., Hong Kong Baptist University

Ph.D. in Computer Science, including a visiting Ph.D. appointment at Michigan State University in 2016.

Selected Work

Selected work across multimodal understanding, biometric security, and applied computer vision.

Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs

Haochen Wang, Yuhao Wang, Tao Zhang, Yikang Zhou, Yanwei Li, Jiacong Wang, Jiani Zheng, Ye Tian, Jiahao Meng, Zilong Huang, Guangcan Mai, Anran Wang, Yunhai Tong, Zhuochen Wang, Xiangtai Li, and Zhaoxiang Zhang. ICLR 2026 Poster

Precise region-level perception, multi-region interaction modeling, and compositional reasoning for multimodal LLMs.

SecureFace: Face Template Protection

IEEE Transactions on Information Forensics and Security, 2021

Face template protection with a focus on privacy, security, and deployable biometric systems.

On the Reconstruction of Face Images from Deep Face Templates

IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019

A widely cited study on reconstructing face images from deep templates and the security implications for biometrics.

LeapDetect: An Agile Platform for Inspecting Power Transmission Lines from Drones

IEEE ICDMW Demo, 2019

An applied computer vision system for transmission-line inspection using drone imagery.

Community

I regularly review for leading journals and conferences in computer vision, multimedia, and AI.

  • IEEE TPAMI
  • IEEE TIFS
  • IEEE TIP
  • IEEE TCYB
  • Pattern Recognition
  • IEEE/CAA Journal of Automatica Sinica
  • ICME
  • ICPR
  • ICASSP

I am interested in hearing from interns working on long-context modeling, efficient AI, and multilingual or multimodal learning. You can reach me by email.