VLM, LLM, and speech
Multimodal and language models for real-time understanding of live content.
I am an algorithm engineer at TikTok LIVE, working on VLM, LLM, and speech models for real-time understanding of large-scale live content.
My interests include long-context modeling, efficient AI through quantization, distillation and hardware-software co-design, and multilingual and multimodal learning.
The common thread across my work is connecting model research with the latency, throughput, and reliability requirements of production-scale live systems.
Multimodal and language models for real-time understanding of live content.
Modeling evolving streams with long temporal context and cross-modal evidence.
Quantization, distillation, and hardware-software co-design for efficient inference.
Learning across languages and modalities for globally distributed live content.
My path combines academic research with product-facing algorithm work across media, vision, and multimodal intelligence.
Algorithm Engineer working on VLM, LLM, and speech models for real-time, large-scale live-content understanding.
Senior Computer Vision Engineer focused on computer vision and AI algorithm and architecture R&D for live video.
Staff Researcher working on computer vision algorithms and R&D platforms for industry and scientific research.
Ph.D. in Computer Science (HKBU, 2018), B.Eng. in Computer Science and Technology (SCUT, 2013), and a visiting Ph.D. appointment at Michigan State University (Feb–Aug 2016).
For a more complete record of publications and activity, these public profiles are the best entry points.