Multimodal live understanding
VLM, LLM, and speech models for real-time understanding of multilingual live content.
I am an algorithm engineer at TikTok LIVE, working on VLM, LLM, and speech models for real-time understanding of large-scale live content.
My current interests include long-context modeling, efficient AI through quantization, distillation and hardware-software co-design, and multilingual and multimodal learning. I care about research that remains effective under real-world latency, throughput, and reliability constraints.
Before TikTok, I worked on real-time edge AI and large-scale content understanding at YY Live, and on computer vision algorithms and platforms at Lenovo. I received my Ph.D. in Computer Science from Hong Kong Baptist University and was a visiting Ph.D. student at Michigan State University.
My work connects fundamental model research with production-scale AI for live content.
VLM, LLM, and speech models for real-time understanding of multilingual live content.
Modeling long, evolving streams where temporal context and cross-modal evidence both matter.
Quantization, distillation, and hardware-software co-design for efficient large-scale inference.
Roles spanning livestream AI, audio-video understanding, and applied computer vision.
VLM, LLM, and speech models for real-time understanding of large-scale live content.
Computer vision and AI algorithm and architecture R&D for live video, including real-time edge AI and large-scale content understanding.
Computer vision algorithms and R&D platforms for industry and scientific research.
Ph.D. in Computer Science, including a visiting Ph.D. appointment at Michigan State University in 2016.
Selected work across multimodal understanding, biometric security, and applied computer vision.
Precise region-level perception, multi-region interaction modeling, and compositional reasoning for multimodal LLMs.
Face template protection with a focus on privacy, security, and deployable biometric systems.
A widely cited study on reconstructing face images from deep templates and the security implications for biometrics.
An applied computer vision system for transmission-line inspection using drone imagery.
I regularly review for leading journals and conferences in computer vision, multimedia, and AI.
I am interested in hearing from interns working on long-context modeling, efficient AI, and multilingual or multimodal learning. You can reach me by email.