This page acts as a topic map for the research areas that connect my publications and engineering work.

VLM and multimodal LLMs

Fine-grained visual grounding, contextual reasoning, and precise region understanding for multimodal models.

Live-content understanding

VLM, LLM, and speech models for multilingual live content under real-time constraints.

Biometric security

Face template protection, reconstruction risk, and privacy-aware biometric representation learning.

Applied computer vision systems

Building practical systems that connect model capability with deployment constraints and measurable user value.

Representative Keywords

If you arrive here from an old search result, the best next stop is usually the publications page.

  • Multimodal large language models
  • Long-context modeling
  • Efficient AI
  • Speech models
  • Computer vision
  • Region-level understanding
  • Biometric security
  • Template protection
  • Privacy
  • Applied AI systems