|
I am a senior postdoctoral researcher at the CS department and the Big Data Institute in the University of Oxford, working with Prof. Yarin Gal and other CeBAM PIs. Previously, I was a senior research fellow at National University of Singapore, where I worked with Prof. Mong-Li Lee, Prof. Wynne Hsu, Prof. Tat-Seng Chua and Prof. Shuicheng Yan. I also worked as a visiting researcher at Microsoft Research Asia, an associate researcher at Skywork AI Singapore, and SEA AI lab, respectively.
My research has been published in top-tier ML/NLP/CV/MM venues, e.g., ICML, NeurIPS, ICLR, ACL, CVPR, ACM-MM, AAAI, WWW, SIGIR, EMNLP, TPAMI, IJCV, TMM, TKDE, TOIS, TNNLS, TASLP. My papers were selected as Most Influential Papers by Paper Digest, and ESI Highly Influential Papers, 2024 WAIC Outstanding Paper Award and also several Best Papers (Nominations as well) on some venues. I was awarded the World AI Conference Bright Star in 2026 & Rising Star in 2023, and also ranked as Top 2% Scientists Worldwide 2024&2025 by Stanford University. I’ve regularly served as (Senior) Area Chair or Senior Program Committee of top-tier conferences. I was on the organization committee of conferences, WSDM, EMNLP, ACL, ACM MM, etc. I serve as the Associate Editor of some journals, e.g., IEEE TAFFC, IEEE TASLP, ACM TALLIP, Neurocomputing. My Ph.D thesis was awarded the Excellent Doctoral Thesis of Chinese Information Processing Society (CIPS).
My research interests lie in NLP, CV, and the intersection of both (i.e., Multimodal/Vision-Language Learning). My long-term goal is to achieve human-level AI centered around multimodal LLMs & generalists. I pay the main focus on building large foundation multimodal models and bridging physical and mental worlds. Know me via some latest series of representative works (see research statement for more):
Recently, I also extensively explore the AI for science, including 1) psychology & social norm studies, 2) bio-/medicine & healthcare & clinics, and 3) material science, by integrating the advanced LLM/agent methodologies.
I am constantly looking for collaborations on the above topics. Remote manner is also supported. For promising students I will provide sufficient GPUs. Hit me up, if you are a Ph.D/master/bachelor student and interested in what I am doing now (with potential vacancies for research interns/RAs/visiting). For students from University of Oxford, I’m particularly looking for collaborations on world modeling and AI scientist. Please describe your research status and attach your resume & statement.
We are excited to release OmniScientist, the first Omni-Modal Omni-Discipline AI Scientist for end-to-end multidisciplinary scientific research. 🚀 Skill and Desktop versions are now publicly available! Demo Page, Paper, Code available on the project page.
• 13 Aug 2026We are excited to release V-RAE, a video autoencoder that constructs generative latent spaces directly from frozen visual representations. Project Page, Paper, Code
• 29 Jul 2026We are excited to introduce Mental World Modeling (MWM), the first theoretical framework for explicitly modeling the mental world beyond conventional physical world modeling. 🚀 Demo and code are publicly available! Demo Page, Paper, Code
• 30 Jun 2026We are organizing an IJCV Special Issue on Multimodal Unified Comprehension and Generation (MUCG). Call for Papers, welcome submissions!
• 10 Jun 2026We are co-organizing the 2nd MLLM for Unified Comprehension and Generation (MUCG 2026) Workshop at ECCV 2026. Call for Papers, welcome submissions!