Spatially Grounded Multimodal VLMs
Developing perception-enhanced vision-language models that generate spatial tokens for complex 2D and 3D spatial understanding.
I am a final-year PhD student at MBZUAI working on spatially grounded multimodal intelligence.
My research builds vision-language and generative models that understand and create 2D, 3D, video, and 4D content.
My representative works include Perceptio, UCanDance, PointNeXt, and 3D-CoMPaT.
I am actively seeking research scientist, postdoctoral, and industry research opportunities in multimodal VLMs, spatial reasoning, world models, and generative AI starting in 2027.

PhD in Computer Vision
Mohamed bin Zayed University of Artificial Intelligence , 2023 - 2027

M.Sc. in Computer Science
King Abdullah University of Science and Technology , 2020 - 2022

B.Sc. in Computer Science and Technology
Southern University of Science and Technology , 2017 - 2021

Exchange Student, School of Computing
National University of Singapore , 2020
Developing perception-enhanced vision-language models that generate spatial tokens for complex 2D and 3D spatial understanding.
Studying controllable video synthesis through long-horizon pose diffusion conditioned on music and motion structure.
Building scalable point-cloud networks and training strategies for robust 3D object understanding and recognition.
Creating compositional 3D datasets and assets that support material, part, and object-level content generation and recognition.
Selected publications and preprints. * denotes equal contribution.
A perception-enhanced VLM using spatial token generation for 2D and 3D spatial reasoning.
Long-horizon pose diffusion for music-driven dance video synthesis.
PointNeXt revisits PointNet++ with improved training and scaling strategies for point-cloud understanding.
A large-scale 3D dataset for compositional material and part recognition.
An improved large-scale 3D vision dataset for compositional recognition.
Semi-supervised few-shot learning with prototypical random walks.










Last updated: March 2026
Featured in 36Kr Europe's documentary coverage of a 72-hour AI survival challenge, discussing AI agents, AI video creation, and human-AI interaction.
Invited by Next Capital to share perspectives on the Middle East 2030 vision and opportunities for Chinese founders in the region.
Guest appearance on 掘金中东, covering KAUST, MBZUAI, Dubai Business Associates, Middle East universities, and life across the region.