Xu Cao
Building AI Agent that understands humans — and grows up with children.
I am a CS PhD candidate at the University of Illinois Urbana-Champaign, advised by Prof. James M. Rehg and Prof. Jimeng Sun. My research sits at the intersection of vision-language models, agentic AI, and social embodied AI, with applications spanning AR smart glasses and pediatric foundation models.
Driven by my own history as a pediatric rare-disease patient, I co-founded PediaMed AI to advance AI built specifically for pediatrics and child development — in particular, intelligent agents for the early identification, diagnosis, and treatment of cerebral palsy, autism, developmental disorders, and rare neurodevelopmental diseases.
Previously, I worked as a machine learning engineer at SambaNova Systems and Tencent on autonomous driving and generative AI, and held research internships at Google Research and NEC Labs America. I received my MS in Computer Science from the Courant Institute at NYU, and my BS in Data Science and Chemistry from Fudan University, graduating ranked 1st in my major.
Research — three pillars
Learning from Children
Cognitive-learning-inspired Multimodal LLMs: studying how children acquire perception, language, and reasoning to build models that learn the same way.
Representative: CogSense (Arxiv) · VCog-Bench (COLM 2025)
Social Embodied AI Agents (core works during my PhD)
Agents that perceive non-verbal social cues (gaze, gesture, and emotion) — moving from social perception toward theory of mind, on platforms like AR smart glasses.
Representative: GazeAnywhere (CVPR 2026) · SocialGesture (CVPR 2025) · TransGesture (NeurIPS 2025)
Pediatric Foundation Models & Healthcare AI
Foundation models for child development and neurodevelopmental disorders, from autism screening to video-based disease progression simulation.
Representative: ChildGait (ECCV 2026) · Disease Progression Simulation (CHIL 2026 Best Paper) · MpoxVLM (ML4H 2024) · EHRMamba (ML4H 2024) · ViTASD (ICASSP 2023) · AggPose (IJCAI 2022)
Startup
PediaMed AI
With Dr. Meihuan Huang, I co-found and co-direct a lab of ten researchers building AI for pediatrics and child development — from cognitive multimodal LLMs to AR agents for children. We also run the AI4CHL (ICLR 2025) and CV4CHL (CVPR 2026) workshop series.
Several projects in cognitive multimodal LLMs and AR agents are open to venture investment — feel free to reach out.
PediaMed AI — selected publications
-
ECCV 2026
ChildGait: Decoding Children's Gait Behavior from Video
-
CVPR 2026Highlight
Enhancing Video Vision Language Model with Hippocampal Sensing
-
Tech Report
-
Tech Report
Toward Cognitive Supersensing in Multimodal Large Language Models
News
- Video-based Disease Progression Simulation won the Best Paper Award at CHIL 2026.
- Serving as an Area Chair for NeurIPS 2026.
- ChildGait — joint work by UIUC, PediaMed AI, and Shenzhen Children's Hospital — is accepted to ECCV 2026.
- Presenting GazeAnywhere at the NVIDIA Tech Talk at CVPR 2026 (June 5).
- Several demos accepted to CVPR 2026 — meet the team in Denver to discuss collaboration and funding.
- Multiple papers accepted to CVPR 2026 with collaborators at UIUC, Google AR, and PediaMed AI.
- Organizing the CVPR 2026 Workshop on Computer Vision for Children (CV4CHL) with Prof. Ismini Lourentzou's team.
- Awarded the Next Generation Trainee Fellowship from CIFAR.
- Two papers accepted to NeurIPS 2025; recognized as a NeurIPS 2025 Top Reviewer.
Selected publications
-
CHIL 2026Best Paper Award
-
CVPR 2026Demo
-
NeurIPS 2025
-
CVPR 2025
SocialGesture: Delving into Multi-person Gesture Understanding
-
COLM 2025
What is the Visual Cognition Gap between Humans and Multimodal LLMs?
-
CVPR 2024
MAPLM: A Real-World Large-Scale Vision-Language Benchmark for Map and Traffic Scene Understanding
-
AAAI 2023IAAI Oral · Innovative Application Award
THMA: Tencent HD Map AI System for Creating HD Map Annotations
-
IJCAI 2022Oral
AggPose: Deep Aggregation Vision Transformer for Infant Pose Estimation