OMEGA LAB / RESEARCH

Research directions

We study how agents perceive, model, reason and act, from robot policies and world models to tool use, multi-agent workflows and the safety of embodied systems.

Demo · UniT

Multimodal Perception

Representing the physical world

We study 3D geometry, multimodal fusion and scene representations, combining visual, geometric and other sensory information to understand objects and their surroundings.

UniT recovers metric-scale geometry from any number of views, shown here reconstructing an indoor scene at HKUST(GZ). The project page includes more reconstructions, results and a live demo.

Explore UniT

Robot Learning & Manipulation

Learning policies for action

We develop vision-language-action models that connect reasoning with manipulation, studying motion generation, policy improvement and efficient inference.

World Models & Prediction

Modeling how the world changes

We investigate predictive world models and future motion, with a focus on how behavior, driving style and physical dynamics inform planning.

Vision-Language Navigation

Grounding language in space

We study how agents follow language instructions in physical environments, connecting spatial perception, motion prediction and planning for ground and aerial navigation.

Intent & Decision-Making

Understanding behavior and decisions

We model intent and human driving behavior, and examine how visual evidence and context shape autonomous decisions under real-world constraints.

Multi-Agent Systems

Coordinating agents and tools

We study how agents use specialized models and tools, share information and coordinate tasks, spanning digital workflows for analysis and generation as well as collaborative aerial and ground systems.

Embodied AI Safety

Stress-testing and hardening embodied agents

We probe how perception, planning and detection models fail under backdoor, adversarial and out-of-distribution threats, designing stealthy yet effective attacks to expose vulnerabilities and developing mitigation strategies for robust autonomous driving and multimodal systems.