OMEGA LAB · HKUST(GZ)

OMEGA Lab

Omnimodal Multi-Embodiment
Generalist Agents

Our research connects multimodal perception, intent and motion modeling, and autonomous decision-making in physical systems and digital workflows.

Explore our research
THE OMEGA PERSPECTIVE

Omnimodal

Modalities & representations

Multi-Embodiment

Land, water, air & robotic systems

Generalist Agents

World models, reasoning & decision-making

  1. Perceive
  2. Model
  3. Predict
  4. Reason
  5. Plan
  6. Act

Omnimodal Multi-Embodiment Generalist Agents

Select any number of modalities to highlight their matching orbits. Select again to deselect. Drag along the orbit to change its position; release to continue orbiting. Arrow keys also adjust position. Escape returns to the overview.

Omnimodal · Multi-Embodiment · Generalist Agents

The three O, ME and GA sections fade, and their letters join to spell OMEGA, then condense into the lab’s Greek Ω logo. A nucleus and orbital paths then appear. After orbiting, the graphic unfolds back into the three sections and repeats. Each modality has its own tilted orbital plane around the central nucleus, with vision, language, geometry and tactile emphasized. The nine modalities and representations are: vision, language, geometry, tactile, audio, radio, proprioception, intent and trajectory. Inside the nucleus, ten kinds of embodiments form a ring around the central Ω: quadrupeds, UGVs, autonomous cars, USVs and UAVs; then dexterous hands, robot arms, bimanual systems, wheeled dual-arm robots and humanoids. Labels and tracks become smaller, softer and occluded as they pass behind the translucent nucleus. Multiple modalities can be selected independently; dragging pauses only the dragged modality until release. The Generalist Agents section forms a six-stage loop: Perceive, Model, Predict, Reason, Plan and Act. At the center, digital tasks and physical tasks connect through an agent and tool use. Software tools and control interfaces let agents act in digital workflows and physical environments, including autonomous navigation. Outward calls and returning results or observations form a shared feedback loop.

FROM THE LAB

News & media

OUR RESEARCH

Research directions

We study how agents perceive, model, reason and act, from robot policies and world models to tool use, multi-agent workflows and the safety of embodied systems.

Explore research

Multimodal Perception

Representing the physical world

We study 3D geometry, multimodal fusion and scene representations, combining visual, geometric and other sensory information to understand objects and their surroundings.

Demo · UniTResearch & publications

Robot Learning & Manipulation

Learning policies for action

We develop vision-language-action models that connect reasoning with manipulation, studying motion generation, policy improvement and efficient inference.

Research & publications

World Models & Prediction

Modeling how the world changes

We investigate predictive world models and future motion, with a focus on how behavior, driving style and physical dynamics inform planning.

Research & publications

Vision-Language Navigation

Grounding language in space

We study how agents follow language instructions in physical environments, connecting spatial perception, motion prediction and planning for ground and aerial navigation.

Research & publications

Intent & Decision-Making

Understanding behavior and decisions

We model intent and human driving behavior, and examine how visual evidence and context shape autonomous decisions under real-world constraints.

Research & publications

Multi-Agent Systems

Coordinating agents and tools

We study how agents use specialized models and tools, share information and coordinate tasks, spanning digital workflows for analysis and generation as well as collaborative aerial and ground systems.

Research & publications

Embodied AI Safety

Stress-testing and hardening embodied agents

We probe how perception, planning and detection models fail under backdoor, adversarial and out-of-distribution threats, designing stealthy yet effective attacks to expose vulnerabilities and developing mitigation strategies for robust autonomous driving and multimodal systems.

Research & publications

PROJECTS & RESOURCES

Projects & open source

Explore our research through project pages, demonstrations and code.

PEOPLE

Our team

Meet the principal investigator and explore our member directory.

View all people
Xinhu Zheng

Xinhu Zheng 郑心湖

Assistant Professor

The Hong Kong University of Science and Technology (Guangzhou)

Intelligent Transportation · Internet of Things

Xinhu Zheng studies multimodal perception, intent and trajectory prediction, and autonomous systems.

PROSPECTIVE STUDENTS & RESEARCHERS

Join OMEGA Lab

Bring your questions about multimodal intelligence and agents. Explore PhD, MPhil and postdoctoral pathways, university support, and how to get in touch.

Opportunities & applications

GET IN TOUCH

Research &
collaboration

For research collaboration, prospective student enquiries and academic exchange, contact Xinhu Zheng.

xinhuzheng@hkust-gz.edu.cn

Office W2 L5 509
The Hong Kong University of Science and Technology (Guangzhou)

University faculty profile