๐Ÿ† The CUHK-X Challenge is LIVE โ€” USD $20K prize pool ยท two tracks on Kaggle ยท finals @ UbiComp 2026, Shanghai Register โ†’

Hardware and Environment Setup

We spans a multi-room home and supports three tasks: HAR, HAU (captioning task), and HARn (question answering task), integrating diverse modalities, including RGB, depth, thermal, infrared, IMU, skeleton, and mmWave, to enable robust perception and reasoning in complex indoor contexts.

CUHK-X was collected using a sophisticated multi-sensor setup ensuring synchronized data capture across all modalities:

Vzense NYX 650

RGB-D camera providing color and depth information

Texas Instruments Radar

mmWave sensing for privacy-preserving motion detection

IMU Sensors

Motion and orientation tracking with high temporal resolution

Thermal Cameras

Heat signature analysis for environmental robustness

Synchronized Recording

Temporal alignment across all modalities for consistent analysis

Hardware setup showing the multi-sensor configuration

Layout with room-wise visual annotations (Bedroom, Kitchen, Bathroom, and Living Room) showing corresponding example images and sensor placements. The icon indicates the location of the ambient sensor:

Background

Dataset Overview

CUHK-X represents a significant advancement in multimodal human activity datasets, featuring:

  • Seven Synchronized Modalities: RGB, Infrared (IR), Depth, Thermal, IMU, mmWave Radar, and Skeleton data
  • Large-Scale: 64,267 annotated action samples from 30 diverse participants
  • Dual Data Structure: Both singular actions (30,000+ samples) and sequential activities for temporal reasoning
  • Rich Annotations: LLM-generated captions with human-in-the-loop validation
  • Environmental Diversity: Indoor and outdoor settings with varying conditions

Modality Specifications

๐ŸŽฅ
RGB Video
Standard color video recordings
๐Ÿ“
Depth
3D spatial information from depth cameras
๐Ÿ”ฅ
Thermal
Heat signature analysis
๐ŸŒก๏ธ
Infrared (IR)
Thermal imaging for lighting robustness
๐Ÿ“ก
mmWave Radar
Privacy-preserving motion detection
๐Ÿฆด
Skeleton
3D pose estimation and joint tracking
๐Ÿ“ฑ
IMU
Inertial Measurement Unit for motion
Dataset overview showing modality examples

Action Categories and Distribution

Categories

Personal Care (6 actions)
Washing face, Brushing teeth, Combing hair, Undressing, Wiping hands, Getting Dressed
Eating and Drinking (6 actions)
Drinking, Eating, Grabbing utensils, Pouring, Stirring, Peeling fruit
Household (5 actions)
Sweeping, Mopping, Washing dishes, Wiping surface, Folding clothes
Working (6 actions)
Typing on a keyboard, Writing, Calling, Checking the time, Reading, Turning a page
Socializing and Leisure (5 actions)
Taking a selfie, Playing board games, Watching TV, Using a phone, Listening to the music with headphones
Sports and Exercises (9 actions)
Walking, Lunges, Sitting down, Lying down, Standing up, Stretching, Jumping jacks, Squats, Running
Caring and Helping (3 actions)
Taking medicine, Checking body temperature, Massaging oneself
Action categories and distribution

Distribution

  • Distribution Feature:> The dataset follows a long-tail distribution (a small number of actions account for a large proportion of occurrences, while most are infrequent), consistent with the common imbalance of real-world datasets.
  • Category Diversity: Covers basic daily activities, work-related tasks, household chores, and physical exercises, providing a rich foundation for human activity recognition.
  • Data Scale:
    • Each participant contributes over 30 minutes of footage with more than 100 samples.
    • Vision modality includes 4,029 clips (total duration: 19 hours and 29 minutes).
Action frequency

Data Visualization

This is an example that includes seven modalities: RGB, IR, Thermal, Depth, Skeleton, Radar, and IMU, which were recorded at the same time. We have a 9-axis IMU, but we've only shown these data for simplicity.