CUHK-X OBSERVATORY

// MOBISYS 2026 Β· MULTIMODAL HUMAN ACTION DATASET

SEVEN WAYS
OF SEEING
A HUMAN ACTION

CUHK-X records people living ordinary moments β€” brushing teeth, folding clothes, doing squats β€” through seven synchronized sensor streams. This observatory lets you walk through CUHK-S, the public, privacy-preserving subset: depth, infrared and thermal video with captions, pace labels, action chains and next-action logic.

PARENT DATASET CUHK-X β€” 64,217 samples Β· 7 modalities Β· 30 subjects β€” available under DUA

SPECTRAL SWEEP SYNC Γ—3 Β· β€”
DEPTH
INFRARED
THERMAL
00:00.0

⟷ DRAG THE DIVIDERS β€” ONE MOMENT, THREE SENSORS

β€”VIDEO CLIPS
β€”SENSOR-STREAM HOURS
β€”SUBJECTS
β€”ACTION CLASSES
β€”SCRIPTED SCENES
β€”LABEL ROWS

01 STATIONS

Three instruments, one per benchmark. Pick a station to start exploring.

STATION 01 Β· HAU Sequence Lab Long scripted scenes played back in depth βŠ• infrared βŠ• thermal sync β€” with captions, pace labels and ordered action chains. 814 TAKES Β· 7 SCENES Β· 3 SYNCED STREAMSENTER β–Έ STATION 02 Β· HAR / HARN Action Atlas Every action class on one wall. Hover to scrub film strips, click to inspect dual-stream clips, read the coverage instruments. 40 CLASSES Β· DEPTH βŠ• IRENTER β–Έ STATION 03 Β· HARN Reasoning Console Watch a depth clip and predict the next action. Real ground truth, real distractors β€” score yourself against the VLM baselines. 1,492 LOGIC LABELS Β· BEAT THE VLM BASELINEENTER β–Έ

02 SENSOR ARRAY

Every sample in CUHK-X is captured simultaneously by seven instruments. Three privacy-preserving streams ship in the public CUHK-S subset.

DEPTH

3-D structure from a Vzense NYX 650 β€” bodies as geometry, not identity.

● IN CUHK-S

INFRARED

Active IR video that keeps seeing when the lights go out.

● IN CUHK-S

THERMAL

Heat signatures of the body β€” robust, anonymous, always-on.

● IN CUHK-S

RGB

Reference color video for grounding and annotation.

FULL CUHK-X

MMWAVE

TI radar point clouds β€” motion sensing through privacy.

FULL CUHK-X

IMU

On-body inertial traces of every gesture.

FULL CUHK-X

SKELETON

3-D pose sequences distilled from the cameras.

FULL CUHK-X

03 ACCESS PROTOCOL

DOWNLOADCUHK-S Β· open
CUHK-S/
β”œβ”€β”€ HAU/        sequential scenes Β· depth+IR+thermal
β”‚   β”œβ”€β”€ data/   user/​scene-env-trial/​*.mp4
β”‚   └── GT/     captions Β· emotion Β· ordered chains
└── HARn/       single actions Β· depth+IR
    β”œβ”€β”€ data/   40 action classes (+4 extra)
    └── GT/     next-action logic + candidates

Full CUHK-X (7 modalities, 64,217 samples) is available for non-commercial research under a Data Use Agreement. Collection was IRB-approved; public clips are de-identified sensor streams.

CITE
@inproceedings{jiang2026cuhkx,
  title     = {CUHK-X: A Large-Scale Multimodal Dataset
               and Benchmark for Human Action Recognition,
               Understanding and Reasoning},
  author    = {Jiang, Siyang and others},
  booktitle = {Proc. ACM MobiSys},
  year      = {2026}
}

Questions & collaborations β€” syjiang [at] ie.cuhk.edu.hk