DEPTH
3-D structure from a Vzense NYX 650 β bodies as geometry, not identity.
β IN CUHK-S// MOBISYS 2026 Β· MULTIMODAL HUMAN ACTION DATASET
CUHK-X records people living ordinary moments β brushing teeth, folding clothes, doing squats β through seven synchronized sensor streams. This observatory lets you walk through CUHK-S, the public, privacy-preserving subset: depth, infrared and thermal video with captions, pace labels, action chains and next-action logic.
PARENT DATASET CUHK-X β 64,217 samples Β· 7 modalities Β· 30 subjects β available under DUA
β· DRAG THE DIVIDERS β ONE MOMENT, THREE SENSORS
Three instruments, one per benchmark. Pick a station to start exploring.
Every sample in CUHK-X is captured simultaneously by seven instruments. Three privacy-preserving streams ship in the public CUHK-S subset.
3-D structure from a Vzense NYX 650 β bodies as geometry, not identity.
β IN CUHK-SActive IR video that keeps seeing when the lights go out.
β IN CUHK-SHeat signatures of the body β robust, anonymous, always-on.
β IN CUHK-SReference color video for grounding and annotation.
FULL CUHK-XTI radar point clouds β motion sensing through privacy.
FULL CUHK-XOn-body inertial traces of every gesture.
FULL CUHK-X3-D pose sequences distilled from the cameras.
FULL CUHK-XCUHK-S/ βββ HAU/ sequential scenes Β· depth+IR+thermal β βββ data/ user/βscene-env-trial/β*.mp4 β βββ GT/ captions Β· emotion Β· ordered chains βββ HARn/ single actions Β· depth+IR βββ data/ 40 action classes (+4 extra) βββ GT/ next-action logic + candidates
Full CUHK-X (7 modalities, 64,217 samples) is available for non-commercial research under a Data Use Agreement. Collection was IRB-approved; public clips are de-identified sensor streams.
@inproceedings{jiang2026cuhkx,
title = {CUHK-X: A Large-Scale Multimodal Dataset
and Benchmark for Human Action Recognition,
Understanding and Reasoning},
author = {Jiang, Siyang and others},
booktitle = {Proc. ACM MobiSys},
year = {2026}
}
Questions & collaborations β syjiang [at] ie.cuhk.edu.hk