CUHK
UIUC
Columbia University
PITT University
The first RGB-free international competition ยท built on the CUHK-X benchmark
Jun 20 โ Sep 15, 2026 ยท Finals @ UbiComp 2026, Shanghai
The CUHK-X Multimodal Human Activity Challenge is now LIVE on Kaggle, and the full CUHK-X dataset will be released after the competition concludes. CUHK-S (a sample subset of CUHK-X) includes only 18 users and excludes the RGB modality. We welcome the community to use it and share feedback.
CUHK-X is a comprehensive multimodal dataset containing 64,267 samples of 40 actions across seven modalities designed for human activity recognition, understanding, and reasoning. Unlike existing datasets that focus primarily on recognition tasks, CUHK-X addresses critical gaps by providing the first multimodal dataset specifically designed for Human Action Understanding (HAU) and Human Action Reasoning (HARn).
The dataset was collected from 30 participants across two indoor environments covering diverse daily scenarios, using a prompt-based scene creation approach that leverages Large Language Models (LLMs) to generate logical and spatio-temporal activity descriptions. This ensures both consistency and ecological validity in the collected data.
CUHK-X provides three comprehensive benchmarks: HAR (Human Action Recognition), HAU (Human Action Understanding), and HARn (Human Action Reasoning), encompassing six distinct evaluation tasks. Our extensive experiments demonstrate significant challenges in cross-subject and cross-domain scenarios, highlighting the dataset's value for advancing robust multimodal human activity analysis.
Hardware and environment setup, seven synchronized modalities, 40 action categories and their real-world distribution.
View the data โThree benchmarks โ HAR, HAU and HARn โ with full experimental results across modalities and models.
See the numbers โStream real CUHK-S clips in the browser: synced depth/IR/thermal playback, a 44-action atlas and a reasoning quiz.
Open the observatory โIf you use CUHK-X in your research, please cite our paper:
@inproceedings{10.1145/3745756.3809209,
author={Jiang, Siyang and Yuan, Mu and Ji, Xiang and Yang, Bufang and Liu, Zeyu and Xu, Lilin and Li, Yang and He, Yuting and Dong, Liran and Lu, Wenrui and Yan, Zhenyu and Jiang, Xiaofan and Gao, Wei and Chen, Hongkai and Xing, Guoliang},
title={A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning},
year={2026},
isbn={9798400720277},
publisher={Association for Computing Machinery},
address={New York, NY, USA},
url={https://doi.org/10.1145/3745756.3809209},
doi={10.1145/3745756.3809209},
booktitle={Proceedings of the 24th Annual International Conference on Mobile Systems, Applications and Services},
pages={352--370},
numpages={19},
keywords={human action understanding, large language models, datasets},
location={University of Cambridge, Cambridge, United Kingdom},
series={MobiSys '26}
}
For dataset access, questions, or collaborations:
We thank all participants who contributed to the CUHK-X dataset collection. Special acknowledgments to the CUHK research team and collaborators who made this comprehensive multimodal dataset possible. The hardware setup and synchronization infrastructure were crucial for achieving the quality and scale of CUHK-X. The CUHK-X dataset creators obtained approval from an Institutional Review Board (IRB) to conduct their study and collect data from human subjects.
We gratefully acknowledge Dr. Jamie Du, Zhijiang Chen, Runju Fan, Yajing Feng, Peipei Li, Yutang Wei, Jiamin Wu, Yixin Xu, and Danni Yuanyong from Huizhou University, as well as Dr. Yunqi Guo from The Chinese University of Hong Kong, for their assistance in collecting the dataset. We also thank Ruijun Xia, Bohan Liu, and Guangyu Chen from The Chinese University of Hong Kong for their help with the paper artifacts.
All datasets used in this study were accessed and used under explicit authorization from their respective owners.
CUHK-X aims to advance research in healthcare monitoring, smart environments, and privacy-preserving human activity understanding. We hope this dataset serves as a valuable resource for the research community to develop more robust and practical human activity recognition systems.