๐Ÿ† The CUHK-X Challenge is LIVE โ€” USD $20K prize pool ยท two tracks on Kaggle ยท finals @ UbiComp 2026, Shanghai Register โ†’
MobiSys '26

A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning

Siyang Jiang, Mu Yuan, Xiang Ji, Bufang Yang, Zeyu Liu, Lilin Xu, Yang Li, Yuting He, Liran Dong, Wenrui Lu, Zhenyu Yan, Xiaofan Jiang, Wei Gao, Hongkai Chen, Guoliang Xing

CUHK-X is a large-scale multimodal dataset of 64,267 samples covering 40 actions performed by 30 participants across two indoor environments. Going beyond recognition-only datasets, it introduces benchmarks for Human Action Understanding (HAU) and Human Action Reasoning (HARn) alongside classic HAR, using a prompt-based scene-creation method that leverages LLMs to generate logically and spatio-temporally consistent activity sequences. Its three benchmarks span six tasks; state-of-the-art models reach 76.52% (HAR), 40.76% (HAU), and 70.25% (HARn), underscoring the difficulty of fine-grained multimodal action understanding.

MobiSys '26 Demo

MACS — a Portable System for Multimodal Activity Recording with Built-In Annotation

Bohan Liu, Siyang Jiang, Mu Yuan, Hongkai Chen, Guoliang Xing

MACS is a portable, self-contained system for capturing synchronized multimodal recordings of human activity. It integrates multiple sensing modalities into a single portable rig and provides built-in annotation, letting researchers label activities at capture time instead of through costly post-processing. MACS serves as the data-collection backbone behind large-scale corpora such as CUHK-X, substantially lowering the cost of building high-quality, well-labeled multimodal activity datasets.