A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
CUHK-X is a large-scale multimodal dataset of 64,267 samples covering 40 actions performed by 30 participants across two indoor environments. Going beyond recognition-only datasets, it introduces benchmarks for Human Action Understanding (HAU) and Human Action Reasoning (HARn) alongside classic HAR, using a prompt-based scene-creation method that leverages LLMs to generate logically and spatio-temporally consistent activity sequences. Its three benchmarks span six tasks; state-of-the-art models reach 76.52% (HAR), 40.76% (HAU), and 70.25% (HARn), underscoring the difficulty of fine-grained multimodal action understanding.