🏆 The CUHK-X Challenge is LIVE — USD $20K prize pool · two tracks on Kaggle · finals @ UbiComp 2026, Shanghai Register →

Benchmarks & Tasks

CUHK-X provides three comprehensive benchmarks that progressively increase in complexity, from basic recognition to advanced reasoning:

🎯

HAR - Human Action Recognition

Objective: Traditional action classification across modalities

  • Cross-subject evaluation (LOSO protocol)
  • Cross-domain performance analysis
  • Long-tail distribution handling
  • Multimodal fusion strategies
🧠

HAU - Human Action Understanding

Objective: Comprehend actions through contextual integration

  • Action Captioning: Generate natural language descriptions
  • Emotion Analysis: Identify emotional states
  • Sequential Reordering: Organize actions chronologically
  • Action Selection: Choose relevant actions from candidates
🔮

HARn - Human Action Reasoning

Objective: Infer intentions and causal relationships

  • Next Action Prediction: Predict likely subsequent actions
  • Temporal Reasoning: Understand action progression logic
  • Contextual Inference: Consider environmental factors
  • Causal Understanding: Link actions to intentions

Novel Data Collection Framework

Leveraging Large Language Models to generate consistent, logical activity descriptions that participants then perform. This approach ensures:

Logical Consistency:

Activities follow natural progression and causality

Spatio-temporal Coherence:

Actions are contextually appropriate

Human-in-the-Loop Validation:

Quality assurance for generated scenarios

Scalable Annotation:

Efficient generation of diverse scenarios

Experimental Results

Key Findings

Our comprehensive evaluation across the three benchmarks reveals several important insights:

🎯 HAR Performance

Modality Accuracy Precision Recall F1-Score
RGB 90.89% 92.24% 91.02% 91.28%
Depth 90.46% 91.76% 90.75% 90.93%
IR 90.22% 91.53% 89.94% 90.46%
Thermal 92.57% 93.54% 93.50% 93.36%
mmWave 46.63% 48.29% 46.63% 44.53%
IMU 45.52% 40.84% 38.00% 38.32%
Skeleton 79.08% 91.46% 79.08% 84.17%

🧠 HAU Performance Highlights

These three images represent the results of action selection, emotion analysis, and action sequence, respectively.
  • QwenVL-7B: Consistently best performer across tasks
  • VLLaVA-7B: Strong performance in depth and IR modalities
  • Emotion Analysis: Up to 77.77% accuracy with thermal imaging
  • Sequential Reordering: 68.5% accuracy for complex temporal reasoning
action_selection emotion_analysis sequential_action

🔮 HARn Insights

  • Reasoning vs Captioning: Reasoning models significantly outperform captioning models
  • Modality Impact: Depth and IR often superior to RGB for reasoning tasks
  • Model Scale: Larger models (7B) consistently outperform smaller ones
  • Context Understanding: Critical for next action prediction accuracy
HARn result