// STATION 03 · HUMAN ACTION REASONING
The hardest task in CUHK-X: watch a depth clip, infer what the person will do next. Play the benchmark yourself — the labels are real.
ROUND — / —
> OBSERVED ACTION — …
> PREDICT WHAT THE SUBJECT DOES NEXT:
How models score on the full CUHK-X (MobiSys ’26 paper). Beat the table above, then try beating it with your own model.
| MODEL | CAPTION BLEU-1 | EMOTION ACC | REORDER ACC |
|---|---|---|---|
| QWENVL-7B | 18.04 | 55.03 | 60.00 |
| VLLAVA-7B | 12.86 | 73.34 | 5.29 |
| INTERNVL-8B | 0.72 | 31.35 | 74.03 |
Even 7-8B VLMs struggle with ordering and intent — that gap is the benchmark.