CUHK-X OBSERVATORY

// STATION 03 · HUMAN ACTION REASONING

01 REASONING CONSOLE HARN BENCHMARK

The hardest task in CUHK-X: watch a depth clip, infer what the person will do next. Play the benchmark yourself — the labels are real.

ROUND — / —

> OBSERVED ACTION —
> PREDICT WHAT THE SUBJECT DOES NEXT:

SCORE 0/0 QWENVL-7B BASELINE · 60.0% REORDERING ACC

02 MACHINE BASELINES

How models score on the full CUHK-X (MobiSys ’26 paper). Beat the table above, then try beating it with your own model.

HAR · CROSS-TRIAL ACCURACYper modality
HAU · VISION-LANGUAGE MODELSselected tasks
MODELCAPTION BLEU-1EMOTION ACCREORDER ACC
QWENVL-7B18.0455.0360.00
VLLAVA-7B12.8673.345.29
INTERNVL-8B0.7231.3574.03

Even 7-8B VLMs struggle with ordering and intent — that gap is the benchmark.

◂ STATION 02 · ACTION ATLAS OVERVIEW ▸