MMGR: Multi-Modal Generative Reasoning Benchmark and Evaluation
EMNLP 2026 (Main, Poster) · MSLD 2026 (Poster) · PDF · Website · Code · Data · Poster
Evaluate generated media as solutions to reasoning problems. MMGR covers 10 tasks and 1,853 instances across abstract reasoning, embodied navigation and physical commonsense. A convincing image or video must also preserve the task’s rules, state and causal structure.
Maze and Sudoku use pixel-based path reconstruction and OCR with constraint checks. Other tasks use task-specific VLM rubrics, with human agreement analysis reported in the paper. For video, a chain-of-frame evaluation checks whether intermediate states make valid progress toward the final outcome.
In the reported zero-shot experiments, Veo-3, Sora-2 and Wan-2.2 all score 0% on Sudoku. The results expose a gap between visual fluency and reasoning correctness. Model coverage varies by task; navigation and physical-commonsense scores should be read alongside the reported evaluator reliability.
Poster
