MMGR: Multi-Modal Generative Reasoning Benchmark and Evaluation

← All research

Reasoning & Agents

EMNLP 2026 (Main, Poster) · MSLD 2026 (Poster) · PDF · Website · Code · Data · Poster

Evaluate generated media as solutions to reasoning problems. MMGR covers 10 tasks and 1,853 instances across abstract reasoning, embodied navigation and physical commonsense. A convincing image or video must also preserve the task’s rules, state and causal structure.

Maze and Sudoku use pixel-based path reconstruction and OCR with constraint checks. Other tasks use task-specific VLM rubrics, with human agreement analysis reported in the paper. For video, a chain-of-frame evaluation checks whether intermediate states make valid progress toward the final outcome.

MMGR's three reasoning domains, image and video generation, and human and automatic evaluation pipeline
Figure 1: MMGR connects controlled reasoning tasks to image and video generation, then checks both the outcome and the solution process. The maze illustrates the evaluation pipeline. Click to enlarge.

In the reported zero-shot experiments, Veo-3, Sora-2 and Wan-2.2 all score 0% on Sudoku. The results expose a gap between visual fluency and reasoning correctness. Model coverage varies by task; navigation and physical-commonsense scores should be read alongside the reported evaluator reliability.

Poster

Download poster (PDF)

Poster for MMGR: Multi-Modal Generative Reasoning Benchmark and Evaluation
Click the poster to view the full-resolution PDF.