- Source
- arXiv
- Published
- Runtime
- 0:00
- Snippets
- 4
A conversation between
NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
§02
Snippets
-
NARU evaluates narrative evolution and cultural understanding jointly across 146.8 hours of Japanese video—tasks existing benchmarks rarely assess together.
Reveals that current MLLMs fail at long-range narrative integration and culturally grounded reasoning, two capabilities essential for real-world video understanding.
-
A hierarchical memory-based pipeline structures annotations into events, narratives, and cultural layers, then synthesizes questions via task-oriented generation and iterative shortcut removal.
Scalable annotation method enables construction of high-quality, multi-faceted benchmarks while maintaining consistency across two native-speaker verification stages.
-
NARU spans four narrative dimensions and five cultural dimensions, with questions grounded in high-context non-English video content.
Benchmarking cultural nuance in non-English media exposes blind spots in MLLMs trained primarily on English-centric datasets and assumptions.
-
Evaluations of eight model configurations reveal substantial, persistent limitations in both long-range narrative integration and culturally grounded reasoning.
Identifies concrete failure modes that future MLLM development must address to move beyond shallow event detection toward genuine story comprehension.
§03
Synthesis
The Gap in Video Understanding
Existing video benchmarks focus on spotting individual events—who did what, when—but miss two interconnected challenges: how narratives evolve across hours of footage, and how cultural context shapes meaning in ways no subtitle makes explicit. The authors find that today's multimodal AI models (MLLMs) struggle badly at both. To measure this systematically, they introduce NARU, a Japanese-language benchmark with 1,481 questions spanning 146.8 hours of video, designed to stress-test narrative tracking and cultural reasoning simultaneously.
Why Japanese? It's a high-context language and culture where implication often outweighs statement. Western benchmarks, mostly in English, tend to focus on literal content; they don't capture the subtlety of what remains unsaid. The choice reflects a real gap: most benchmarks cluster in English-heavy domains and miss what matters in non-English media.
How the Benchmark Works
The authors built NARU by decomposing the annotation problem into layers. Raw videos first get structured as events (discrete moments with actors and actions), then woven into narratives (how events connect across time), then tagged for cultural meaning (unspoken social norms, emotional registers, aesthetic choices). This hierarchy lets annotators impose order on chaos—a crucial step when dealing with 146.8 hours of content.
Questions didn't just appear. The team used task-oriented synthesis: given an event or narrative segment, automatically generate candidate questions, then iteratively remove "shortcuts"—queries solvable by surface-level matching without true understanding. Two native-speaker verification rounds (68 annotators total) caught inconsistencies and cultural blind spots.
The benchmark spans four narrative dimensions (tracking character arcs, causal chains, temporal shifts, and thematic threads) and five cultural dimensions (social dynamics, emotional subtext, aesthetic conventions, historical allusion, and values). This structure reveals what breaks: is a model failing because it loses the thread across hours, or because it misreads a culturally loaded gesture?
What Breaks, and Why It Matters
Testing eight model configurations—ranging from smaller to larger MLLMs—exposed consistent weaknesses. Models handle short-range event retrieval reasonably but collapse on questions requiring synthesis across 30 minutes or more. Cultural reasoning is worse: models either ignore context entirely or hallucinate "explanations" that feel plausible but miss the actual cultural logic at play.
The results matter because long-form video is increasingly central to how people consume media, yet we have no systematic way to evaluate whether AI can follow it. NARU provides that measuring stick. By grounding questions in real Japanese media and real cultural nuance—rather than synthetic data—it forces models to grapple with something they can't pattern-match past: sustained, context-dependent meaning.
This isn't a complaint; it's a roadmap. The benchmark isolates failure modes precisely enough that researchers can target them. If you want MLLMs that actually understand narrative cinema, documentaries, or any footage longer than a TikTok, NARU is now the test.
Mine your own.
Lode is a workbench, not a feed. Paste a YouTube URL. The model proposes a transcript, a set of quote-grounded snippets, a synthesis essay, and the fan-out. You decide what stays.