Benchmark · research site · release v0.1.0
Beyond the Frame
Benchmarking long-horizon coherence in generative visual narratives
The Visual Story Benchmark (VSB) generates stories with exact, machine-checkable world state, asks an image model to draw every scene, reads the images back into structured state, and scores ten consistency dimensions separately. No single evaluator decides coherence, and no single number hides the dimensions.
01 — What is measured
Ten dimensions, two kinds of fact
Every scene carries requirements derived from a replayed state machine. A requirement is either stated in that scene's text (text–image adherence) or must be remembered from earlier scenes (coherence). Scene 1 states everything; later scenes read like a screenplay and state only what changes.
02 — Primary leaderboard
Coherence under structured-memory prompting
03 — RQ1: coherence decay
How fast do memory facts get lost?
Without any memory (independent scenes) drift compounds scene after scene; with the structured-memory prompt the curve is flat because every persistent fact is restated. The gap between the two panels is the value of state prompting.
04 — RQ4: do the evaluators work?
Recovering injected faults
05 — Explore
Story browser
Forty procedurally generated stories with full ground truth, side-by-side model outputs, and rule overlays.
Browse storiesFailure gallery
Every violation, labelled with an eleven-class taxonomy, searchable by model, category and scene.
Open galleryInterventions
Six prompting conditions: independent scenes, previous text, structured memory, previous image, summarised state, full context.
Compare conditionsRate stories
Blind, quality-controlled human evaluation interface. Export your ratings as JSON for the reliability pipeline.
Start rating