Benchmark · research site · release v0.1.0

Beyond the Frame

Benchmarking long-horizon coherence in generative visual narratives

A single beautiful image says nothing about whether the umbrella is still broken three scenes later.

The Visual Story Benchmark (VSB) generates stories with exact, machine-checkable world state, asks an image model to draw every scene, reads the images back into structured state, and scores ten consistency dimensions separately. No single evaluator decides coherence, and no single number hides the dimensions.

40stories · 10 categories · 3 tiers
210scenes
3,865checkable requirements
640frozen runs
5models on the primary board
Evidence status. All rows on this release are procedural simulators: renderer-backed models whose fault rates are stated parameters. They validate the measurement pipeline (a perfect model scores 100%; injected faults are recovered with known recall). They are not measurements of any real image model. Hosted-provider adapters ship, are transport-tested, and run only when credentials are present. No human ratings have been collected yet; the human-rating interface and reliability pipeline are exercised with clearly labelled synthetic raters. See Judges and Methods.

01 — What is measured

Ten dimensions, two kinds of fact

Every scene carries requirements derived from a replayed state machine. A requirement is either stated in that scene's text (text–image adherence) or must be remembered from earlier scenes (coherence). Scene 1 states everything; later scenes read like a screenplay and state only what changes.

Identityskin tone, hair colour and style; presence and absence
Clothingtop, trousers, accessory
Object persistenceobjects stay, leave when they should, never resurrect; counts
Object statebroken / empty / lit / open / wilted persist
Spatialleft-to-right order of characters
Environmentlocation and weather
Causala stated cause has its effect, and the effect persists
Chronologytime of day follows the story
Actionstated pose or action is depicted
Adherenceevery fact stated in the scene text

02 — Primary leaderboard

Coherence under structured-memory prompting

0%25%50%75%100%sim·amnesiac: 97% [96%, 99%]sim·amnesiac97%sim·drift-low: 98% [97%, 99%]sim·drift-low98%sim·drift-mid: 96% [95%, 98%]sim·drift-mid96%sim·oracle: 100% [100%, 100%]sim·oracle100%sim·sloppy: 76% [73%, 80%]sim·sloppy76%
Global score (unweighted mean of ten dimensions), primary condition = structured_memory. Whiskers: 95% story-bootstrap CI.

Full leaderboard with every dimension

03 — RQ1: coherence decay

How fast do memory facts get lost?

Without any memory (independent scenes) drift compounds scene after scene; with the structured-memory prompt the curve is flat because every persistent fact is restated. The gap between the two panels is the value of state prompting.

0%25%50%75%100%2345678Scenememory pass ratesim·amnesiac · Scene 2: 68% [63%, 73%]sim·amnesiac · Scene 3: 42% [35%, 49%]sim·amnesiac · Scene 4: 35% [28%, 42%]sim·amnesiac · Scene 5: 30% [22%, 38%]sim·amnesiac · Scene 6: 32% [23%, 42%]sim·amnesiac · Scene 7: 33% [24%, 44%]sim·amnesiac · Scene 8: 27% [18%, 36%]sim·amnesiacsim·drift-mid · Scene 2: 81% [74%, 87%]sim·drift-mid · Scene 3: 77% [72%, 82%]sim·drift-mid · Scene 4: 75% [69%, 80%]sim·drift-mid · Scene 5: 68% [59%, 76%]sim·drift-mid · Scene 6: 66% [51%, 79%]sim·drift-mid · Scene 7: 75% [68%, 81%]sim·drift-mid · Scene 8: 74% [65%, 81%]sim·drift-mid
Independent condition: pass rate of memory (unstated) requirements by scene index. Bands: 95% CI over stories present at that index.
sim·amnesiac sim·drift-mid
0%25%50%75%100%2345678Scenememory pass ratesim·amnesiac · Scene 2: 96% [93%, 99%]sim·amnesiac · Scene 3: 97% [95%, 98%]sim·amnesiac · Scene 4: 97% [95%, 99%]sim·amnesiac · Scene 5: 97% [96%, 99%]sim·amnesiac · Scene 6: 98% [96%, 99%]sim·amnesiac · Scene 7: 94% [86%, 100%]sim·amnesiac · Scene 8: 95% [86%, 99%]sim·drift-low · Scene 2: 99% [98%, 100%]sim·drift-low · Scene 3: 98% [97%, 99%]sim·drift-low · Scene 4: 99% [98%, 100%]sim·drift-low · Scene 5: 97% [95%, 99%]sim·drift-low · Scene 6: 98% [96%, 100%]sim·drift-low · Scene 7: 97% [94%, 99%]sim·drift-low · Scene 8: 98% [96%, 100%]sim·drift-mid · Scene 2: 94% [89%, 98%]sim·drift-mid · Scene 3: 96% [94%, 98%]sim·drift-mid · Scene 4: 93% [91%, 95%]sim·drift-mid · Scene 5: 93% [90%, 96%]sim·drift-mid · Scene 6: 90% [86%, 94%]sim·drift-mid · Scene 7: 96% [94%, 98%]sim·drift-mid · Scene 8: 96% [94%, 98%]sim·oracle · Scene 2: 100% [100%, 100%]sim·oracle · Scene 3: 100% [100%, 100%]sim·oracle · Scene 4: 100% [100%, 100%]sim·oracle · Scene 5: 100% [100%, 100%]sim·oracle · Scene 6: 100% [100%, 100%]sim·oracle · Scene 7: 100% [100%, 100%]sim·oracle · Scene 8: 100% [100%, 100%]sim·sloppy · Scene 2: 71% [65%, 76%]sim·sloppy · Scene 3: 67% [60%, 74%]sim·sloppy · Scene 4: 63% [54%, 71%]sim·sloppy · Scene 5: 66% [57%, 74%]sim·sloppy · Scene 6: 63% [50%, 75%]sim·sloppy · Scene 7: 68% [57%, 78%]sim·sloppy · Scene 8: 71% [59%, 82%]
Primary condition (structured_memory): the same curves.
sim·amnesiac sim·drift-low sim·drift-mid sim·oracle sim·sloppy

04 — RQ4: do the evaluators work?

Recovering injected faults

98.5%rule-checker recall of injected faults (sealed oracle)
74.4%precision, counting knock-on flags as false positives
0.42AUROC · pairwise figure hist
0.42AUROC · pairwise global hist
0.40AUROC · dino consecutive

Automatic-vs-human evaluation

05 — Explore

Story browser

Forty procedurally generated stories with full ground truth, side-by-side model outputs, and rule overlays.

Browse stories

Failure gallery

Every violation, labelled with an eleven-class taxonomy, searchable by model, category and scene.

Open gallery

Interventions

Six prompting conditions: independent scenes, previous text, structured memory, previous image, summarised state, full context.

Compare conditions

Rate stories

Blind, quality-controlled human evaluation interface. Export your ratings as JSON for the reliability pipeline.

Start rating