Leaderboard · release v0.1.0 · condition structured_memory

Component metrics first

Each cell is the macro-average over stories of the per-story pass rate in that dimension. Hover a cell for its 95% story-bootstrap CI. The global column is an unweighted mean of the ten dimensions and exists only for ranking; its sensitivity to weights is reported below.

Evidence status. Every model here is a procedural simulator with stated parameters (see Methods). The board demonstrates that the measurement stack separates models with known differences; it does not rank real systems.
Primary leaderboard. Coherence = memory-requirement pass rate; Adherence = stated-requirement pass rate.
ModelIdentityClothingObject persistenceObject stateSpatialEnvironmentCausalChronologyActionAdherenceCoherenceGlobal
sim·amnesiacsim96.3%97.2%96.5%91.9%93.5%99.8%100.0%99.7%98.3%98.7%96.3%97.5%
sim·drift-lowsim99.7%97.5%96.9%89.5%98.1%99.6%93.6%99.2%99.2%99.4%98.1%98.2%
sim·drift-midsim96.1%95.0%95.4%95.6%97.1%97.2%80.0%97.7%98.2%97.7%95.1%96.4%
sim·oraclesim100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
sim·sloppysim75.5%67.2%74.9%74.2%53.0%80.3%62.9%84.0%81.6%83.3%68.9%76.4%
By tier (global score), first-failure position, pixel pairwise similarity, and generation cost.
ModelEasyMediumHardFirst failure (rel. pos.)Pairwise figure sim.API callsCost (USD)p50 latency
sim·amnesiac96.7%97.8%97.7%0.680.5181260$0.000.004 s
sim·drift-low98.4%97.9%98.7%0.810.515210$0.000.004 s
sim·drift-mid98.8%96.8%93.1%0.600.5061470$0.000.004 s
sim·oracle100.0%100.0%100.0%1.000.521210$0.000.005 s
sim·sloppy79.3%78.6%69.1%0.240.372210$0.000.004 s

Paired comparisons

Story-level paired bootstrap on the global score (2,000 resamples). A comparison is significant when the 95% interval excludes zero.

ABΔ global95% CIP(Δ>0)Win rateSig.
sim·oraclesim·drift-low1.8%[0.8%, 3.0%]1.00075%yes
sim·oraclesim·amnesiac2.5%[1.4%, 3.9%]1.00081%yes
sim·oraclesim·drift-mid3.6%[2.2%, 5.2%]1.00084%yes
sim·oraclesim·sloppy23.6%[20.0%, 27.4%]1.000100%yes
sim·drift-lowsim·amnesiac0.8%[-1.0%, 2.4%]0.80556%no
sim·drift-lowsim·drift-mid1.9%[-0.0%, 3.5%]0.97568%no
sim·drift-lowsim·sloppy21.8%[18.1%, 25.8%]1.000100%yes
sim·amnesiacsim·drift-mid1.1%[-0.8%, 2.9%]0.87564%no
sim·amnesiacsim·sloppy21.1%[17.6%, 24.6%]1.000100%yes
sim·drift-midsim·sloppy20.0%[16.3%, 24.1%]1.00098%yes

Score weighting and its sensitivity

Weights are equal (0.100 each) by construction. We perturbed the weights 1000 times with a flat Dirichlet and re-ranked. Mean Kendall τ between the equal-weight ranking and perturbed rankings: 0.891.

ModelEqual-weight rankP(rank unchanged)
sim·oracle1100%
sim·drift-low256%
sim·amnesiac354%
sim·drift-mid491%
sim·sloppy5100%

Prompt sensitivity

The same stories with a deterministic surface paraphrase of every scene text (synonyms only; every fact kept). A benchmark that moves a lot here is measuring wording, not coherence.

Model · conditionPlainParaphraseΔ95% CI
sim·drift-mid|structured_memory96.4%96.4%0.00%[0.00%, 0.00%]

All conditions

Intervention arm: the same models under the other five prompting conditions. See Models for paired analysis.
Model · conditioncharacter identityclothingobject persistenceobject statespatial consistencyenvironmentcausal continuitychronologyaction consistencytext image adherencecoherenceglobal
sim·amnesiac · full_context98.4%96.4%94.1%88.6%96.3%97.8%84.3%98.9%97.6%98.7%96.0%96.8%
sim·amnesiac · independent60.7%53.0%53.3%61.6%42.8%77.5%63.6%90.0%61.4%90.2%45.1%68.3%
sim·amnesiac · previous_image90.5%85.7%86.8%80.9%91.5%74.6%65.7%94.1%79.5%98.5%78.3%86.8%
sim·amnesiac · previous_text70.6%60.3%59.0%70.8%50.0%88.1%50.0%97.9%69.9%91.1%59.3%75.5%
sim·amnesiac · summarized_state96.9%96.8%96.9%95.2%94.9%98.5%98.6%100.0%98.1%98.9%96.3%97.8%
sim·drift-mid · full_context94.4%94.0%98.0%89.1%90.9%96.5%91.4%99.0%94.5%98.4%92.8%95.4%
sim·drift-mid · independent83.1%74.5%75.2%75.2%83.5%87.6%57.1%98.6%86.8%94.6%74.6%85.0%
sim·drift-mid · previous_image90.6%91.4%93.3%95.8%95.3%91.8%95.7%96.0%88.9%97.3%88.4%93.2%
sim·drift-mid · previous_text88.8%86.2%88.3%88.1%86.7%92.7%90.0%98.8%91.9%97.1%85.2%91.7%
sim·drift-mid · summarized_state95.5%92.5%99.1%90.5%90.4%95.3%85.7%97.4%98.7%97.9%93.3%95.6%