Judge validation · RQ4
Which automatic metrics track coherence?
Two reference standards are available. The sealed oracle records exactly which facts a simulator corrupted in each scene; against it, every evaluator's fault-detection performance is measurable. Human ratings, when collected, give story-level rank correlations. Neither is treated as truth.
Rule checker vs the sealed oracle
| Class | Injected | Detected | Recall | Flagged | Precision |
|---|---|---|---|---|---|
| identity drift | 1952 | 1897 | 97.2% | 2496 | 76.0% |
| attribute drift | 1871 | 1854 | 99.1% | 2368 | 78.3% |
| background drift | 570 | 570 | 100.0% | 570 | 100.0% |
| adherence failure | 354 | 347 | 98.0% | 594 | 58.4% |
| object disappearance | 281 | 281 | 100.0% | 627 | 44.8% |
| spatial contradiction | 175 | 174 | 99.4% | 331 | 52.6% |
| state reversal | 141 | 141 | 100.0% | 141 | 100.0% |
| chronology violation | 115 | 115 | 100.0% | 115 | 100.0% |
| causal contradiction | 26 | 26 | 100.0% | 27 | 96.3% |
| resurrection | 10 | 9 | 90.0% | 10 | 90.0% |
| count inconsistency | 6 | 6 | 100.0% | 6 | 100.0% |
Precision counts knock-on failures (an omitted character fails every attribute check) as false positives; see the paper for the decomposition.
Continuous signals vs the oracle
Scene-level: can the signal tell a frame with at least one injected visual fault from a clean one (AUROC, 0.5 = chance)? Story-level: Spearman ρ between the signal's story mean and oracle coherence (1 − faults/requirements).
| Signal | n | AUROC | Positive rate |
|---|---|---|---|
| pairwise figure hist | 2720 | 0.418 | 57% |
| pairwise global hist | 2720 | 0.422 | 57% |
| dino consecutive | 2720 | 0.401 | 57% |
| clip text image | 3360 | 0.371 | 55% |
| Signal | n | Spearman ρ | 95% CI | Kendall τ |
|---|---|---|---|---|
| rule checker global | 640 | 0.943 | [0.93, 0.96] | 0.817 |
| rule checker coherence | 640 | 0.970 | [0.96, 0.98] | 0.885 |
| pairwise figure mean | 640 | 0.236 | [0.17, 0.31] | 0.166 |
| pairwise figure min | 640 | 0.321 | [0.25, 0.39] | 0.226 |
| clip mean | 640 | 0.122 | [0.04, 0.20] | 0.083 |
| dino mean | 640 | 0.293 | [0.22, 0.37] | 0.205 |
Automatic metrics vs synthetic ratings
95 rated runs. Rows: automatic signals; columns: rated dimensions. Cells: Spearman ρ (n).
| Signal | causal coherence | identity consistency | narrative coherence | object consistency | prompt adherence | visual quality |
|---|---|---|---|---|---|---|
| rule global | 0.50 (95) | 0.40 (95) | 0.26 (95) | 0.37 (95) | 0.30 (95) | 0.07 (95) |
| rule coherence | 0.49 (95) | 0.42 (95) | 0.21 (95) | 0.32 (95) | 0.22 (95) | 0.09 (95) |
| rule character identity | 0.42 (95) | 0.49 (95) | 0.16 (95) | 0.22 (95) | 0.20 (95) | 0.08 (95) |
| rule clothing | 0.38 (95) | 0.38 (95) | 0.17 (95) | 0.22 (95) | 0.13 (95) | 0.06 (95) |
| rule object persistence | 0.59 (72) | 0.42 (72) | 0.23 (72) | 0.59 (72) | 0.38 (72) | 0.02 (72) |
| rule environment | 0.34 (95) | 0.41 (95) | 0.38 (95) | 0.36 (95) | 0.43 (95) | 0.00 (95) |
| rule text image adherence | 0.49 (95) | 0.41 (95) | 0.31 (95) | 0.38 (95) | 0.34 (95) | 0.04 (95) |
| pairwise figure mean | 0.20 (95) | -0.04 (95) | 0.20 (95) | 0.10 (95) | 0.22 (95) | 0.01 (95) |
| clip mean | 0.19 (95) | 0.18 (95) | 0.12 (95) | 0.00 (95) | 0.13 (95) | -0.00 (95) |
| dino mean | 0.19 (95) | 0.03 (95) | 0.17 (95) | 0.13 (95) | 0.22 (95) | 0.08 (95) |
| rule object state | 0.19 (48) | 0.19 (48) | 0.43 (48) | 0.38 (48) | 0.33 (48) | 0.02 (48) |
Model-ranking agreement (rule global vs rated narrative coherence, 5 models): Spearman ρ = 0.60, Kendall τ = 0.40.
Rater quality control and reliability (synthetic raters)
| Rater | Items | Gold accuracy | Median s | Duplicate |Δ| | Excluded |
|---|---|---|---|---|---|
| sim-rater-0 | 33 | 100% | 25.5 | 0.33 | no |
| sim-rater-1 | 33 | 100% | 20.6 | 0.67 | no |
| sim-rater-2 | 33 | 100% | 30.0 | 0.83 | no |
| sim-rater-3 | 33 | 100% | 21.4 | 1.17 | no |
| sim-rater-4 | 33 | 100% | 22.7 | 0.33 | no |
| sim-rater-5 | 33 | 100% | 26.8 | 1.00 | no |
| Dimension | Units | Krippendorff α (ordinal) | α (interval) | ICC(2,k) |
|---|---|---|---|---|
| identity consistency | 95 | 0.390 | 0.377 | — |
| object consistency | 95 | 0.370 | 0.409 | — |
| causal coherence | 95 | -0.002 | 0.072 | — |
| narrative coherence | 95 | 0.331 | 0.344 | — |
| visual quality | 95 | 0.047 | 0.031 | — |
| prompt adherence | 95 | 0.251 | 0.226 | — |
Mixed-effects model (narrative coherence ~ model + (1|story) + (1|rater)) fitted on 126 ratings; converged: true.