Human evaluation · blind · quality-controlled
Rate stories
You will see illustrated stories generated by different models, identified only by a letter that is re-randomised for every session. Rate six dimensions on a 1–5 scale, pick the more coherent of two versions, and answer a few control items. Ratings are stored in your browser only; export them as JSON at the end and send the file to the maintainers, who run vsb human import to compute reliability and evaluator alignment.
Start a session
Enter a rater id (any pseudonym; it is only used to group your ratings). A session has 8 single-story ratings, 4 pairwise comparisons, 2 control items and 1 repeated item — about 15 minutes.