Creative writing benchmark
Every model writes the same prompts three times, once per instruction layer. Scores are mine, assigned by reading the output — not a model grading a model. How it's scored.
Model defaults, nothing supplied. Catches raw tics and house personality.
No models scored on this pass yet.