Creative writing benchmark
Every model writes the same prompts three times, once per instruction layer. Scores are mine, assigned by reading the output — not a model grading a model. How it's scored.
One style doc, identical for every model. Tests steerability.
No models scored on this pass yet.