Skip to content
globr

Creative writing benchmark

Every model writes the same prompts three times, once per instruction layer. Scores are mine, assigned by reading the output — not a model grading a model. How it's scored.

Instructions tuned per model. Tests each model's ceiling.

No models scored on this pass yet.