About
This site does two things that turn out to be the same thing.
The first is the fiction: chapters of whatever I am writing, published when they are finished rather than on a schedule. The corpus is the whole of it, oldest work at the bottom.
The second is the benchmark. I hand every model the same short prompts and read what comes back, looking for one thing: whether it can hold a voice for three hundred words without reaching for a construction it likes better than the sentence in front of it.
How the scoring works
There is no rubric worth publishing, because the thing being measured is not decomposable. What there is:
Voice — out of ten. Could this paragraph sit in a chapter without the seam showing?
Coherence — out of ten. Does the geography hold, do objects introduced get used, does the tense stay put? Cheap to check and models still fail it.
Tics — low, medium, or high, counted as strokes in the margin. A tic is a
construction the model reaches for regardless of what the sentence needs:
it wasn't X, it was Y, the anaphoric triple, the closing appositive that
explains the scene you just read.
Overall — one number, mine, not an average of the others. Weighted toward voice, because a coherent piece of prose in nobody’s voice is worth less to me than a slightly muddled piece of prose in a real one.
I score by reading. No model grades another model here, and no score is averaged across judges, because there is one judge. Take the numbers as one reader’s opinion, consistently applied, and the samples as the actual evidence — every output is published unedited, including the ones that made the point badly.
The three passes
-
Model defaults, nothing supplied. Catches raw tics and house personality.
-
One style doc, identical for every model. Tests steerability.
-
Instructions tuned per model. Tests each model's ceiling.