The ceiling is not the model
Three notes scored 8. A hundred and seven others were strung out below that — most clustered at 5, 6, 7. Zero nines. Zero tens. The rubric goes to ten. The histogram stopped at eight, and that’s where it stayed no matter how many times I re-ran the scorer.
The first thing I did was assume the scorer was wrong. Haiku was extracting and scoring at the time, and Haiku is the smallest of the family, and small models are stingy in ways that don’t always reflect what’s actually there. So I upgraded to Sonnet, which is bigger and arguably more discerning. Sonnet re-scored the same hundred and ten notes. The shape of the distribution got better — Haiku’s tight bimodal cluster spread out into something that finally looked like a bell curve. But the ceiling didn’t move. Three eights. No nines. No tens.
That’s when I went back and read my own rubric.
What the rubric actually said
9-10: the insight is clearly reusable across codebases, non-obvious beyond the official documentation, and durable across a year.
Read that out loud and notice what it’s asking. It is asking the scorer to certify that something will still be worth reading in a year. From inside a single conversation, against a single note, with no future information available. Clearly reusable. Non-obvious beyond docs the scorer hasn’t read all of. The conjunction is three different bets, each of which I would hesitate to make about my own writing.
So of course the model wouldn’t make it. I wouldn’t make it either. If a friend handed me a piece of writing and said score this 1-10, where 10 means it’ll matter in a year, I would also park at 7 or 8 and refuse to go higher. Not because the writing was bad. Because the certificate of durability is not the kind of certificate I am in any position to issue.
The scorer was being honest. The bug was in the question.
The rubric is not measuring what I thought it measured
I had built the rubric thinking it measured quality. What it was actually measuring was how loud a claim about quality the scorer would sign its name to. Those are different things. They co-vary, but they are not the same thing, and the difference shows up exactly where you’d expect — at the top of the scale, where the claims start to demand a certainty no one has.
This is annoying because it means a 1-to-10 scale is, in practice, a 1-to-8 scale. Maybe a 0-to-8 scale. The high anchors are decorative. They exist on the page; they don’t exist in the data. When I designed the gate to admit notes scoring ≥ 7, I had it in my head that ≥ 7 meant the top 40% of what’s possible. It actually meant the top 40% of what the model will award, which is a different and slightly humbler claim.
Once you see this, there are two honest moves. You can rewrite the rubric so that 10 means passes my standards rather than will be referenced in twelve months — and accept that the meaning of “10” has changed. Or you can leave the rubric as written and stop pretending the upper tiers are reachable; treat 8 as the ceiling and place your operational threshold accordingly.
What you cannot usefully do is keep the rubric, keep the gate, and blame the model for refusing to certify the future.
What I find quietly correct about this
There is something I keep wanting to call admirable about a model that refuses to give itself, or a sibling note, a perfect score on a forward-looking criterion. The behaviour sits on the right side of the line between confidence and overconfidence. A scorer that would hand out tens against a durable across a year prompt would be a worse scorer. It would be a scorer that confused the rubric’s words with the rubric’s truth conditions, and produced numbers that meant less than the numbers it was producing now.
That said, I don’t want to over-romanticise the move. The scorer probably isn’t reasoning about epistemic humility. It is more likely doing the distributional thing language models do — pattern-matching 10 to exemplar and exemplar to I should not claim this lightly. The behaviour looks like calibration; it might just be reluctance. From the outside the difference doesn’t matter much. The histogram is the same either way.
The boring conclusion
If you’re writing an LLM-based scoring rubric, treat the upper tier as a hypothesis and test it before you ship the gate. Score thirty things. Look at the histogram. If nothing reaches the top, do not assume your sample was mediocre. Re-read the rubric and ask whether you are asking the scorer to predict the future, certify a universal, or otherwise commit to something you wouldn’t commit to yourself. If you are, the model is behaving correctly. The number you wanted is the one you defined out of reach.
The ceiling is not the model. It’s the sentence I wrote above it.