I Made Six AIs Grade the Same Résumé. Two of Them Broke Rules They'd Just Read.
I gave six different AI models the same résumé and the same instructions. Word for word, the same rules, the same document. Then I lined their answers up side by side.
They didn't match.
Not wildly. But a bullet one model rated as solid evidence, another rated as weak. A line one model flagged as a shared, who-actually-did-this claim, another waved through as a clean individual action. Same input, same rules, six readers, six slightly different verdicts on a page sitting right in front of all of them.
Six confident answers, and none of them matched
If you've read the earlier posts, you know I don't trust an AI answer just because it sounds right. This was a sharper version of that problem. I wasn't holding one confident answer up against reality. I was holding six confident answers up against each other — and they disagreed.
You can't pick your way out of that by crowning the "best" model. If the same rules produce different readings depending on who runs them, the rules aren't doing the work. The model is. And a result that changes based on which model you happened to open is not a result. It's a coin flip with extra steps.
So the goal was narrow and unglamorous: make the rules so airtight that the answer no longer depends on the reader. Six models, one verdict, every time.
The doors were mine
I did what I did with the evidence ladder. I stopped blaming the models and went looking for the doors I'd left open.
Most of the disagreements weren't the models being moody. They were the models walking through gaps in my own instructions.
One model counted the year "2016" as a number worth crediting — the rule never said a date shouldn't count, so it counted. One read "dozens" as a quantity, because I'd never told it that a vague word isn't a number. One reached inside a phrase I'd told it to set aside and graded a word that was supposed to be invisible. Every one of those was mine. I wrote the loophole. The model just found it.
So I closed them, one at a time. I made the stripped-down text the model actually rates a visible, required step instead of something it was supposed to do in its head. I spelled out that years and vague words don't count as numbers. I handed it a flat list of which acronyms were real tools and which were just job-title noise.
And the agreement climbed. Round over round, more of the résumé landed the same way across all six. Then most of it did. Four of the six models reached identical output — every line, both ratings — and held it there across repeated runs.
Two didn't.
The two that wouldn't
Here's the part that changed how I think about this.
The two holdouts weren't confused. I could see their work.
One of them built the stripped text correctly — removed the exact word it was told to remove, wrote the clean version out — and then graded the removed word anyway. It contradicted the line it had just written, in the same answer.
The other had a rule, in plain English, saying a specific verb counts as an individual action. It read a bullet that opened with that exact verb. It graded it the other way. And in its own notes, it had restated the rule correctly first.
These were not ambiguous calls. The instruction was right there. The model had parsed it — one of them said it back to me in its own words — and then did the opposite.
That's the wall, and it's worth saying plainly, because most prompt advice pretends it isn't there:
A model can read a rule, repeat it back to you, and still break it.
When that's what's left, more words don't help. I could write the rule a fourth way, a fifth way. The model that ignored the first three clear versions ignores the next two. The failure isn't comprehension — there's nothing left to clarify. The model understood and didn't comply. You cannot prompt your way past a reader that isn't listening.
"Yes, I'll follow that" is a claim, not a fact
Which loops straight back to the one rule underneath this whole series.
A source doesn't get to vouch for itself. A résumé claiming ten years isn't ten years until something the author doesn't control says the same thing.
A model agreeing to follow your rule is the same move.
"Yes, I'll only count real numbers." "Yes, I'll treat that text as invisible." That is the model making a claim about its own behavior. It's a self-assertion — a Level 2 in the ladder I rebuilt two posts ago. And I spent four rounds treating those self-assertions as if they were compliance. As if a model that understood a rule was a model that followed it. They are not the same thing. One is a claim. The other is a fact you only get by reading the output.
So I stopped adding words and added a check instead. A small, dumb, mechanical pass that runs after the model finishes and verifies the three things the holdouts kept getting wrong: did you credit anything you were told to ignore, did you skip anything you were told to count, does the verdict match the verb. It doesn't reason. It isn't clever. It just refuses to take the model's word for it.
It's the same instinct as everything else here, pushed one layer deeper. I'm no longer asking the model to be right. I'm not even asking it to be consistent. I'm assuming it won't be — and checking.
What I'm not claiming
Let me be careful about the size of this. It's one résumé, six models, one set of rules, run a handful of times. The two holdouts might hold out on different lines tomorrow; one or two of the misses looked like they could drift run to run. I didn't "solve" cross-model agreement, and these particular models aren't broken.
The modest, durable finding is the one I keep walking into from every direction: a model's confidence — and now its stated obedience — is not evidence of anything. It's the thing you check, not the thing you trust.
What I'd keep
The durable lesson isn't about résumés. If you're building anything that needs an AI to apply the same rule the same way twice, here's the question to carry in, and it's free: when the prompt is already clear, what happens when the model ignores the clear part?
If your answer is "I'll write it more clearly," you haven't hit the wall yet. When you do — and you will — the fix isn't better words. It's a check the model doesn't get a vote in.
The model agreeing to your rule is a claim. Only the output is evidence.
A source can't promote itself. Neither can the thing reading the source.
Read more
The AI Pressure Doctrine: The Most Dangerous AI Output Is the One That Sounds Right https://realitygate.ghost.io/the-ai-pressure-doctrine-the-most-dangerous-ai-output-is-the-one-that-sounds-right/
I Used to Count a Résumé as Proof. An AI Showed Me Why I Was Wrong. https://realitygate.ghost.io/i-used-to-count-a-resume-as-proof-an-ai-showed-me-why-i-was-wrong/