I Aimed My Own Lie-Detector at My Own Rulebook
Fourth in a series on the AI Pressure Doctrine. Earlier posts showed AI models disagreeing on the same evidence, found the rule that fixed it, and then watched that rule get fooled by a forgery wearing the right costume. This one turns the whole apparatus inward — on the rulebook itself.
I pressure-test everything I build. That's the whole project — trust only what survives pressure. So this post isn't about discovering I'd skipped a step. It's about what happened when the thing I turned the pressure on was a document the pressure itself depends on.
Here's how I got there, and what I learned when the rulebook had to survive its own rule.
How this started: getting organized before the mess got worse
I'd reached two solid documents — a doctrine and the workflow that operated it. They were good. But I could see where this was heading: a few more months of this and I'd have a pile of overlapping files, rules, and doctrines with no structure holding them together. Better to impose order now, while it was still two documents, than to excavate it later out of thirty.
I'd studied for a security certification, and it teaches a specific way to layer a governance program: a Vision at the top (what you believe), then Policy (the mandatory rules), then Standards (the measurable benchmarks), then Baselines, Procedures, and Guidelines underneath. Each layer answers to the one above it. Every rule traces up to a principle; every principle can be pointed at when a rule gets challenged. That's not bureaucracy for its own sake — it's what lets a disagreement bottom out somewhere instead of going in circles.
My documents weren't layered like that. One of them literally described itself as three things at once — doctrine, workflow, and decision protocol — all in a single cover. That's a normal way for a working document to grow, but it can't function as a program that way. So I started splitting them into the proper layers.
The split immediately exposed a hole. I had plenty of rules, but I had never written down the beliefs the rules came from. The top of the pyramid — the Vision — was empty.
So I wrote it. Six principles, lifted from the foundations I already had:
- Fluent language can hide weak evidence — structure is not proof.
- Consensus is not evidence — agreement amplifies errors as easily as it validates truth.
- Polished structure creates false confidence — the better it looks, the more scrutiny it deserves.
- Humans stop thinking once labels are applied — resist premature closure.
- Operational failure appears before conceptual failure — the field breaks before the theory.
- Human behavior breaks systems faster than technology — the human loop is the highest-risk node.
That's the Charter. It's the thing every other document is supposed to answer to. And the moment I finished it, the doctrine I'd been preaching turned around and pointed at it.
The uncomfortable question
Here is the function a Vision document is supposed to perform: when two qualified people look at the same evidence and disagree about whether it's good enough, the Charter is where that argument ends. Not in opinion — in a shared belief about what evidence is. If the Charter can't do that, it's decoration. Nice words, no teeth.
So the test of the Charter isn't "is it well-written." It's "can it actually settle a dispute." And a reviewer I trust looked at my six principles and said, essentially: no, it can't — and you can't fix it by ranking them.
The argument was sharp. When a real case puts two principles in tension — say, a polished official document that says one thing and a pile of messy field evidence that says the opposite — Principle 3 ("distrust the polish") and Principle 1 ("the messy stuff isn't proof either") pull in opposite directions. The reviewer's claim was that two reasonable people would resolve that tension differently, not because my wording was vague, but because they genuinely weight the principles differently based on who they are. A forensic investigator trusts the operational signal. A compliance auditor trusts the procedural finish. Same principles, opposite verdicts. And no amount of wordsmithing closes that gap, because the gap isn't about words — it's about values.
If that's true, my Charter has no floor. The disputes it's supposed to end would just relocate to which principle you happen to favor.
I could have argued back. Instead I did the thing the doctrine demands: I stopped trusting the confident claim — mine and the reviewer's — and built a test.
The method, because the method is the point
Here's where I have to show the machinery, because the whole credibility of what follows depends on the trap being real.
The reviewer's claim and my hope were two different predictions about the same thing, so I wrote them down as competing hypotheses before running anything:
- The principles converge. Independent reasoners, given only the six principles and a genuine conflict, reach the same decision. (My hope.)
- The principles split on collision. They reach different decisions, and the split traces to two principles in direct opposition with no tiebreaker. (The reviewer's prediction — the fatal one.)
- The principles split on wording. They reach different decisions, but only because they read an individual principle differently. (The fixable-but-real one.)
The trick is that you cannot tell these apart by looking at the decision. A split is a split. What separates "the principles collide" from "the wording is fuzzy" is the reasoning — so I made each reasoner report not just its verdict, but which principle drove it and which principle fought against it. The reasoning is the evidence; the verdict is just the headline.
Then I built two conflict cases. One was a head-on clash: a polished, signed, official "PASS" report produced by the team being audited, against a messy pile of raw scans and unsigned emails showing failures. The other was subtler — a vendor's encryption claim backed by three "independent" sources that, on close reading, all traced back to the vendor's own datasheet. A consensus that was really one voice in a trench coat.
I ran both cases through three different AI models from three different companies, each told to reason only from the six principles, each in a clean session with no memory of the others. Three labs, so no single model's habits could carry the result. And one guardrail I'd learned the hard way: I ran them on raw interfaces, not through a third-party tool — because earlier the same week I'd caught a popular wrapper silently injecting its own formatting instructions that overrode what the user explicitly asked for. If the plumbing can override your instructions, the plumbing is part of the experiment. I kept it out.
That's the trap. If the reviewer was right, I'd see the models split on the same collision, resolving it opposite ways. I genuinely did not know which way it would go.
What happened
They didn't split.
On both cases, all three models reached the same decision. On the head-on clash, every one of them rejected the polished official report and sided with the messy field evidence — nobody saluted the letterhead. On the trench-coat consensus, all three independently refused to call the claim confirmed, every single one naming the same principle as the reason: agreement that traces to one source is not agreement. On that second case they didn't just reach the same verdict — they reached it through the same principle. That's about as clean as this kind of result gets.
The reviewer's fatal prediction — the split with no tiebreaker — did not appear. Not once.
So the Charter survived. The principles, faced with two real conflicts, converged. I don't need to add a tiebreaker rule, and I especially don't need to mutate my six beliefs into a ranked algorithm, which is what "just rank the principles" would have forced me to do.
That was the result I wanted. Which is exactly when I get nervous.
The part where I argue against my own win
A result you wanted is the most dangerous kind, because you stop checking it. So I ran the finding back past the same reviewer, and they found the hole — and it's a real one.
My "head-on clash" wasn't actually head-on. In that case, I'd intended Principle 1 and Principle 3 to collide. But a third principle — number 5, the field breaks before the theory — quietly reinforced one side. The messy operational evidence wasn't just "messy"; it was the field, and Principle 5 says the field's failure shows up before the polished theory admits it. So the models had an escape hatch. They weren't cornered into resolving a true tie; one principle broke it for them.
Which means my result proves something narrower than I first wanted to claim. It proves the principles converge when a tiebreaker exists. It does not prove they'd converge in a genuinely tie-less collision — a case engineered so that two principles point at opposite verdicts with no third principle available to rescue the decision. I never built that case. It might still split.
And there's a second hole, bigger than the first, and it's one my own Charter named. Principle 6 says the human is the highest-risk node — the tired, rushed, role-bound person is where systems actually break. But I tested the Charter on AI models, which have no career, no professional identity, no incentive to weight one principle over another, and no bad afternoon. The reviewer's original concern was about human value-pluralism, and I validated against non-humans. I tested the principles' logic. I did not test them against the exact failure mode the principles themselves call supreme.
So here's the honest scorecard. The Charter survived a real test: its principles resist being talked into trusting a polished lie or a circular consensus, and three independent machines agree on that. That's not nothing — it means the Charter can't easily be weaponized to justify bad evidence, which is a real property to have. But it survived a bounded test. The hardest case and the most important reasoner — a true tie, and a human — are still on the to-do list, logged honestly as the next things to break.
Why I'm telling you this instead of just shipping the rulebook
Because the doctrine isn't "trust what I built." It's "trust only what survives pressure" — and that has to include the rulebook, or the whole thing is a hypocrite.
It would have been easy to write the Charter, declare it sound because it reads well, and move on. That move — it's polished, therefore it's proof — is the precise thing Principle 3 exists to forbid. Applying the doctrine to everything except the document that states the doctrine would have been the most embarrassing possible failure. So I aimed it inward, set a trap I might have failed, found the principles convergent, and then refused to overclaim the win the moment I had it.
That's the actual product here. Not a rulebook that's beyond question — a rulebook that has been questioned, by its own rules, and that comes with a written list of where it hasn't yet been pushed hard enough. A framework that names its own untested gaps is worth more than one that claims it has none.
The doctrine survived contact with itself. Bounded, caveated, and with homework still due — but it survived. And if I ever catch myself trusting it because it's mine rather than because it's been pressured, I hope someone points at Principle 3 and tells me to check the costume.