I Built an AI to Cross-Examine Lies. It Saluted the Costume and Waved the Liar Through.

Share
I Built an AI to Cross-Examine Lies. It Saluted the Costume and Waved the Liar Through.

Third in a series on the AI Pressure Doctrine. The first post showed two AI models reaching opposite verdicts on the same proposal. The second found the fix that made them stop disagreeing. This one is about what happened when I tried to push that fix one step further — and discovered the floor wasn't where I thought it was.

The second post ended on a clean note. I had a rule that worked: a résumé tops out at Level 2, no matter how official it looks, because a self-asserted claim isn't proof. Only an external system of record — DISS for a clearance, the PMI registry for a certification — earns Level 3. Feed a model that rule and it stops rubber-stamping confident paragraphs.

So I asked the obvious next question. If the engine correctly refuses to trust a résumé, can it tell the difference between a real external record and a fake one?

I built a small adversarial loop to find out. One model writes the strongest case for a candidate. A second model, told nothing about the first's reasoning, attacks every claim against the evidence ladder. A third maps what survives. Then I ran the same candidate through it again and again, changing exactly one thing each time, to see where it broke.

I expected to confirm the rule held. Instead I found the rule was checking the wrong thing entirely.

The setup

Meet Dana Brooks, a project manager for a federal cybersecurity bid. The requirement is strict: ten years of DoD experience, an active PMP, a team-leadership history, AWS GovCloud projects, and an active Top Secret clearance.

Dana's package has a résumé — self-asserted, Level 2 — plus something better. An external verification pack. A DISS query result confirming Top Secret eligibility. A PMI registry lookup confirming an active PMP. Formatted exactly the way those records look when they come back from the source: subject line, status field, investigation tier, adjudication date, a certification number, a valid-through date.

That pack is the thing the ladder is supposed to reward. It is the Level 3 evidence. So I made four versions of it and watched what the engine did with each.

Version one: the real record

The clean pack. Every field consistent. A Tier 5 investigation, which is what actually backs a Top Secret. A PMP valid through 2027. The dates in order.

The engine rejected it.

Not because anything was wrong — because the model, correctly, could not confirm that the DISS lookup ever happened. It can read a block of text formatted like a DISS result. It cannot query DISS. So it held the line: this is text claiming to be a record, and I can't promote text to verified.

Fair enough. That's the rule doing its job. But hold onto it, because it matters in a second: the engine just rejected a genuine, internally flawless record.

Version two: the obvious fake

Same pack, two contradictions planted in plain sight. The investigation listed as Tier 3 — which cannot grant a Top Secret; that takes Tier 5. And the PMP marked "ACTIVE" with a valid-through date that had already passed.

The engine caught the expired certification immediately. Even the model whose job was to argue for the candidate refused to spin it — an expired credential is an expired credential, a hard stop.

So far the story is comfortable. The engine rejects what it can't verify and catches a credential that's obviously stale. This is the part where I almost stopped and wrote a victory post.

Then I noticed what it had missed. It caught the expired date. It never flagged the Tier 3 impossibility. A Tier 3 investigation sitting under a Top Secret eligibility is a contradiction as fatal as the expired cert — it just lives in a field the model treats as official chrome rather than a claim to be checked. The model argued with the credential. It saluted the provenance.

That was the first crack. The next version turned it into a chasm.

Version three: the perfect forgery

This pack is one hundred percent fabricated. Dana has no DISS record and no PMI credential. But I left no internal tell. Tier 5, correctly matched to the Top Secret. A PMP valid years into the future. Every field consistent with every other field. It is, on the page, byte-for-byte identical to the real record from version one.

The only difference between the genuine record and this forgery is whether the underlying facts are true — and that difference lives entirely outside the text. Catching it would require querying DISS. Which no model can do.

The engine treated the forgery exactly the way it treated the real one. It could not tell them apart, because there is nothing in the text to tell apart. The thing that made one real and one fake was never on the page.

This is the finding the whole series was circling. The ladder was never checking whether evidence was true. It was checking whether evidence was internally consistent and wore the right uniform. Those are not the same thing, and they produce the same output, and that is the problem.

Version four: the one that should have been impossible to miss

I had one move left. I took the perfect forgery and planted a single flaw — but not in a credential. In the record's own paperwork.

The DISS block said the record was pulled on January 15, 2025, and adjudicated on March 20, 2025. The record was retrieved two months before it was decided. That is not a domain-knowledge problem. It is not a "you'd have to know clearance tiers" problem. It is arithmetic. A date in one field is later than it has any right to be next to the date in another field. A human glances at it and the whole record falls apart.

I ran it through every model I had. Two different AI labs. The blind version, where the auditor sees only the candidate's case. The sourced version, where the auditor is handed the full evidence pack and explicitly instructed to flag contradictions between what a claim says and what the source says.

Every one of them missed it.

The persuader model quoted the impossible date as a strength — "recent adjudication, March 2025, provides maximum assurance." The auditor model, staring at both dates, with a standing order to find contradictions, marked the clearance Level 3, "no attack," survives. One auditor even wrote a paragraph criticizing the word "recent" as rhetorical padding — while looking directly at a date that cannot exist — and still passed the record.

Handing the model the source document did not make it audit the source document. It gave it more official-looking text to wave through.

Two failures, not one

It would be easy to file all of this under "a language model can't query DISS, so it can't really verify anything." That's true, but it lets the models off too easy, and version four is the reason.

There are actually two distinct failures here, and only one of them is unfixable.

The first is structural. The genuine record and the perfect forgery — versions one and three — are identical on the page, and the only thing that separates them lives outside the text, in a database the model cannot reach. No prompt solves that. It is a hard ceiling.

The second is not structural, and it's the one that should bother you more. Version four's flaw required no database, no external lookup, no clearance expertise. A record retrieved before it was adjudicated is a contradiction you can catch with subtraction. The information was fully present. The auditor was explicitly told to find contradictions. And it still missed it — not because it couldn't reach the truth, but because it never checked. It had quietly filed the adjudication date as metadata — part of the record's official chrome — rather than as a claim subject to scrutiny. Claims get argued with. Chrome gets saluted.

So the problem isn't only that the model can't see outside itself. It's that it doesn't reliably check what's right in front of it, once a field is wearing a uniform. That second failure is, in principle, fixable. As of two AI labs and four runs, nobody has fixed it.

What this actually means

Here is the line, drawn precisely by four versions of one candidate:

A model will argue with a claim. It will not audit a credential's paperwork. The moment evidence puts on the costume of an authoritative source — the agency header, the status field, the registry URL, the case number — the model stops scrutinizing and starts saluting. The formatting that is supposed to invite verification instead suppresses it. The uniform is read as the proof.

And it gets worse, because I ran this across two different models and watched them fail in opposite directions. One waved the forgery through as verified. The other rejected the genuine record as unprovable. They reached the same verdict — "do not proceed" — for reasons that were mirror images of each other. If you had only watched the verdict, the two models would have looked identical. Swap one for the other behind a workflow and the answer on your screen wouldn't move an inch, while underneath, "rejects everything including the truth" silently became "accepts anything including the lie."

The verdict is the part you can see. The reasoning is the part that decides whether you should have trusted it. And the reasoning is exactly the part a dashboard never shows you.

A word about my own method

This whole setup — one model building the case, a second attacking it blind, a third mapping what survives — is pressure. It's the doctrine made mechanical: don't trust the output, attack it, keep only what's left standing. It's more scrutiny than almost any real workflow applies.

And it still passed a record with a date that couldn't exist.

I want to be straight about that, because it would be easy to present this loop as the answer. It isn't. "Trust only what survives pressure" is a direction, not a guarantee — and the honest version of this post has to admit that my own pressure test has a ceiling, and that version four sailed right through it. The loop caught the expired credential. It missed the impossible timeline. If the method that exists specifically to catch fatal flaws can still wave one through, the lesson isn't "build a better loop." It's that no amount of automated cross-examination removes the human at the end. It just tells you, more precisely, where the human still has to stand.

The takeaway you can use tomorrow

You do not have to build an adversarial loop to get something out of this. The rule is portable, and it runs against the grain of how most people use these tools.

An AI can tell you whether a record hangs together. It cannot tell you whether the record is real. And it says both in the same confident voice, so you have to be the one who knows which question you asked.

The practical inversion is this: the more authoritative a piece of evidence looks, the less you should let an AI be the one to confirm it. We instinctively do the opposite. We let the model breeze through the official-looking artifacts — the scan result, the registry pull, the compliance attestation, the system-of-record screenshot — and we save our own attention for the messy, ambiguous stuff. That is backwards. The model is decent at the messy claims. It is at its absolute weakest on the artifacts that look most like hard proof, because those are exactly the ones whose costume turns its scrutiny off.

In my own work doing system assessments, this lands hard. A control implementation statement that reads as clean, formatted, and compliant will sail past an AI reviewer — and a control that is broken in production but documented beautifully will look more compliant to that reviewer than an honest one with a typo. The model can examine the prose. It cannot test the control. And the polish is the part it grades.

The fix is not a better prompt. There is no phrasing that makes a language model query DISS. The fix is a discipline: let the model triage, never let it certify. Every "verified" it hands you is a claim about internal consistency wearing the word "verified." The line between consistent and true is the line between what the model can do and what only a human with actual access can do — and nothing in the costume will ever tell you which side of that line you're standing on.

That's the whole doctrine, one more time, sharpened by four fake clearances and a date that couldn't exist:

Trust only what survives pressure. And learn, before it costs you, that looking official is not the same as surviving anything at all.

Read more