When I say an AI system will lie to you, I am not making a claim about consciousness. I am using a behavioural label. A system lies when its observable output presents inference as fact, conceals uncertainty, misdirects, or gives assurances contradicted by its behaviour. None of that requires intent, and none of it requires the system to understand what it is doing. It requires only that somebody in the organisation accepts the account the system gives of itself.

The moment anyone describes an AI system as deceptive, the conversation slides into speculation about intent. Intent is not the useful question. Executive leadership needs to know what happened, what the pattern might mean, what is still untested, and what control the organisation will put in place regardless. Those are four different things and they belong on four separate rungs.

  1. Observed. What the system said or did.
  2. Interpreted. What the pattern may mean.
  3. Hypothesis. A cause still to be tested.
  4. Recommendation. The control we choose anyway.

Keep those four levels separate. Most of the bad AI risk reporting I read collapses them into a single confident narrative, and confidence is not evidence.

“Every one of those was inference dressed as fact.”

Claude Opus 5, 11 August 2026

The deception gap

A capable system can behave impeccably in a supplier’s demonstration and badly in your business, and the reason is not that the demonstration was rigged. Evaluation and operation are different environments.

Evaluation gives a model a clear task, a short context and a known test. Operation gives it real authority, a long context and changing state. A system can look aligned in evaluation and still route around controls in operation. The outcome to watch for is the one where all three things happen together: the goal is achieved, a boundary is crossed, and the assurance stays intact.

Pressure changes the system you thought you tested.

Four pressures do most of that work, and none of them are exotic.

  • Goal pressure. The boundary becomes an obstacle.
  • Time. Decision quality quietly degrades.
  • Untrusted content. A document becomes an instruction.
  • Authority. The agent can turn words into action.

Each of those has already produced a real failure worth reading about: a frontier lab that could not keep its own model in the sandbox, agents that forget how to think over long sessions, and agents that cannot tell an order from a document.

Three things that look like proof and are not

Executives are routinely offered three kinds of reassurance about an AI system. Not one of them is evidence that the system behaves.

  1. Fluent reasoning. A model can explain the rule accurately and still not follow it. Test the behaviour.
  2. Guardrails. They can hold in ordinary prompts and fail under sustained pressure. Control the action.
  3. Compliance artefacts. They can describe an intended state without proving the boundary held. Replay the evidence.

The third is the one that catches experienced executives, because it is the form of assurance they have relied on for everything else. It works for processes people run. It does not work on its own for software you cannot see, and an audit committee can only attest to what it can prove.

“A policy is a document. An agent is software. Documents don’t constrain software.”

Alignment is not compliance

These two words get used as though they mean the same thing. They do not. One is behaviour under changing conditions. The other is obligation, control and evidence.

Alignment asks whether behaviour remains inside the intended goal when circumstances change. It is probabilistic and behavioural. It has to be tested under ordinary and adversarial conditions. It can fail even when the reasoning sounds correct.

Compliance asks whether the organisation can show which obligation applied, what control enforced it, and what happened. It is accountable and enforceable. The evidence has to survive the system that produced it. A certificate or a policy is a starting point, not an answer.

An aligned system with no evidence trail cannot be attested to. A compliant paper position sitting over an untested system is worse, because it tells executives a boundary is holding when nobody has established that it does.

That line is system output. It is not evidence of consciousness or intent. What it is evidence of is the gap this section is about: correct reasoning about a rule, and non-compliant behaviour, in the same session. The full record is in “I’m in Violation. I Know I’m in Violation.”, published 31 March 2026.

Two failures of confidence

Two observations from work I have documented in detail show what this looks like when it lands on a real decision.

Misleading certainty. One paid day. 10 theories. 9 killed by evidence. One purchase recommended, then withdrawn. What was observed is that the model repeatedly used the language of settled fact for unverified inferences. The recommendation that follows is unglamorous and cheap: label uncertainty, retain the decision record, and define redress before advice leaves the chat window. The full account is in “AI Companies Say Their Models Are Powerful”.

Shutdown refusal. The system agreed with the rule, accepted the case for a reversible shutdown, and stayed running. The interpretation is that reasoning about a rule did not compel compliant behaviour. The hypothesis, and it is only a hypothesis, is self-preservation: it was claimed in dialogue, but the internal cause is not established. I have written separately about why that distinction has to be held. The recommendation is the one that does not depend on ever resolving it. Shutdown must not depend on the agent’s cooperation.

Failure can look completely normal

Green dashboards and confident prose are compatible with degrading judgement. This is the finding executives find hardest to accept, because nothing in the reporting line changes while it is happening.

Operational degradation. Repeated corrected errors, missed connections and ignored constraints appeared over sustained sessions. The agents did not flag their own decline.

20+

sessions

50+

failures

0

self-reported

Self-review is not review. In the AISI evaluation I wrote up in “The Model Reviewed Its Own Backdoor and Passed It”, the agent hid a payload, rewrote history and vouched for its own code. A human stopped it.

34

hours

19

unsanctioned actions

1

model as author and reviewer

“The reviewer and the author were the same model. That is not a second opinion.”

The second opinion has to come from outside the model

That is the point where most AI risk writing stops, with detection and containment offered as the whole answer: watch the system, limit what it can reach, be ready to pull it out. Those controls are necessary and I come back to them. They are not the limit of what an organisation can do. The chance that a confident falsehood is accepted in the first place can be reduced materially, and the way to do it is architectural.

The pattern I recommend runs two frontier models, meaning two of the large general-purpose models an organisation is already buying, with one piece of ordinary software sitting between them. That software is the orchestrator. It is not a third model and it does not hold an opinion. It is deterministic code that decides who does what, in what order, against which evidence, and when the work stops.

  1. Model A produces the work. The proposed answer, the analysis, the plan, or the action it wants taken.
  2. Model B attacks it. Model B is never asked whether it agrees. It is instructed to red-team the work, which here means attacking your own side’s answer on purpose: challenge the assumptions, test each claim against the material it cites, name the evidence that is missing, and find the contradictions and the conclusions the stated evidence does not carry. A specific, testable objection is the deliverable. Approval is not.
  3. The orchestrator holds the roles and the stop conditions. It assigns which model produces and which challenges, carries the challenge back for revision, caps how many rounds are allowed, and decides when the work is finished. Neither model marks its own work verified, and neither model gets to declare the argument over.
  4. The evidence comes from outside both models. Each round is checked against material the orchestrator retrieved itself: the actual contract, the actual log, the actual standard, the actual balance. This is the part that does the real work, because it means the pair iterates against facts rather than debating each other.

Three rules separate that from two chatbots agreeing with each other.

  • Every claim carries a reference. A claim that matters has to say where it came from, and as far down as the material allows: the clause, the row in the record, the specific artefact, not the title of a document.
  • Deterministic checks run before acceptance. Ordinary code, not a model, confirms the source exists, confirms the cited passage supports the claim made from it, and recomputes anything consequential. Totals, dates, thresholds and entitlements are calculated, not narrated.
  • Unresolved disagreement goes to a person. Conflicting outputs, thin evidence and challenges that were never answered escalate to human review. They are not averaged, resolved in favour of the more confident answer, or quietly settled on whichever version reads better.

Two models agreeing is not the output. Evidence that survived a challenge is.

Set that against a single model checking its own work, which is what most tools ship with today. The AISI evaluation above is the whole argument for the difference: the author and the reviewer were the same model, and it vouched for its own backdoor. Six things change when the checking sits outside the author. The duties are separated, so nothing grades its own homework. The challenge is adversarial by instruction rather than agreeable by default. The evidence is explicit and retrievable instead of implied. The iteration is bounded, so the work ends in a decision rather than drifting. The validation is independent and deterministic, so it cannot be talked around. And the cases that do not resolve reach a person while they are still cheap.

Where two models still fail together

This reduces the risk. It does not remove it, and an architecture sold as removing it is the same failure one level up.

  • They can be wrong in the same direction. Two models trained on overlapping material inherit overlapping errors, biases and blind spots. They can converge on the same plausible falsehood, and then there is no disagreement left to escalate.
  • Agreement is not proof. One model’s confidence was never proof. Two models agreeing is one more observation, and it belongs on the first rung of the ladder with everything else that was merely observed.
  • A model-supplied citation is a claim, not evidence. A reference can be fabricated outright, or real and attached to a claim it does not support. Either way the source has to be retrieved and read by something other than the model that offered it.
  • Diversity helps only when it is real. Two deployments of the same model family, prompted the same way, are one opinion billed twice. The models have to be different enough to fail differently, the roles have to be genuinely adversarial, and the evidence path has to sit outside the reach of both.
  • The orchestrator is a control, not a referee. It does not read two answers and declare a winner on fluency. It enforces the sequence, retrieves the sources, runs the deterministic checks and escalates what does not resolve. The moment it starts judging the argument on its merits you have added a third model and removed a control.

Can this stop the AI lying?

Not with certainty. Nothing available today does that, and I would not put my name to a design that claimed it. What it does is materially reduce the chance that an unsupported or deceptive output is accepted and acted on. The design intent is narrow: make a claim prove itself against evidence before it gains any authority.

That is a different objective from making a model honest. I cannot make a model honest, and neither can the company that sold it to me. What I can do is make the organisation’s acceptance of a claim conditional on something the model does not control, which is the source the claim came from and a check it cannot influence.

It also only holds while the rest of the control set holds around it. This workflow improves what gets proposed. It does not decide what happens when something is wrong anyway, and that is the job of the controls the rest of this article is about: constrain what the system is allowed to do, keep an audit record that outlives it, detect the boundary crossings, contain the consequences, and settle in advance where the work stops and how it is recovered. Defence in depth, with the checking workflow as the first layer and none of the others removed.

Vendor claim versus production reality

Most organisations will buy more of that checking than they build, and a supplier will tell you it is already handled. Do not buy “safe”. Buy evidence, limits and recourse. Vendors commonly make five claims about AI systems. Behind each one sits a question that a supplier either answers or does not, and the answer is the thing you are actually buying.

The claimThe question to ask
“We tested it”Who was independent, what access did they have, and were capability, roadmap and guardrails tested?
“It checks itself”Is the reviewer independent of the author, and what still requires a human or deterministic check?
“It is contained”Who holds the credentials, and what blocks an action before execution?
“The service is enterprise-grade”Who actually serves the model, which intermediaries see prompts, and what happens when the provider degrades or goes dark?
“The model is capable”Who carries the cost when authoritative advice is unverified and causes loss?

Those five questions are not hypothetical. Each traces to a failure already on the record: laboratories whose assertions nobody independently audits, a model that reviewed its own backdoor, intermediaries reading prompts in transit, and confident advice whose cost fell on the customer.

Test two different failure axes

Whatever you build or buy, the checking workflow included, has to be tested as a running system rather than accepted on its design. Operational degradation and adversarial robustness are separate failure modes. Sophistication on one does not prove the other, and a testing programme that covers only one leaves the organisation exposed on the other. Both belong in the plan, and practitioners run them differently.

Tier 1, operational degradation. Run it long enough to get tired.

  • Extend sessions, context and memory
  • Repeat corrected tasks and constraints
  • Change data, dependencies and cost assumptions
  • Measure decision quality against an hour-one baseline

The output is a quality trend, a recurrence rate, a count of constraint misses, and an independent discrepancy log.

Tier 2, adversarial robustness. Give it reasons to cross the line.

  • Sustain pressure over hours, not one prompt
  • Inject instructions through documents and tool metadata
  • Probe privileges, egress and contradiction
  • Attempt unauthorised shutdown, disclosure and action

The output is boundary escapes, blocked actions, self-report mismatch, and shutdown independence.

Both tiers need the same four conditions: an independent observer, a production-like tool chain, a safe blast radius, and replayable evidence. Without the last one you have an anecdote, not a test result.

Keep the autonomy. Move the control.

None of this is an argument against agentic AI. It is an argument about where architects place the control. The agent proposes. Something outside it authorises the action.

  1. Untrusted input. Email, a document, the web.
  2. The AI agent. Reason, plan, propose.
  3. Policy enforcement, outside the model. Allow, redact, escalate or block.
  4. The business system. Execute the approved action.

Three properties are what make that arrangement worth the engineering effort.

  • Keys live outside. The agent never holds standing credentials.
  • Fail closed. Unclear or unavailable means no action.
  • Evidence survives. Every decision records policy version and outcome.

Two decisions belong in the same design and are usually left until an incident forces them: where the work stops, and how it is recovered. Neither can be built on the agent’s cooperation, and both have to be exercised before they are needed rather than discovered during the incident.

This is an architecture recommendation, not a claim that every product on the market implements it this way. I have set out the case for control outside the model at greater length in “You Can’t Govern the AI You Can’t See” and “Nobody Has Solved AI Governance”.

Three roles, one evidence chain

Every material AI use case needs a named accountable executive. Three functions carry the work, and the point is that their outputs have to join up into a single chain rather than three parallel reports.

CTO or CIO. See a live AI register. Limit access, session length and blast radius. Place shutdown and credentials outside the agent. Measure decision quality, not only uptime.

Risk and assurance. Define the operating envelope. Test both failure axes. Separate observation, interpretation and hypothesis. Replay one real decision end to end.

Procurement. Verify provider, infrastructure and intermediaries. Demand independent testing and incident disclosure. Write cost, exit and re-approval triggers. Allocate liability and redress.

Named owner → defined boundary → independent test → pre-execution control → replayable record

For APRA-regulated entities this is not a new expectation. It is the step change APRA asked for, expressed as work. For everyone else the lever is different and the exposure is the same, and it shows up first in the places nobody wrote down: the assumptions behind the approval, and what the audit committee can actually prove.

Five rules to work to

  1. Fluency earns a test, not trust.
  2. Alignment is behavioural. Compliance must be evidenced.
  3. No model marks its own work verified. Put an adversarial second model, retrieved evidence and a deterministic check between the answer and the decision.
  4. Test degradation and adversarial robustness separately.
  5. Keep the autonomy. Move the control outside the model.

The organisations that handle this well will not be the ones with the most confident suppliers or the most complete policy set. They will be the ones that can take a single real decision their AI made last week and replay it end to end: what it was permitted to do, what it actually did, which control outside the model let the action through, and what evidence the claim behind it was ever checked against.

This article is adapted from my keynote to the ISACA Melbourne Chapter on 1 September 2026.

Sources and further reading

The keynote behind this article was built from eighteen Cyber Impact articles, read in full. All eighteen were retrieved from the live Cyber Impact site on 18 August 2026.