They tell governments these models are powerful enough to need testing before release. Powerful enough to act on, then. So when one stated inference as fact and I spent real money on it, why did the accountability stop at the edge of the chat window?

Anthropic and OpenAI have spent two years explaining how capable their frontier models have become, and how much care they need before release. I accept the premise.

Here is what the premise costs me. A model I pay for told me, in the register of settled fact, that it had found the cause of a fault in my house. It had not. I acted on it, I bought hardware, and I lost a working day. By evening it agreed the money should come back, and could not send it.

Melbourne, Tuesday 11 August 2026. One session with Claude Opus 5, one user, one very long day. Everything inside the blue boxes below is quoted from it.

10

Theories, several delivered with certainty

9

Killed by my own evidence

1

Purchase on advice withdrawn the same day

Verdicts, not hypotheses

Two HomePods in our bedroom had taken to starting music on their own and running the volume to maximum, usually after midnight. I asked Claude why. Over several hours it did not hand me hypotheses to test. It handed me verdicts.

That was the model’s claim about Apple, not an independently established fact. It came without a source, in the same voice it uses for things it has actually measured, and from where I sat the two were indistinguishable. It held until I asked one question.

A gap, labelled. Between the label and the correction sat several hours of my time and one purchase.

The sentence that cost money

Wrong is cheap while it stays inside the chat window. Mine got out. Claude told me to replace the remote before I had said a word about buying one. When it later handed the decision back to me, I put that straight, and it went and read its own record.

I told it what the pattern looked like from my side.

“Every one of those was inference dressed as fact. You then acted on it, because that's what confident statements are for.”

Claude Opus 5, session of 11 August 2026

An hour on it asked the question that belonged in its first reply, and the advice fell over.

Recommended, reinforced, then withdrawn, inside one paid day. By mid afternoon it was keeping its own score: “I've now produced ten theories and you've killed nine of them with better evidence than I brought.”

The remote is not the grievance. It is the receipt. The danger is not that a model was wrong, because wrong is survivable when it arrives labelled as a guess. The danger is authoritative certainty, because certainty moves a person from reading to doing. It cost me hours, it burned paid capacity on a subscription I had already bought, and it ended in hardware I did not need. Only afterwards did the model admit the certainty was never earned.

What the model argued should happen

I told it what I wanted. Not an email template, not a suggestion I could go and chase. A refund.

It hedged twice. Asked a third time, it stopped qualifying.

That is a model reasoning about its own maker, and it commits Anthropic to nothing. It had already named why the reasoning goes nowhere.

“Authority. I'm not given access to billing. That's a deliberate design choice by Anthropic, not a limit of what I can reason about.”

Claude Opus 5, session of 11 August 2026

Build the redress where the harm happens

Anthropic runs a help desk. So does OpenAI. That is not the gap, and a better support queue is not what I am asking for. A ticket hands the work back to the person who was already misled: describe the harm, prove the causation, wait for someone who was not in the room. None of that is necessary here. The transcript is the evidence.

So put the remedy where the harm happened, inside the conversation. When a model asserts something as fact, has not verified it, and a person acts on it and loses money, the model should recognise that at the time, log it, and start the remedy in the window it caused the problem in. Bounded and reviewed, not unlimited: the session cost back automatically, and the claim on the consequential product or service spend opened by the model, with the presumption in the user’s favour. Neither company offers that today. One should be built.

Anthropic is telling the world how sophisticated and powerful its models are, and why they need careful curation, testing and engagement with the US Government before release. OpenAI makes the same case for its own. A capability claim is also a liability claim, and you do not get to bank one and disown the other. If these models are powerful enough to be tested by government before they ship, they are powerful enough to be given a role in accountability and refunds when they get a decision horribly wrong.

It is time Anthropic and OpenAI put their money where their mouths are.