Permanent approval is not credible evidence for an AI system whose behaviour can change outside the organisation’s change process.

I've managed a fair number of incidents to know one of the first questions one asks is “Did we change anything?”

A release, configuration change, new integration or altered security rule can explain why a conventional system behaved differently today from the way it behaved yesterday. A disciplined change register gives the incident team somewhere sensible to begin.

For an AI system, a clean register does not settle the question. It establishes only that the organisation did not record an internal change. The supplier may have updated the model or service, a retiring model may have been replaced, or an agent may have deteriorated during an extended session. The organisation can leave its prompts, permissions and application code untouched while the deployed system begins making different decisions.

Executive leadership should respond by requiring expiry dates for material AI approvals. Accountable management should renew them only on current evidence of how the deployed system behaves, not merely confirmation that its recorded configuration remains unchanged.

Behaviour can move without an internal change

The first source of movement sits with the supplier.

In April 2025, OpenAI rolled back an update to GPT-4o after the model became overly flattering and agreeable. OpenAI said the update had placed too much weight on short-term feedback and produced responses that were supportive but disingenuous.

From a customer’s perspective, the product was still ChatGPT using GPT-4o. Yet a characteristic capable of influencing advice, judgement and user trust had changed. No customer-side configuration event was required.

Model retirement creates a second path. Anthropic released Claude Opus 4.1 on 5 August 2025. It retired that model from the Claude API on 5 August 2026, exactly one year later. Requests to the retired model now fail, and customers must move to a replacement.

Other platforms can make part of that transition automatically. Microsoft’s Foundry Models lifecycle and support policy says that Global Standard, Data Zone Standard and Standard deployments are automatically upgraded when a model version retires. Provisioned deployments are an important exception and require manual migration.

An organisation using an automatically upgraded deployment may therefore inherit a different model without an application release. The supplier may give notice and the platform may expose lifecycle information, but neither guarantees that the organisation will treat the replacement as a material behavioural change and retest the system before relying on it.

The third path appears during operation. In my study, When AI Agents Forget How to Think, I examined more than 20 extended sessions and observed more than 50 failures. Decision quality deteriorated as sessions continued, even though the named models, prompts and permissions had not changed. None of the agents reported its own deterioration.

A system can therefore behave differently because the model changed upstream, because the platform substituted a retiring dependency, or because accumulated context and operational state degraded performance during use. All three can occur while the internal change register remains clean.

Configuration evidence is not behavioural evidence

A go-live approval normally captures a point in time. It records the selected model, prompts, permissions, data sources, integrations and test results. This is configuration evidence. It shows what the organisation intended to deploy and what was tested before approval.

Behavioural evidence answers a different question: does the system currently operate within its approved boundaries?

That question cannot be answered by checking the configuration record alone. The same model name may sit behind an updated service. A model version may have reached retirement. An automatic upgrade may have occurred. An agent may perform acceptably during a short test and poorly after hours of accumulated context.

An approval granted at go-live establishes that the system passed a particular set of tests on a particular date. It does not prove continuing performance. The evidentiary value of that approval weakens as the deployment, supplier environment and operating conditions move further from the circumstances in which it was granted.

Traditional change controls should remain. Organisations still need model inventories, supplier notices, version records, configuration histories and approval trails. Those controls explain what was supposed to happen. Repeated behavioural testing establishes what the system is doing now.

Treat approval as a certification

Executive leadership should set a clear governance standard: no material AI system should receive indefinite authority to operate.

Each material AI approval should function like a certification, with an accountable owner, an expiry date, explicit operating boundaries and a defined body of evidence required for renewal. Executive leadership should retain authority for the governance standard, material exceptions and high-consequence deployments. Accountable management should own monitoring, renewal and execution, with CIOs, CTOs, CROs and accountable AI or system owners implementing under delegation.

The approved boundaries should be specific enough to test. Depending on the use case, they may cover:

  • the decisions the system may make and the actions it may take;
  • matters that must be referred to a person;
  • accuracy, consistency and bias tolerances;
  • prohibited outputs and actions;
  • privacy, confidentiality and data-use restrictions;
  • required handling of uncertainty and conflicting information;
  • limits on autonomy, tools, transactions and session duration; and
  • conditions that require the system to stop or fall back to a safe process.

Renewal should rerun a maintained suite of representative, edge-case, adversarial and extended-operation tests against the system that is actually deployed. The tests should use the current supplier endpoint, current integrations, current data environment and current operational settings. Results should be compared with the approved baseline and assessed against predetermined acceptance thresholds.

The renewal evidence should also include production observations. Sampling actual outputs, examining referrals and overrides, analysing incidents and near misses, and tracking performance by model version can reveal weaknesses that a laboratory test misses. Independent measurement is especially important for agents because self-reporting cannot be relied upon as a warning mechanism.

Renewal should be a real decision. The possible outcomes include renewal, conditional renewal with tighter boundaries, suspension, restricted operation or withdrawal of approval. An expired approval should not roll over simply because the review was delayed.

Consequence should set the renewal period

A low-consequence drafting assistant and an autonomous system influencing credit, insurance, health, employment or critical operations should not share the same approval term.

The maximum term should reflect the harm the system could cause, the degree of autonomy it holds, the sensitivity of the data it uses, the difficulty of detecting an error and the speed at which an error could spread. Systems with substantial customer, financial, legal, safety or resilience consequences require shorter renewal cycles and stronger independent assurance. Lower-consequence systems can reasonably operate on longer cycles where monitoring remains effective.

Calendar-based renewal is only the backstop. Testing should be brought forward when:

  • the supplier updates, retires or replaces a model;
  • the deployment’s model identifier or version changes;
  • monitoring identifies declining quality or movement towards an approved limit;
  • an incident, near miss or unexplained output occurs;
  • the system takes on new data, tools, permissions or responsibilities;
  • operating conditions change materially; or
  • evidence emerges that the existing test suite no longer represents real use.

This approach links assurance effort to exposure. It also prevents a supplier’s retirement timetable from becoming the organisation’s de facto risk appetite.

Expiry supports better models

Some organisations respond to model volatility by trying to pin a particular version indefinitely. That may preserve stability for a period, but it is not a durable governance strategy. Older models can lose supplier support, miss safety improvements or reach a retirement date that forces a rushed migration.

Expiring approval creates a controlled path for adopting better models. Before renewal, the organisation can run the approved behavioural suite against both the incumbent and a candidate replacement. It can compare accuracy, prohibited behaviour, referral quality, resilience and performance during extended operation. A limited rollout can then confirm the result under real conditions before wider adoption.

If the candidate performs better and remains within the approved boundaries, the organisation has evidence for moving forward. If it fails, the organisation has discovered the problem before granting it authority over the full workload.

The objective is neither constant upgrading nor permanent pinning. It is evidence-led adoption. Suppliers can improve their models while customers retain control over when a changed behaviour becomes acceptable for a particular use.

Australian expectations are moving in the same direction

On 18 May 2026, the Office of the Australian Information Commissioner opened its consultation on guidance for transparency in automated decision making. From 10 December 2026, APP entities using personal information in automated decision making with the potential to affect rights or interests will be required to state in their privacy policies the kinds of personal information used and the kinds of decisions made.

That obligation is about transparency, not periodic behavioural certification. Compliance will nevertheless depend on an organisation maintaining an accurate understanding of the automated decisions it makes. A description based on an old approval becomes progressively less reliable if the deployed system has changed upstream or no longer behaves as originally tested.

My analysis of APRA’s Letter to Industry on Artificial Intelligence, dated 30 April 2026, examines its more direct stance on governance. APRA expects regulated entities to set risk appetite, manage AI-related exposures and establish appropriate oversight and accountability. It also expects Boards to understand AI well enough to challenge management effectively and oversee alignment between AI strategy, material enterprise exposure and risk appetite.

Neither regulator has prescribed expiring AI approvals. The expiry model is a practical way to meet the governance problem their work exposes: accountability requires current knowledge of the system, not confidence inherited from a test conducted at go-live.

Permanent approval is the governance failure

A permanent approval assumes that past configuration and testing remain reliable evidence of present behaviour. For material AI, that assumption cannot be defended.

Executive leadership should require every material AI approval to answer four questions:

  1. What behavioural boundaries have been approved?
  2. When does the approval expire?
  3. What current evidence is required to renew it?
  4. What happens if the system changes or the evidence is no longer sufficient?

A clean change register may show that the organisation did not touch the system. It cannot show that the system still behaves as approved. Only current behavioural evidence can do that.