Knowledge Centre
AI AuditAI··4 min read

What is an AI Audit?

The uncomfortable truth about an AI audit is that the model is the easy part. The weights can be produced, the code can be read, the metrics can be rerun. What an auditor actually wants is something most organisations cannot produce on demand: the record of how the system came to be trusted, and of what it has done with that trust since.

The four questions an auditor asks

Strip away the frameworks and an AI audit converges on four plain questions. What does the system do: not the design intent, the operational reality, including the uses it has quietly acquired since launch. On what evidence does it act: what data it consumes, of what quality, from where. Who decided: who approved this system for this purpose, who accepted its known limits, and, for consequential outputs, who made the final call. What happened: the outcomes, the incidents, the overrides, and what the organisation did when the system was wrong.

Note that these are questions about an organisation, not about mathematics. An auditor is rarely trying to falsify the model. They are testing whether the humans around it exercised judgement, and whether that judgement left a trace.

Why systems without decision records cannot answer

Most organisations can produce artefacts: the model file, the pipeline, the dashboard, the policy document. What they cannot produce is the reasoning. The decision to deploy lives in a slide deck built for a steering committee. The acceptance of a known limitation lives in a meeting that was never minuted. The overrides live in the memories of the people who made them, some of whom have left. When the audit arrives, the organisation does not retrieve its answers; it reconstructs them, and weeks of archaeology produce a narrative that is part memory and part hope. Auditors can usually tell the difference. An AI audit is not a test you sit on the day. It is a record you either kept or did not, and by the time the auditor arrives it is too late to start keeping it.

How provenance answers the four questions

Decision provenance is the preserved chain from evidence to conclusion to decision to outcome, and it answers the audit’s questions in order. What does the system do: the decision record shows what was actually recommended and actually decided, case by case, not what the design document intended. On what evidence: each fact in the chain carries a quality state (measured, modelled, inferred, stated or unmeasured), so “what evidence” arrives together with “how good”. Who decided: a human made each consequential call, and where they overrode the recommendation, the record holds who, when, and against what evidence. What happened: outcomes are scored against both what was recommended and what was decided, so the audit sees not only the calls but how they turned out.

Provenance has one further property auditors care about: it cannot be quietly rewritten. In ONX, facts are versioned, and scenario runs and candidate options are immutable once created, so the record of what was known and considered at the time is the record, not a later reconstruction of it. A decision audit trail assembled this way is not produced for the audit. It is the ordinary exhaust of deciding.

One concrete example

Clearly illustrative, with no customer implied. Two firms each use a screening system in hiring, and each receives the same questionnaire from a client’s procurement team: describe the system, its data, its oversight, and its track record. The first firm convenes a working group. Documents are assembled from old decks, the original project lead is emailed at her new employer, and the answers, when they finally ship, read like what they are: a history written from memory. The second firm opens the record: this is what the system does, with the recommendation-by-recommendation history; this is the evidence it acted on, each fact labelled by quality; these are the humans who decided, with the overrides and their grounds; and this is what happened, scored. The two firms’ models might be equally good. Their audits are not remotely alike, and the difference was settled years earlier, in what each chose to record.

Audit readiness and decision intelligence

Audit regimes differ by industry and jurisdiction, and nothing here is legal advice. But the direction is consistent: from the EU AI Act’s expectations around high-risk systems to client procurement teams asking sharper questions, the demand is for evidence of governed operation, not assurances of it. That is why audit readiness is less a compliance project than a property of decision intelligence done properly. ONX treats candidate matching in Hiring as high-risk under the strict reading of the EU AI Act, with human review gates and a human in the loop for the calls that matter, and keeps its AI explainable and challengeable. More fundamentally, it keeps the chain an auditor asks for as a side effect of normal operation: evidence with quality states, humans deciding, overrides recorded, outcomes scored. The audit stops being an event. It becomes a query.

Common questions

What is an AI audit?

An AI audit is an examination of an AI system and, more importantly, of the organisation around it. It converges on four questions: what does the system actually do in operation, on what evidence does it act, who decided that it should run and who made the consequential calls, and what happened as a result. Auditors are rarely trying to falsify the mathematics. They are testing whether the humans around the system exercised judgement, and whether that judgement left a trace.

What questions does an AI auditor ask?

Four, in essence. What does the system do: the operational reality, not the design intent. On what evidence: what data it consumes, of what quality, from where. Who decided: who approved the system for this use, who accepted its limits, and who made the final call on consequential outputs. What happened: outcomes, incidents, overrides, and what the organisation did when the system was wrong. Every framework elaborates these; none escapes them.

Why do organisations struggle with AI audits?

Because they can produce artefacts but not reasoning. The model file, the pipeline and the policy document all exist, but the decision to deploy lives in a slide deck, the acceptance of a known limitation was never minuted, and the overrides live in the memories of people who may have left. When the audit arrives, the organisation does not retrieve its answers; it reconstructs them, and a reconstruction is slower, costlier and less convincing than a record.

How does decision provenance help in an AI audit?

Decision provenance is the preserved chain from evidence to conclusion to decision to outcome. It answers the audit’s four questions directly: the record shows what was recommended and decided case by case, each fact carries a quality state, every consequential call names the human who made it along with any override and its grounds, and outcomes are scored against both the recommendation and the decision. Kept as a side effect of normal operation, provenance turns the audit from an archaeology project into a query.

Part of the pillarEnterprise Decision Intelligence, the complete philosophy in one essay

Related reading

See a decision run live

Watch evidence land, options reorder against the binding constraint, and the outcome get scored.