An AI governance audit tests whether the organization can show that its AI rules work in practice. It should trace selected systems and use cases from policy to ownership, approval, controls, monitoring, incidents, and retained evidence. The result is not a generic maturity score. It is a defensible record of what works, what does not, and who will fix it.
An internal governance audit can prepare leaders for regulatory, customer, board, or assurance questions, but it is not automatically a legal opinion or independent certification. Define the purpose and audience before work begins.
Set the scope around decisions
Choose which business units, AI systems, vendors, agents, models, data classes, and workflows are in scope. Record the governing policies, risk tolerance, evaluation criteria, period under review, and evidence deadline.
Sample use cases by consequence, not convenience. Include at least one high-impact workflow, one third-party service, one employee-led use case, one system with sensitive data, and one system that can take or recommend an action.
Build an evidence map
For every requirement, identify the control, owner, frequency, evidence source, and expected result. Typical evidence includes:
- AI inventory and approved-use register
- business and control ownership records
- risk classifications and approval decisions
- vendor assessments and contract terms
- data-flow, access, retention, and permission records
- model or workflow evaluation results
- human-review procedures and sampled approvals
- employee guidance and training completion
- monitoring, exception, incident, and change logs
- decommissioning or exit plans
A policy statement is evidence that a rule exists. It is not evidence that the rule was followed.
Test seven control areas
Governance and accountability
Confirm that owners know their responsibilities and can explain current decisions. Check whether exceptions and accepted risks have named approvers and expiration dates.
Inventory and classification
Test whether the inventory includes actual tools and workflows, not only centrally purchased software. Compare it with procurement, identity, browser, API, and team-level evidence where appropriate.
Data and access
Sample permissions, inputs, retention settings, connectors, logs, and deletion behavior. Verify that stated data restrictions match technical configuration and user practice.
Evaluation and human oversight
Review whether test cases represent real tasks and important failure modes. Sample human approvals to see whether reviewers received enough context and could meaningfully reject or modify the action.
Vendor and supply-chain controls
Confirm that product tiers, subprocessors, model providers, contract commitments, and change-notification processes are known. Check that vendor approval is tied to a defined use rather than blanket use of a brand.
Monitoring and incidents
Verify who watches for failures, which thresholds trigger review, how users report issues, and whether incident exercises have tested containment, rollback, communication, and recovery.
Workforce capability
Check whether employees and managers can apply the rules to realistic scenarios. Training attendance alone does not show that people can classify data, choose an approved tool, review output, or escalate uncertainty.
Grade findings by exposure and evidence
Use a small number of clear finding types: effective, partially effective, missing, or not applicable. Add severity based on potential impact, likelihood, breadth, and detectability. Record the exact evidence reviewed and the reason for the conclusion.
Avoid inventing certainty. If evidence is missing, say that the control could not be verified. If a control exists but is not consistently followed, distinguish design weakness from operating failure.
Turn findings into a closure plan
Each open finding needs an accountable owner, target date, interim containment, remediation action, dependency, and closure evidence. Leaders should explicitly accept, mitigate, transfer, or avoid the risk.
Re-test closure. A revised policy, new dashboard, or completed training session may be part of remediation, but it does not prove that the underlying workflow now operates as intended.