An AI risk assessment should help a decision-maker decide whether a use case can proceed, what controls it needs, and who accepts the remaining risk. It should not become a generic questionnaire completed once and forgotten.
The best assessments begin with the workflow, not the model.
Step 1: Define the intended use
Write a precise statement of purpose. Identify users, affected people, input data, output, system connections, frequency, geographic scope, and the decision or action the output influences. Record foreseeable misuse and explicitly excluded uses.
Step 2: Establish impact
Ask what happens when the system is wrong, delayed, manipulated, unavailable, or used outside its intended context. Consider financial, legal, safety, privacy, employment, customer, reputational, and operational consequences. Include the scale of exposure and whether harm can be reversed.
Step 3: Examine data and access
Record data sources, classification, provenance, permissions, retention, transfers, and downstream reuse. For agents, list every tool and system the agent can access. Broad access can turn a modest model error into a significant business event.
Step 4: Map failure modes
Test likely and severe failures, including inaccurate output, bias, prompt injection, data leakage, inappropriate automation, stale knowledge, brittle integration, user overreliance, and silent degradation. Add workflow-specific failures rather than relying only on a standard list.
Step 5: Design controls
Controls may include input restrictions, identity and permission boundaries, grounding, output validation, human review, approval thresholds, rate limits, logging, monitoring, fallback procedures, and user training. Each control needs an owner and testable evidence.
Step 6: Evaluate residual risk
After controls, decide whether the remaining risk is acceptable for the intended use. Record the approver, conditions, launch scope, monitoring thresholds, and events that require reassessment.
Use proportionate depth
A low-impact drafting assistant should not face the same process as an autonomous agent that changes customer records. Use risk tiers to scale the evidence required. Higher autonomy, sensitive data, consequential impact, broad reach, and weak reversibility should trigger deeper review.
Keep the assessment alive
Reassess when the model, prompt, data source, connected tool, user population, output purpose, or regulatory context changes. Production monitoring should feed back into the risk record. Incidents and near misses are evidence, not separate paperwork.
A strong decision record
At minimum, the final record should show the intended use, risk tier, accountable business owner, key failure modes, controls, test results, residual risks, approvers, launch conditions, monitoring plan, and next review date.
The assessment is complete when an accountable person can explain why the use is permitted and point to evidence that the controls work.