Large language models introduce familiar model, technology, data, third-party, and operational risks in a new combination. A defensible program connects each use case to its business owner, risk tier, controls, evidence, and decision rights.
Use this checklist during intake, architecture review, vendor due diligence, validation, release approval, and periodic monitoring. Each section can remain open while another is reviewed so teams can compare requirements across the lifecycle.
How to use this checklist
- Document the business purpose, users, affected customers, decisions, and prohibited uses.
- Assign an executive owner, product owner, control owners, reviewers, and incident decision-maker.
- Classify inherent risk based on customer impact, autonomy, data sensitivity, external use, and decision materiality.
- Define when human approval is mandatory and who may override or stop the system.
- Inventory prompts, grounding sources, training or tuning data, outputs, logs, and retained records.
- Approve data classifications, customer-consent requirements, retention periods, and deletion procedures.
- Restrict model, tool, retrieval, and administrative access using least privilege and separation of duties.
- Test for sensitive-data leakage, prompt injection, unauthorized retrieval, and cross-user exposure.
- Record the model, version, provider, hosting arrangement, configuration, tools, connectors, and dependencies.
- Complete vendor diligence covering security, privacy, resilience, subcontractors, data use, model changes, and exit rights.
- Define retrieval boundaries, approved tools, output filters, fallback behavior, and failure containment.
- Document how model or provider changes trigger reassessment, testing, approval, and communication.
- Establish test cases for accuracy, groundedness, completeness, fairness, safety, privacy, and policy compliance.
- Include adversarial prompts, ambiguous requests, missing context, tool failures, and unsupported claims.
- Set measurable thresholds and document who accepts residual risk when a threshold is not met.
- Validate human-review procedures and ensure reviewers receive enough context to challenge outputs.
- Confirm security, privacy, legal, compliance, model-risk, business, and technology approvals required by risk tier.
- Version prompts, retrieval sources, tools, policies, guardrails, model configuration, and release evidence.
- Use controlled pilots, limited populations, feature flags, rollback plans, and manual alternatives.
- Train users on appropriate reliance, escalation, prohibited data, and reporting obligations.
- Monitor usage, exceptions, harmful or unsupported outputs, overrides, latency, cost, complaints, and control failures.
- Define thresholds for alerting, restricted use, rollback, suspension, escalation, and regulatory assessment.
- Connect incidents to root-cause analysis, remediation, retesting, approvals, and customer response.
- Retain the model inventory, decisions, tests, approvals, logs, incidents, changes, and periodic reviews.
Minimum evidence package
Before release
- Use-case and risk assessment
- Architecture and data-flow record
- Testing results and exceptions
- Approvals and operating procedures
After release
- Monitoring and threshold reports
- Changes and version history
- Incidents and remediation
- Periodic review and risk acceptance
Put the checklist into your review process
Request the working checklist for use during intake, validation, change approval, and ongoing oversight.
Purchase the LLM Model Risk Checklist


