AI Governance · 9 min read
AI adoption does not end at go-live
An AI tool can perform well in a pilot and still fail at enterprise level. Successful adoption requires an organisation to be able to evidence why the tool exists, how it was approved, what risks it creates, how people oversee it and what happens when it changes.
Enterprise AI is often presented as a route from idea to pilot and then into production. In a regulated environment, production is not the final milestone. It is the point at which assumptions meet real users and data, operational pressure and customer consequences. That is why AI assurance doesn't end at go-live and is a vital part of the full lifecycle.
Models change. Prompts and reference material are revised. Suppliers update their services. Colleagues find new ways to use a tool. The customer population and the circumstances in which a system operates evolve. A control that was appropriate at launch can become incomplete later.
The Mills Review (July 2026) looked at the potential for AI to reshape financial services by 2030, with recommendations for the Financial Conduct Authority (FCA) on how to respond. It suggested point-in-time validation may be insufficient for adaptive systems. Instead, it found that firms will increasingly need more dynamic model governance, monitoring and assurance, supported by end-to-end controls across the AI lifecycle. It also confirmed that senior managers remain accountable for taking reasonable steps to control AI-driven outcomes under the Senior Managers Regime, even as systems become more autonomous. Assurance that stops at go-live meets neither expectation.
The practical question is not does the tool work but rather can we demonstrate that it is approved, safe, controlled, explainable and being used for its intended purpose. That question moves the conversation away from a polished demonstration and towards evidence.
AI adoption: 6 questions you should be able to answer
Before AI can be classed as an enterprise-grade tool, a team should be able to answer six questions.
- What purpose does the AI serve?
- What could go wrong?
- What controls prevent, detect or correct failure?
- What evidence is there for the effectiveness of those controls?
- What outcomes are customers and colleagues experiencing?
- Who is accountable for decisions, issues and change?
These questions apply across generative AI, predictive machine-learning systems and rule-based automation, but the tests will differ. A generative system needs controls for inaccurate or invented content, output validation, and prompt and data handling. A predictive model raises questions about training data, bias, drift and explainability. A rules engine requires version control, exception testing and escalation. AI is a broad church so assurance isn't a generic checklist.
The adoption test: Purpose · Risk · Controls · Evidence · Outcomes · Accountability. If you're unable to explain and evidence all of these areas, the AI can't yet be classed as an enterprise-grade tool.
Assess the AI lifecycle, not a snapshot
A tool looks different depending on when you check it. Real assurance looks back to how it was approved, sideways to the controls it depends on and forward to how it's monitored and changed. The framework below turns that into six practical stages.
| Lifecycle stage | Assurance focus |
|---|---|
| 1. Inventory | Record the business owner, approved purpose, affected teams and where the use case sits in the organisation's AI inventory. |
| 2. Approval | Evidence the proposal, risk classification, required approvals, dependencies and any conditions attached to use. |
| 3. Build or configure | Control data sources, prompts, rules, access, security, integrations and supplier settings. |
| 4. Test before launch | Test accuracy, fairness, security, resilience, edge cases and customer-outcome scenarios. |
| 5. Operate | Use trained colleagues, meaningful human review, audit trails and clear limits on approved use. |
| 6. Monitor and change | Review management information, incidents, complaints, drift and changes to the model, supplier, data or purpose; retire the use case when appropriate. |
The links between these stages matter. A live control that can't be traced back to an approved requirement or a named risk isn't an admin failure, it's one of governance - the organisation can't say why the control exists or whether it still works.
AI risk classification: it's about the impact, not the technology
Risk classification sets the depth of assurance, the evidence required and the level of approval. It doesn't decide whether basic controls apply - they always apply. A low-risk drafting aid still needs an approved account, a trained user and limits on sensitive data. An operational decision tool needs documented validation and ongoing monitoring. Anything touching customer decisions, financial outcomes, regulatory assessments or sensitive data needs deeper testing, stronger oversight and senior challenge.
Don't confuse simple technology with low risk. A customer-facing or sensitive-data use case stays high risk even if the tool looks basic - the risk lives in the purpose, the data and the potential consequences, not the technology. Review the classification whenever the use case, data, model, supplier or customer impact changes.
Human oversight only counts if a human can change the outcome
'Human-in-the-loop' is often written into an approval as though the phrase itself is a control. It isn't. The Information Commissioner's Office's (ICO's) own evidence shows how often this goes wrong in practice. Its March 2026 review of automated decision-making in recruitment, drawing on more than 30 employers, found that most didn't recognise they were making solely automated decisions at all - many assumed a person nominally sitting in the process was enough. The ICO's position is clear: a nominal reviewer is not meaningful human involvement. A reviewer must have the authority, discretion and relevant information to change the outcome, not just endorse it. Where that standard isn't met, the process is treated as solely automated, regardless of how it looks on paper.
Oversight only means something when the reviewer is trained, can see the source information, understands the acceptance criteria, has the authority to challenge or stop the output, and leaves a record of the decision made. That's a different thing entirely from asking a busy colleague to sign off an output they can't explain, with no record of any corrections or overrides. If the person can't influence the result, the human is present but the oversight isn't.
AI data risk isn't confined to the model
Data risk doesn't sit inside the AI model alone. It can enter through a prompt, an uploaded file, a connected source, a supplier setting, a generated output or a retention rule. Assurance has to follow the data before, during and after processing - not just check the model once.
Input controls should limit use to approved, necessary data, backed by authorised accounts and clear upload rules. Processing controls should cover access, supplier retention and training settings, secure integration and audit records where needed. Output controls should cover verification, the intended audience, storage and deletion, transparency where required, and a record of corrections and overrides.
Since 5 February 2026, the Data (Use and Access) Act 2025 has replaced article 22 of the UK general data protection regulation (GDPR) with new articles 22A to 22D, changing the legal test for automated decision-making itself. But the legal test changing doesn't mean the need for governance, safeguards and accountability has - the ICO puts governance and accountability at the centre of AI data-protection risk and expects an embedded privacy framework running through development, use and oversight. In practice, that means data protection can't be a one-off launch check. It has to stay visible in day-to-day operation and every change made after go-live.
Test what happens to the customer, not just what the technology does
In regulated financial services, an output can be technically correct and still land badly for the customer. Testing needs to look at consistency across similar cases, assumptions that don't hold up, vulnerability indicators, edge cases and whether the communication is clear, fair and suitable for that customer's circumstances. The FCA's Consumer Duty requires firms to avoid foreseeable harm and actively deliver good outcomes - and to monitor the outcomes customers actually receive. An AI assurance framework needs to connect the technical evidence to the customer evidence: complaints, overrides, repeat contact, harm indicators, how the tool performs for customers with characteristics of vulnerability, and what happens when a weakness is found. A good average accuracy score means nothing if the errors cluster in the cases where customers need the most care. Outcome testing exists to answer three questions: who is affected, how are they affected, and can the organisation step in before a recurring weakness becomes a systemic one.
Assurance is shared, accountability isn't
The Institute of Internal Auditors' Three Lines Model provides a useful structure. Business teams operate and monitor the tool day to day. Risk, compliance, QA and governance provide oversight and challenge. Internal audit provides independent assurance over the framework itself and how it operates in practice. Each line has a distinct role, but all three need access to the same evidence. A monitoring plan needs to state what's sampled, how it's tested, which criteria apply, where the thresholds sit and what happens on a breach. 'QA will review' isn't a monitoring plan. Permanently green MI, with no thresholds and no issue history, isn't reassuring - it's a red flag.
Accountability doesn't belong to the AI, and it can't be allowed to fall through the cracks between teams. Someone owns the purpose, the performance, the data, the decisions, the incidents and the changes, for as long as the use case exists. The failure I see most often isn't a broken model but rather a use case that's outgrown its approval - a new data source nobody logged, a new team using it slightly differently, a supplier update nobody tracked. The warning signs show up early: unapproved tools or personal accounts, scope creep, missing prompt or decision records, human review with no real power to change anything, changes released with no regression testing or re-approval. A red flag doesn't mean the tool has to stop. It means the organisation needs more evidence, sharper challenge or escalation. What matters is that the issue gets recorded, owned and followed through - not waved through as the cost of moving fast.
The AI QA mindset
- Trace the use case back to its approval, risk rating and controls.
- Test inputs, outputs, human decisions, exceptions and monitoring.
- Challenge whether the evidence and oversight are meaningful.
- Escalate misuse, evidence gaps, customer-harm indicators and recurring weaknesses.
Enterprise AI should be treated as a controlled operating capability, not a collection of tools. Success isn't the number of licences issued or pilots launched. It's whether the organisation can show real value and good outcomes while understanding and controlling the risk it's taken on.
Go-live isn't where assurance ends. It's where the evidence starts building.
Adoption, evidenced
See how Better Outcomes handles the review + calibration loop.
The reviewer/override cycle described in the essay is a first-class capability. Prefer to talk architecture and controls? Book a structured technical review.
