All insights

Quality & Compliance · 6 min read

Stop scoring calls.  Start scoring outcomes.

Christian Henson/Chief Technology Officer, Better Outcomz·July 2026
Stop scoring calls.  Start scoring outcomes.

One call does not a customer journey make. If quality assurance is based on one interaction, it tells you whether a process was followed - not whether the customer got the right outcome.

Throughout my time working in regulated financial services, I have watched quality assurance operate primarily as an inspection exercise: a small number of calls selected, a scorecard applied, feedback given to the adviser. That approach has value but it was built around the practical limits of manual review, not a customer's overall experience.

The problem isn't commitment from quality assurance teams; it's limited visibility. No reviewer can spot a pattern across a customer journey if they're only looking at one interaction. A single call may look compliant while the bigger picture tells a different story.

The gap between a compliant call and a good outcome

An adviser saying all the right things on a call does not guarantee a good outcome for the customer. Consider these scenarios:

  • A promised action goes uncompleted.
  • The CRM record conflicts with what was discussed.
  • A vulnerability disclosed early in the journey isn't picked up by the next person.
  • Advice is technically correct but never properly understood.

The Financial Conduct Authority's (FCA's) March 2026 Consumer Finance Regulatory Priorities report draws on its own Financial Lives Survey: 23% of the 1.3 million credit holders who received support from lenders in the two years to May 2024 said they hadn't been able to select an option that suited their needs. Although that figure doesn't, by itself, identify where the journey failed, it strongly suggests interaction-level compliance can't be treated as evidence of a suitable customer outcome.

A good outcome, in a regulated environment, doesn't mean giving the customer the answer they wanted. It means treating them fairly, considering their circumstances, explaining options clearly, completing agreed actions and being able to evidence why the eventual outcome was appropriate. That requires quality management to look beyond the transcript at what happens before, during and after the call.

Why sampling is no longer enough

Sampling was a rational response to the cost of review - manually assessing every call was never a practical option. The result: organisations learned a great deal about a small number of interactions but very little about a customer's overall experiences.

The Mills Review (July 2026) looked at the potential for AI to reshape financial services by 2030, with recommendations for the FCA on how to respond. It suggested sampling may not provide sufficient quality assurance for many AI models. Instead it found that firms will increasingly need more dynamic model governance, monitoring and assurance, supported by end-to-end controls across the AI lifecycle.

Although the review addresses AI rather than call sampling specifically, the same conclusions apply. And AI makes it possible to implement the recommendations where previously call sampling was being used. AI is capable of evaluating more calls, live chats, messages and AI-generated responses against a consistent framework, highlighting interactions that need attention. That doesn't make human reviewers less important, it gives them better coverage and allows them to spend their time on judgement, calibration and improvement rather than initial discovery.

This is why we built Better Outcomes around customer outcome assurance rather than conventional call sampling. The objective isn't more scores. It's a clearer view of what customers experienced and where intervention would make a meaningful difference.

A journey, not a transcript

Better Outcomes brings together multiple interactions including calls, live chat and messaging with CRM-linked evidence, so quality assurance follows the customer's full journey rather than treating each contact as an unrelated ticket. This matters most for vulnerability and conduct risk. Vulnerability is rarely disclosed in one clear phrase and it isn't static; it shows up through changes in language, repeated contact and incomplete actions. Comprehensive assessment of the full customer journey gives quality teams the context to spot that progression and judge whether the organisation responded appropriately.

My test for meaningful quality assurance: Can we explain what the customer experienced, what outcome they received, what evidence informed our assessment and what we would do differently next time?

The same standard has to apply to people and AI

As organisations bring chatbots, copilots and more capable AI agents into customer journeys, quality assurance can't stop at human advisers. An automated response creates the same types of conduct, vulnerability and customer-outcome risks as a human conversation - sometimes at far greater scale. Human and AI-led interactions should be evaluated against the same underlying outcome standards. The individual measures may differ, but the customer should receive consistent treatment regardless of whether a person, an automated service or a combination of both handled the interaction.

Accountability doesn't move either way: the Mills Review confirmed that the Senior Managers Regime continues to apply as AI takes on more of the work, meaning senior managers are accountable for taking reasonable steps to control AI-driven outcomes, even where model behaviour sits partly outside the firm's direct control. That's a technology-strategy problem as much as a compliance one.

Evidence matters more than the QA score

A percentage on a dashboard isn't quality assurance. A useful score has to include its working: the relevant part of the conversation, the connected case information, the rule or outcome expectation being applied and the reason for the contact. Reviewers need to be able to apply their own professional judgement to quality-assurance findings. AI can provide consistency and scale but it should be open to challenge. That's why governance can't be bolted on at the end. Scorecards, prompts, evaluation criteria and model changes need clear ownership. Accuracy has to be measured against human review, and firms need to monitor for false positives, missed issues and any uneven performance across different customer groups.

QA should lead to action

Quality assurance has limited value if it ends at a dashboard. The evidence should lead to action: targeted adviser coaching, changes to customer communications, better workflow controls, clearer policies or a redesign of the journey itself. Repeated findings often point to a process problem, not an individual performance problem.

Better Outcomes isn't about using AI to mark people down; it's about giving quality assurance, operations and product teams a shared view of factors affecting customer outcomes and helping them fix the system advisers and AI agents operate in.

From call scoring to customer outcome assurance

The future of quality assurance isn't autonomous control, with machines making unchecked decisions about whether a customer's journey was good or bad. The stronger model is augmented quality assurance: AI provides breadth, consistency and evidence; experienced people provide context, challenge and accountability.

We should no longer be asking 'Did the adviser follow the process?' But rather 'What did the customer experience? What outcome did they receive? Can we evidence why?' That's the shift Better Outcomes is built to support and, increasingly, the direction in which regulatory expectations are moving.

See it in practice

Better Outcomes evaluates every eligible interaction against your outcome framework.

The article's argument, running in production. Book a 30-minute session and we'll show Better Outcomes against a representative workflow.