Every regulated advice firm knows the arithmetic of manual QA: a reviewer can only listen to so many calls, so most calls are never heard. AI call monitoring changes that arithmetic — but only if its findings can be trusted, checked and acted on. Here is how it works in practice.

The limits of manual call QA

Reviewing a call properly takes at least as long as the call itself, plus time to take notes and write up findings. A 30-minute fact find is the best part of an hour of a reviewer's time. However diligent the team, that caps coverage at a small sample — and a sample is poor at catching problems that are rare but serious. (We worked through the numbers in how many calls should you review?)

Manual review also varies from person to person and from week to week, and findings tend to arrive long after the call — often at a file review or in a complaint, when it is hardest to put things right.

How AI call monitoring works, step by step

  1. The call arrives. Finished calls are collected automatically from your dialler, or recordings are uploaded from any other system.
  2. It is transcribed. Speech becomes a speaker-labelled, timestamped transcript, usually within minutes.
  3. It is classified and scoped. The call type — initial contact, fact find, follow-up — decides which rules apply. A call-selection policy decides whether it is assessed at all, and excluded calls are logged with the reason.
  4. It is assessed against your rules. Each rule checks one thing: exact wording that must be said, a topic that must be covered, signs of vulnerability, how objections were handled, or whether what was said matches the documents on file.
  5. Every verdict is evidenced. Passed, needs review or failed — with the quoted words, a written explanation and a confidence level.
  6. People review what's flagged. Reviewers confirm or override results, and the original verdict is kept for the audit trail.
  7. Results flow on. Alerts for failed rules, MI by call type and adviser, and reports or client files exported when they're needed.

After the call, or during it?

Some tools prompt advisers in real time. For compliance QA, assessing the finished call has real advantages: the whole conversation is available, so a risk mentioned early can be judged against what the adviser did later; nothing competes for the adviser's attention while they're with a client; and there's no trade-off between speed and accuracy. COSA works this way — it picks up each call once it ends and can't make, join or interrupt one.

Can you trust the verdicts?

Treat any accuracy figure quoted in isolation with caution. Performance depends on audio quality, accents and crosstalk, how precisely rules are written, and the mix of calls a firm handles. What matters is how the system performs on your calls, against your rules — and whether you can check each result.

  • Calibrate before you rely on it. Run a sample of past calls through the system, have experienced reviewers assess the same calls, and compare. Where they disagree, look at the evidence: often the rule needs sharper wording.
  • Demand evidence per verdict. A result you can't trace to the words on the call is a liability in a regulated firm.
  • Watch the overrides. The rate and direction of reviewer overrides tell you whether the system is too strict or too lenient on particular rules.
  • Keep a human sample of passes. Periodically reviewing a few calls that passed checks that nothing is slipping through.

Designing rules that work

The quality of automated QA depends on the rules you give it. Good rules describe observable behaviour, check one thing each and apply only to the call types where they make sense.

Rule typeUse it forExample
VerbatimWording that must be said exactly"This call is recorded for training and compliance purposes."
Non-verbatimTopics that must be covered, however phrasedInterest roll-up and its effect on the estate explained
VulnerabilitySigns of vulnerability and the adviser's responseBereavement disclosed; pace and support offered
Objection handlingWhether concerns were dealt with in substanceConcern about inheritance acknowledged and answered
Document checksWhat was said versus the documents on fileFigures discussed match the illustration

For a deeper guide, see writing compliance rules an AI can check.

A rollout plan

  1. Week 1 — connect and scope. Connect your call source, agree which call types and lengths are in scope, and confirm who needs access and in which role.
  2. Weeks 1–2 — write rules. Start with the rules that carry the most risk — disclosures, vulnerability, key product risks — rather than trying to encode your whole manual on day one.
  3. Weeks 2–3 — calibrate. Assess a sample of past calls, compare with your reviewers, refine wording, and reassess.
  4. Weeks 3–4 — go live. Review flags daily, share MI weekly, and agree how findings turn into coaching and remediation.
  5. Ongoing — improve. Add rules as confidence grows, and reassess calls when your standards change.

What to measure

  • Coverage: the share of in-scope calls assessed, and excluded calls by reason.
  • Compliance rate: overall, by call type and by adviser.
  • Needs-review volume and time to review: how quickly flagged calls reach a decision.
  • Most-missed rules: where training or process changes will have the biggest effect.
  • Overrides: how often, on which rules, and in which direction.
  • Vulnerability: signals found, how consistently they were handled, and what followed.

Data protection and security

Call recordings are some of the most sensitive data a firm holds, and they often contain special category data such as health information. Before you start, agree the lawful basis, complete a data protection impact assessment, and check the supplier's access controls, data isolation, audit logging and retention. Our security page sets out how COSA handles each of these.

This article is general information, not legal or regulatory advice.