Most compliance teams sample calls because they have to: there aren't enough hours to listen to them all. The question is what a sample can and can't tell you. The answer comes from some simple arithmetic — and it is less reassuring than most firms assume.

The question a sample has to answer

There are two different reasons to sample calls, and they need very different sample sizes:

  1. Detection — "if something is going wrong, will we find at least one example of it?"
  2. Estimation — "what proportion of calls meet our standards, and is it getting better?"

Both matter under the Consumer Duty. Detection is about stopping harm; estimation is the MI your board sees.

Detection: will a sample catch a rare failure?

If a fraction p of calls contain a particular failure and you review n calls chosen at random, the chance of finding at least one is 1 − (1 − p)n. Here is what that looks like:

Share of calls with the failure10 calls25 calls50 calls100 calls200 calls
0.5% (1 in 200)5%12%22%39%63%
1% (1 in 100)10%22%39%63%87%
2% (1 in 50)18%40%64%87%98%
5% (1 in 20)40%72%92%99%>99%

Read across the "1 in 50" row. A failure affecting one call in fifty — an unexplained early repayment charge, a missed recording notice, a bereavement nobody responded to — has a 60% chance of not appearing at all in a sample of 25 calls. Even a sample of 100 misses it entirely more than one time in eight.

Turned around, this is the sample you need for a 95% chance of catching at least one example:

Share of calls with the failureCalls to review for a 95% chanceFor a 99% chance
5% (1 in 20)5990
2% (1 in 50)149228
1% (1 in 100)299459
0.5% (1 in 200)598919

And that is for one kind of failure, measured across the whole firm. Every separate issue you want to detect — each disclosure, each vulnerability driver — faces the same odds.

Estimation: how precise is your compliance rate?

If your sample says 10% of calls fall short on a rule, how sure can you be? The 95% margin of error for a proportion is roughly ±1.96 × √(p(1 − p)/n):

Calls reviewed25501002004001,000
Margin of error on a 10% rate±11.8 pts±8.3 pts±5.9 pts±4.2 pts±2.9 pts±1.9 pts

With 50 calls, a reported 10% failure rate could really be anywhere from about 2% to 18%. A quarter-on-quarter "improvement" from 12% to 9% on samples that size is indistinguishable from noise.

The adviser problem

Boards and team leads usually want to know which advisers need support. Sampling makes that very hard. Spread 50 calls across ten advisers and each is judged on five calls. An adviser who gets one call in ten wrong will show no problems at all in their five calls nearly 60% of the time — and a single bad call can make a good adviser look like an outlier.

What sampling costs

Reviewing a call properly takes at least its own length, plus time to note and write up findings. As an illustration, at around 45 minutes per 30-minute fact find, 50 reviews a week is more than 37 hours of reviewer time — roughly one full-time person — and still only reaches the detection rates in the table above. Try your own numbers in the ROI calculator.

Is a sample still defensible?

The FCA doesn't prescribe a coverage figure. What the Consumer Duty does require is that firms monitor the outcomes customers receive, identify where things go wrong — including for particular groups, such as customers with characteristics of vulnerability — and evidence it to the board. A firm relying on a sample should be able to explain why the sample is big enough to do that.

That explanation was easier to give when reviewing every call wasn't possible. Now that it is, "we couldn't have listened to them all" is a weaker answer than it used to be.

A better model: review every call first, then sample

  1. Automated first pass on every in-scope call, against your own rules, with the quote and reasoning for every verdict.
  2. People review what's flagged — the calls that need judgement — within a day.
  3. A small random sample of passes gets a human review too, to check nothing is slipping through and to keep reviewers calibrated.
  4. Every excluded call is logged with its reason, so coverage is explicit.

In this model the sample stops being your evidence and becomes your quality check on the evidence. It is how Compare Retirement went from sampling around a quarter of calls to overseeing all of them, with QA on flagged calls taking minutes rather than hours.

This article is general information, not legal or regulatory advice. Figures in the tables assume calls are chosen at random from all calls. Real samples often aren't — particular teams, call lengths or weeks get reviewed more — which makes them less representative than the tables suggest.