Call quality
Manual reviews cover 1–2% of calls. What are you missing in the other 98%?
In most call centers, quality control looks the same: a coach or quality manager listens to a few random recordings per agent per month, fills in a scorecard and discusses the results in a feedback session. The problem isn’t the quality team’s competence — it’s arithmetic. Below we do that arithmetic in full: coverage, the probability of catching a critical error, and cost in euros — then put manual review and AI analysis of 100% of calls side by side, criterion by criterion.
Where does the 1–2% come from?
Let’s count with a simple example we’ll reuse throughout this article. A team of 30 agents makes about 60 calls a day each — that’s 1,800 calls daily and, over 20 working days, roughly 36,000 a month. An experienced quality coach can reliably review and score a dozen or so calls a day: the call itself takes several minutes, then comes the scorecard and notes.
Even doing nothing else, they’ll evaluate 300–400 calls a month. That’s the famous 1–2% — assuming no training sessions, calibrations, meetings or holidays. Note that this isn’t a policy anyone chose; it’s simply the ceiling of what one person can physically listen to.
What a random sample can’t show
A 1–2% sample might suffice if errors were spread evenly. But the most dangerous events — skipped identity verification, misleading a customer about contract terms, a missing mandatory clause — are rare. Run the numbers: if a critical error occurs in 0.5% of calls, that’s about 180 affected calls out of our 36,000. A reviewer sampling 350 calls a month will, on average, stumble on fewer than two of them — and any single bad call has a roughly 99% chance of never being heard by anyone.
In practice this means the organization learns about serious errors not from monitoring, but from a customer complaint or a regulator’s letter — weeks after the fact.
Sample-based monitoring doesn’t answer the question “how is my team performing”. It answers “how did a few randomly picked calls go”.
Randomness hurts agents too
From the agent’s perspective, sample-based evaluation can be simply unfair. One weaker call drawn for review can weigh down a monthly score, while dozens of good ones go unnoticed. It’s hard to build a feedback culture on that — the result depends on the draw, not the work. Agents know it, and it undermines trust in the whole quality program.
What manual review actually costs — a worked example in euros
Coverage is one side of the equation; cost is the other. Three explicit assumptions — swap in your own numbers, the arithmetic stays the same:
- Volume: 30 agents × 60 calls a day × 20 working days = 36,000 calls a month.
- Throughput: one quality specialist reliably scores about 350 calls a month (the 300–400 range from above).
- Cost: a fully loaded quality specialist — salary, employer contributions, tooling — costs €3,000 a month. In Western Europe this figure is often higher, in Central Europe lower; adjust to your market.
Now the math. One scored call costs €3,000 ÷ 350 ≈ €8.60. One full-time reviewer covers 350 ÷ 36,000 ≈ 1% of traffic. Want the “industry standard” 2%? That’s two specialists — €72,000 a year to hear 700 of 36,000 monthly calls. And full manual coverage of 100%? You’d need 36,000 ÷ 350 ≈ 103 reviewers, about €309,000 a month — at the same €3,000 fully loaded cost, more than three times the payroll of the 30-agent team they’d be checking. Nobody does that, of course. Which is exactly why 98% of calls go unheard.
Manual review vs AI analysis of 100% of calls
Here’s the same comparison laid out criterion by criterion, using the example above:
| Criterion | Manual review (random sample) | AI analysis of 100% of calls |
|---|---|---|
| Coverage | 1–2% of calls; ~350 of 36,000 a month per reviewer | Every call — 36,000 of 36,000 |
| Time from call to score | Days to weeks — whenever the next review cycle reaches it | Minutes after the call ends |
| Critical errors | Caught only if the call happens to be drawn — ~99% chance a given bad call is never heard | Flagged the same day, with a transcript quote |
| Scoring consistency | Varies by reviewer, day and fatigue | Same written criteria applied to every call; a human can verify and correct any score |
| Basis for agent evaluation | A handful of sampled calls per agent per month | A trend across all of the agent’s calls |
| Cost at 36,000 calls/month | ~€8.60 per scored call; 100% coverage ≈ 103 reviewers ≈ €309,000/month | Software subscription — no extra reviewers needed as volume grows; pricing discussed at a demo |
| Quality team’s role | Listening and filling in scorecards | Calibrating criteria and coaching where the data points |
One honest caveat on the right-hand column: CallSea doesn’t publish a price list — pricing is discussed at a demo. But the structural difference holds regardless of the number: software doesn’t need one additional reviewer for every additional 350 calls.
What analyzing 100% of calls changes
Automatic AI call analysis flips the sampling logic. Every call is transcribed and scored against the same criteria — checklists, point scales, critical-error rules, plus an overall 0–100 AI score with a written rationale — within minutes of hanging up.
- Critical errors surface the same day, with the specific call and a transcript quote you can bring to the feedback session — see how critical-error detection works.
- Agent evaluation rests on complete data — per-agent trend dashboards across hundreds of calls instead of an impression from three sampled recordings.
- The quality team stops listening and starts managing quality: calibrating criteria, working with results, coaching where the data points.
The human role doesn’t end — it changes character. AI sifts through 100% of calls and shows where to look; the coach makes decisions, runs coaching and calibrates the scoring criteria. Every AI score can be verified and corrected by a human. And one boundary worth knowing: CallSea’s models read only the transcript text — they do not analyze agents’ tone of voice or emotions.
Where to start?
The best first step costs nothing: write down your current scorecard and the rules you evaluate calls by today — they become the criteria for AI scoring. The second step is technical and smaller than it sounds: getting recordings into the system, for example via automated SFTP ingestion. How to design criteria so automatic scoring is unambiguous is a topic for a separate guide, and if you read Polish, our post on designing scoring criteria covers it in detail.
Frequently asked questions
What percentage of calls should a call center monitor for quality?
There is no regulatory minimum. Most teams sample 1–2% simply because that is what a human team can physically score — one specialist reliably covers about 350 calls a month. The more useful question is which calls need human attention: with automated scoring of 100% of calls, coaches review the flagged conversations instead of a random draw.
Is AI call scoring as accurate as a human reviewer?
It is more consistent: the same written criteria are applied to every call, with no fatigue and no drift between reviewers. Accuracy depends on how unambiguous your criteria are, which is why calibration on historical recordings matters. In CallSea a human can verify and correct any AI score, so human oversight stays in the process — and the models read only the transcript text, they do not analyze agents’ tone of voice or emotions.
Do we still need quality coaches if AI analyzes every call?
Yes — the role changes rather than disappears. AI does the sifting: it transcribes, scores and flags calls. Coaches decide what the results mean, run coaching sessions and calibrate the scoring criteria. The shift is from listening to a 1–2% sample to working with complete data on where the team actually loses quality.