Knowledge Centre
Quality Assurance ScoreOperations & Workforce··4 min read

What is a Quality Assurance Score?

A team can hold an excellent quality assurance score for a year while its customers quietly conclude that the service is getting worse. Both things can be true at once, because a QA score does not measure what customers experience. It measures how closely sampled work matches the organisation’s own definition of good.

What a quality assurance score measures

A quality assurance score is a graded evaluation of a sample of work against an internal scorecard. In operations the sampled work is usually recorded calls, chats, emails or completed cases. An assessor scores each item against weighted criteria: were the required statements made, was the process followed, were the notes complete, was the customer handled well. Some criteria are pass or fail; serious breaches often fail the whole item automatically. The item scores aggregate into a percentage for a person, a team and a programme, tracked week by week.

Three phrases in that description carry the weight, and each limits what the number can claim. It is the organisation’s own standard: the scorecard encodes what the firm believes good looks like, which is a hypothesis, not a fact. It is a sample: a handful of interactions per person per period, from which a general judgement is extrapolated. And it passes through a human assessor: two trained people can watch the same call and score it differently. None of this makes the score useless. It makes the score evidence, of a particular kind and quality, rather than truth.

Three ways the number bends

  • Sampling error dressed as precision. A few evaluations per person is a small sample, yet the resulting percentage is reported to a decimal place and compared month to month as if it were exact. Much of that movement is the luck of the draw. The smaller the sample, the more the score describes which interactions happened to be picked.
  • Calibration drift between assessors. Standards diverge person by person unless assessors regularly score the same items and reconcile their differences. Once they drift, a team’s improvement can be an assessor change, a site’s advantage can be a lenient grader, and cross-team comparison quietly loses its meaning.
  • A rubric that is not the customer. Scorecards reward what is easy to grade: the greeting, the disclosure, the tidy notes. The customer experiences something the scorecard rarely asks about: whether the issue was actually resolved, how much effort it took, and whether it stayed fixed. This is the gap between QA scores and customer experience, and it can widen for years behind a stable number.

The pattern behind all three is the same. The score is treated as a measurement of quality when it is a measurement of agreement with a rubric, taken on a sample, by people who drift. Firms that remember this read the score alongside customer signals. Firms that forget it end up defending the rubric against the customer.

Coaching input or policing tool

The same instrument does opposite work depending on what is wired to it. As a coaching input, QA scoring is at its best: an assessor and an agent sit with a specific interaction, the criteria give the conversation structure, and the small sample matters less because the purpose is teaching, not judging. As a policing tool, with pay or discipline hanging from the number, everything degrades at once. Agents contest scores instead of learning from them. Assessors soften to avoid conflict. Scores inflate year on year while nothing improves. When a measure becomes a target it stops being a good measure, and a QA score is an unusually soft target because the organisation grades its own homework.

There is a simple tell. Ask what changed in the last quarter because of QA findings. If the answer is a list of individual conversations, the programme is policing people. If the answer includes process changes, scorecard revisions and training built from recurring findings, the programme is doing its real job: turning sampled evidence into organisational learning.

One concrete example

Clearly illustrative, with no customer implied. An outsourced support operation of a few hundred agents holds a strong QA score all year. At renewal, the client arrives with its own survey, which says customers are frustrated, chiefly about contacting support repeatedly for the same issue. The scorecard had graded every sampled contact in isolation: greeting, empathy, notes, closure. Every contact in a frustrating sequence of five could pass individually. Nobody was scoring whether the issue stayed resolved, because the rubric never asked. The score was true, and beside the point. The useful response is not harsher grading; it is changing the question: adding repeat contact as an outcome check, calibrating assessors against resolved versus reopened, and treating the scorecard as a hypothesis to revise rather than a standard to defend.

The decision a QA score should trigger

Read properly, a QA score is not a verdict on people; it is a prompt for decisions. A falling score should trigger a diagnosis: is the work worse, the sample unlucky or the assessor stricter? A stable score alongside worsening customer signals should trigger a rubric review, because the instrument has stopped tracking the thing it exists for. Recurring findings should trigger process change, coaching time and training, upstream of any individual. A quality assurance score is your own standard, sampled and humanly applied: weigh it as evidence, do not enforce it as a verdict.

In a decision-intelligence view, that is exactly how it enters the picture: as evidence with a stated quality, placed honestly in the evidence hierarchy next to stronger and weaker signals, feeding judgements such as delivery confidence rather than standing alone as a scoreboard. The decisions it triggers, revising a rubric, reallocating coaching, fixing the process that keeps generating the same finding, deserve to be recorded with their reasoning and checked against what actually happened, including what happened to the cost of serving the customers involved. A QA programme that closes that loop gets better every quarter. One that only publishes the percentage gets better at publishing the percentage.

Common questions

What is a quality assurance score?

A quality assurance score is a graded evaluation of a sample of work, most often recorded customer interactions or completed cases, against an internal scorecard. Assessors score each sampled item on weighted criteria such as process adherence, accuracy and communication, and the results aggregate into a percentage for a person, team or programme. It measures adherence to the organisation’s own standard, on a sample, through human judgement, and each of those three qualifiers limits what the number can claim.

Why do QA scores and customer satisfaction disagree?

Because they grade different things. A QA scorecard measures compliance with an internal standard, and it tends to reward what is easy to grade: the greeting used, the notes completed, the script followed. Customers judge whether the issue was resolved, how much effort it took and whether it stayed resolved. An interaction can pass every scorecard item while failing the customer, so a high QA score alongside falling satisfaction usually means the rubric is asking the wrong questions, not that the assessors are careless.

What is calibration drift in quality assurance?

Calibration drift is the gradual divergence of assessors from a shared standard, so the same interaction earns different scores depending on who happens to grade it. Regular calibration sessions, where assessors score the same items and reconcile their differences, keep the standard shared. Without them, an apparent improvement in a team’s quality can be nothing more than a change of assessor, and comparisons across teams or months quietly stop meaning anything.

Should QA scores be linked to discipline or pay?

With great caution. The same instrument does opposite jobs depending on what hangs from it. Used as a coaching input, a QA score opens a specific, low-stakes conversation about the work itself. Wired to bonuses or discipline, it invites gaming: disputed scores, pressured assessors and gradual inflation, until the number stops carrying information. When a measure becomes a target it stops being a good measure, and a QA score is an unusually soft target because the organisation is grading its own homework.

Part of the pillarEnterprise Decision Intelligence, the complete philosophy in one essay

Related reading