# The Rubric Your AI Judge Is Running Has Never Been Tested Against a Human. That Is Not a Small Problem.

> When an AI scoring system ships without rubric validation, every number it produces is an opinion with a confidence interval nobody measured.

By Greg McCallum, founder of PeerLab. Published 2026-09-18 in AI Scoring & Assurance. Source: https://www.peerlab.ai/blog/ai-judge-rubric-you-inherited-was-never-tested/

You have an AI scoring your calls. It produces a number. That number influences coaching decisions, headcount decisions, sometimes performance management. The question nobody in the room has answered is: what, precisely, is that number measuring, and how do you know the criteria it is executing were ever right?

This is not a question about model accuracy. It is a prior question. A model can be extremely faithful to a rubric that was written in an afternoon, never challenged, and never checked against what a trained human actually notices in a good conversation. Transcription accuracy at 97% tells you nothing about whether the underlying criteria are coherent. An AUC score tells you the model is consistent, not that it is correct. You can have a very precise instrument pointed at the wrong thing.

## Where Rubrics Come From (and Why That Should Worry You)

Most organisations running AI scoring did not write their rubric from first principles. They inherited it. It came from a vendor's default template, from a QA framework built for a contact centre in 2019, from a sales methodology the previous VP liked, or from a competency library nobody has touched since it was imported into the LMS. The criteria were plausible enough that nobody pushed back. The tool was configured, the pilot ran, the dashboard looked professional, and it went live.

If you run enablement or QA at scale, ask yourself this: could you answer, under mild questioning, why each criterion is worded the way it is, what observable behaviour it maps to, and what the last time was that two independent humans scored the same call against it and compared notes? If the answer is uncertain on any of those, the rubric has not been validated. It has been deployed.

## The Three Failure Modes

Across rubric teardowns and scoring disagreements, three problems appear repeatedly. They are not exotic edge cases. They show up in almost every rubric that has not been deliberately tested.

**Criteria that conflate two distinct behaviours**

A criterion reads: "Rep establishes credibility and builds rapport early in the call." That is two things. Credibility comes from reference points, track record, relevant context. Rapport comes from tone, reciprocity, responsiveness to the other person's signals. A rep can do one well and the other badly. When a single score covers both, you do not know which behaviour drove it, which means you cannot coach from it. The AI will score it however its training data happened to weight the blend. You will never know which half is moving.

**Criteria that require context a transcript cannot supply**

"Rep appropriately adjusts pitch based on prospect reactions." Adjust how? Based on what reactions? A transcript shows words. It does not show the three seconds of silence before the prospect said "interesting." It does not show that the rep steamrolled a genuine buying signal because they were on a script cadence. A human scorer watching a recording might catch it. An AI scoring a transcript cannot, and if the rubric does not flag that this criterion requires audio or video, the model will produce a score anyway. Quietly. With apparent confidence.

**Criteria where the 'good' anchor description is contested among experts**

Ask five experienced salespeople to score the same two-minute discovery exchange against a criterion like "Rep demonstrates strong questioning technique." You will not get agreement. You will get a debate about whether open questions are categorically better, whether hypothesis-led questions are appropriate at this stage, and whether the rep asking fewer, sharper questions outperforms the one asking more questions that the rubric's author would have recognised as textbook. The anchor for "good" is contested. The AI has picked one interpretation, silently, and is running it at scale across your whole team.

These three failure modes do not require the model to be broken. The model can be working exactly as designed. The rubric is the problem.

## Why Nobody Notices Until It Is Too Late

The output of AI scoring looks authoritative. It is numerical, it is fast, and it is consistent, which most humans are not. Consistency gets mistaken for validity. When a manager sees that the system scored a call 6.2 and the rep scored it 8, the working assumption is that the rep is being generous with themselves. The possibility that the rubric has a contested criterion at position four does not come up.

The second reason is that rubric validation requires effort that has no immediate visible output. Running the model on another thousand calls produces a dashboard update. Running a blind-marking exercise with three senior practitioners produces an argument, a set of handwritten notes, and a revised criterion document that will not impress anyone in a steering committee. The incentive structure pushes toward deployment and away from audit.

## A Sampling Protocol You Can Run in a Day

You do not need a psychometrician. You need a morning, three experienced people, and a willingness to sit with uncomfortable disagreement.

1. Pull ten scored calls from the last thirty days. Aim for a spread: two from the top decile, two from the bottom, six from the middle.
2. Strip the AI scores. Do not let your panel see them.
3. Have each panel member score all ten calls independently, against the existing rubric, with no discussion.
4. Compare scores. For any criterion where the range across panel members is greater than two points on a ten-point scale, flag it immediately. That criterion is not reliably interpreted.
5. For each flagged criterion, ask the panel to write, in one sentence, what behaviour they were looking for. Read those sentences aloud. If they differ, you have found a contested anchor.
6. Now show the AI scores. For the calls where AI and all three humans agree, note what the criterion is measuring. For the calls where the AI diverges significantly from all three humans, note what information might be missing from a transcript that a human could infer.

The output of this exercise is not a verdict on whether your AI scoring is good or bad. It is a map. Criteria that survive it, where humans agree with each other and the AI tracks them, are solid. Criteria that fail it need to be split, reworded, or removed before they produce another thousand scores.

The [Human vs AI Scoring Agreement Checker](https://www.peerlab.ai/library/tools/human-vs-ai-scoring-agreement-checker/) is built for exactly this comparison step, if you want to run it more systematically across a larger sample.

## What Validity Actually Means Here

A criterion is valid if it is observable, unambiguous, and consistently interpreted by trained scorers without prior discussion. That is not a high bar. Most criteria in inherited rubrics fail it. The exercise above tells you which ones.

Rubric validity is a prior question to everything else in your AI scoring stack. The model's accuracy, the calibration of its confidence thresholds, the fairness of the scoring distribution across different rep demographics, all of those questions assume you have a rubric worth executing. If you have not answered the prior question, the downstream numbers are not data. They are a consistent story with an unexamined premise.

The thousand judgements that went out last month were not neutral. They shaped how managers coached, how reps saw themselves, and possibly how some of those reps were evaluated. Worth knowing whether the criteria behind them had ever been tested before they went to work.
