Peer Book a call Plug Peer in

The human check on AI scoring.

An add-on to Peer: the human check on the calls Peer already grades.

AI can score a person in seconds. It cannot check itself. A standing panel grades a sample blind and shows you, criterion by criterion, where it agrees with the AI and where it diverges.

98%
Inter-rater agreement within one band on a recent cohort. The marking is consistent, not opinion.
Blind
Assessors never see the AI's score, the report they're checking, or the person's real name. That's what makes the comparison mean anything.
1 in 5
One conversation in five is marked twice, by assessors working apart. The reliability figure is published each quarter.
The problem

A score can't check itself.

More and more organisations let AI judge how people perform: roleplay tools, call-scoring engines, interview sifts, capability platforms. Fast, and it scales. But when a client asks "how do we know these scores are right?", the honest answer is that the system that produced them cannot be the one that checks them. Neither can the party whose programme is being assessed.

The gap

No second read

An AI vendor grades and reports on its own output. Nobody outside it reads the same calls, so the number is trusted, not checked.

The risk

Scores about real people

These grades end up in performance reviews, capability reports and client deliverables. If one is challenged, someone has to stand behind it.

What's missing

The human line

A trained assessor, working blind, marking a sample against the same rubric. And a measured agreement figure showing where the AI can be relied on and where it can't.

What it is — and what it isn't

Built for a baseline, not a verdict.

It is

A baseline by team and by competency

Hundreds of whole calls a quarter, graded to one standard, show where the gaps concentrate — which team, which stage of the call, which skill.

It isn't

A verdict on an individual from one call. A single call surfaces roughly a third of the skills in the standard; nobody should be ranked or coached off one human read, and the portal says so wherever a number could be misread.

It is

The human check on AI grading

Peer grades every call; this is how you know whether to rely on it — criterion by criterion, with an agreement figure.

It isn't

A coaching platform, or real-time. It runs after the call, on a sample. Coaching depth per rep comes from Peer, which grades every call.

It is

Evidence that survives being questioned

A grade a revenue officer, HR or a client can challenge is answered with the moment, the rubric line and the assessor reference.

It isn't

A substitute for grading everything. Human hours are finite; the sample is designed to be reliable at team level, not exhaustive.

How it works

Blind marking, then a measurement.

You decide what share of each client's conversations gets a blind human mark: a spot check, half, or all of them. We handle the rest.

1

Conversations come in

Recorded roleplays or real calls, with the AI's scores alongside. Everything is anonymised before an assessor sees it. See Data & privacy.

2

The panel marks a sample

Trained assessors mark against your own rubric, blind: no sight of the AI's score or the person's name. One in five is marked twice, by two assessors working apart. Every conversation is also marked against the Peer Standard, at no extra cost.

3

Agreement, quantified

Per criterion, we show where the human marks and the AI agree, where they diverge, and by how much, with the evidence clip behind every judgement.

4

A reliability statement you can share

Each quarter, a signed statement of how consistently the panel agreed: something you can put in front of your own clients as proof the scoring has been checked by people. See an example →

Two marks, two questions

Your rubric says whether the scoring was applied consistently. The Peer Standard says where your people actually stand.

A rubric is local: it compares your people to your own expectations, so a strong result proves internal consistency and nothing more. The Peer Standard does not change between clients, which makes it the one part of your report that means something outside your organisation.

Both run on every conversation you send, at no extra cost: the human panel, and Peer’s specialist AI graders, built to score demonstrated selling skill against the Standard, each grade carrying the moment that earned it.

So the check runs both ways. The panel stops the machine drifting; the machine stops the panel drifting. When they disagree, you see the disagreement rather than an average that hides it.

Read the full Peer Standard →

Cycle sizes and cadence

Sized to the team, priced per cycle.

Quarterly is the default and what the reliability statement is built around. Twice-yearly and annual cadences are available for a lighter programme; a one-off Baseline can precede any of them.

Standard · £6,500 per cycle

150 whole calls

One in five double-marked. Teams up to about 100 reps — every rep reached about twice a year on a quarterly cadence.

Enterprise · £11,500 per cycle

300 whole calls

Per-team reliability reporting. For 250+ reps — about one call per rep per quarter.

Baseline · £40 per call

Three calls per rep, one-off

For a chosen cohort, minimum 150 calls. The first read before quarterly cycles begin, and the only design that gives a per-rep starting point. 60 reps ≈ £7,200; 250 reps ≈ £30,000.

Sizing guide

What a quarterly Standard cycle gives you, by headcount

Up to 100 reps — about 1.5 calls per rep per quarter, a genuine team baseline: Standard, quarterly.

100–250 reps — 0.6 to 1.5 per rep per quarter, team-level only: Standard plus a one-off Baseline, or Enterprise.

250+ reps — under 0.6 per rep, too thin to call a baseline: Enterprise (about one per rep per quarter), with a Baseline for any cohort you need read in depth.

Calibration setup — rubric mapping, anonymisation pipeline and marker training — is included for Peer customers. Download the service sheet (PDF) →

Who asks for it

If a Peer grade decides something about a person, this is the line under it.

Peer grades every call. Once a grade influences who gets coached, promoted or performance-managed, someone inside the business will be asked to justify it. These are the people who ask for the human check.

Revenue leaders

A baseline by team and by competency that you can defend to a board, a revenue officer or the sellers themselves — with the moment behind every grade.

Enablement, L&D and people teams

Grades that reach reviews and development plans need a human line under them. Hold the agreement figure before the question arrives.

Procurement, compliance and works councils

Where a capability score about a named individual has to survive being questioned, the clip, the reasoning, the rubric version and the assessor reference all travel with it.

Data, privacy & anonymity

Built so there's nothing sensitive to expose.

This section answers the questions a client's data or compliance team will ask before a single conversation is assessed. It's written to be shared. Send it on, or point a prospect straight here.

Exposure & anonymity
Are our clients' real conversations exposed to your assessors? +
What we assess

We assess the conversations you send us. We don't run the roleplays or record the calls ourselves. So the answer depends on what you send, and in both cases it's designed so there's nothing sensitive to expose.

Roleplays. We're often asked to assess roleplays against an AI buyer. There, the person is talking to a simulated counterparty, not a real customer, so there is no genuine third party on the line and no information about your clients' customers is present at all. The assessment is of how the person handled the conversation: the skill, not the deal.

Real sales calls. Where you ask us to assess real commercial conversations, which we also do, that's where our protocols come in: full anonymisation of the people involved, and (where the content warrants it) a sanitisation step that strips business and personal detail before an assessor sees anything. Both are covered below.

Are the individuals being assessed identifiable? +

No. Every person is given a code, for example P-07. The assessor sees only that code and the anonymised transcript. The AI's scores are attached to the record after the human mark is submitted and locked, never before. They never see a name, an email address, a job title, or any other identifying detail.

Personal details can occasionally surface in the course of a conversation (a first name spoken aloud, say), but they are never the focus, are not surfaced deliberately, and results are only ever reported against the code. The key that maps a code back to a person is held by you, not by us, so on our side the marking cannot be tied to a named individual at all.

What if a conversation could contain sensitive personal or business information? +
Sanitisation gate

Before a conversation reaches an assessor, it can be run through a series of AI agents that sanitise and anonymise it: stripping business context, pricing, product detail, and personal information such as names, phone numbers, addresses and email addresses. Only the cleaned version is delivered for marking. For real-conversation work this is the default; for roleplays, where there's no real party on the line, it's applied where a client wants the extra safeguard.

Two things worth being explicit about. First, sanitisation completes before any assessor sees the material: the cleaning is automated, and whoever processes it is named in our DPA and contractually barred from training on your data (more on that below). Second, we're open about the trade-off: the sanitising agents can be over-zealous, occasionally removing context that a grade actually depends on, or cutting a conversation in a way that makes some criteria harder to assess. Because business or personal detail sometimes ties directly to a judgement, heavy sanitising can very slightly reduce what an assessment can see. It's a deliberate exchange of a little precision for a lot of privacy, applied where it's warranted rather than blanket.

How your data is handled
Does our data ever go to a third-party AI, or get used to train models? +
Self-hosted · no training

No, on both counts. This is where a lot of similar offers quietly fall down, so we're specific about it. The agents that sanitise and anonymise conversations run before any human assessor is involved, so identifiable material never reaches a person here. Any AI provider involved in that step is named as a sub-processor in our DPA, is bound by enterprise terms that prohibit training on your content, and holds it only for the limited period recorded there. Nothing you send us is ever used to train a model, ours or anyone else's, and no copy survives on our side once your report is issued.

The only AI scoring in the picture is the one you supply alongside the conversation, the output we're checking. Our role adds a human panel and, where used, in-house sanitisation. No part of the pipeline hands your material to an outside model.

Who can actually see the data, and under what terms? +

Access is limited to a small, named panel of trained assessors, each under a written confidentiality agreement, on least-privilege access that is authenticated and logged. Assessors work through an anonymised interface and cannot export raw material.

Each client's data is held in its own separate silo: separate storage and separate export, never pooled with, or visible to, another client. We use your data solely to perform the assessment you've asked for: never to benchmark other clients, populate a shared library, or any purpose beyond the engagement.

Where is it stored, and can you keep it in a particular region? +

Conversations and results are held on established cloud infrastructure with encryption in transit and at rest. Access is authenticated, logged, and limited to the people who need it. Data is retained only as long as needed to deliver and stand behind the assessment, and is deleted on request, with a deletion certificate provided. Retention periods can be fixed in the contract.

Hosting-region and residency requirements are scoped per engagement: tell us where your data must and must not sit, and we'll confirm what we can commit to in writing before a single conversation is processed, rather than claim a blanket guarantee that wouldn't hold for every case.

Legal, security & compliance
Do you sign a DPA, and what's the GDPR position? +

Yes. You are the data controller and we act as your data processor under a written Data Processing Agreement. The data is yours, processed only on your documented instructions and only to perform the assessment. Our standard DPA covers purpose limitation, confidentiality, security measures, sub-processor terms, assistance with data-subject requests, international transfers, breach handling and deletion. Read the full DPA →

We maintain a current list of sub-processors and share it on request, with notice of any change. Where a processing activity or transfer requires them, standard contractual clauses (or the UK IDTA) are included. If your legal team would rather work from your own paper, send it over. We'll review and sign from there.

What are your security controls, and do you hold certifications? +

Controls in place today: encryption in transit and at rest; authenticated, logged, least-privilege access; per-client data isolation; automated sanitisation ahead of any human review, with every processor named in the DPA and barred from training on your content; and a confidentiality agreement on every assessor.

On formal certification, we'll be straight with you. We don't yet hold SOC 2 or ISO 27001. Our controls are modelled on the ISO 27001 framework and formal certification is on our roadmap. In the meantime we're glad to complete your security questionnaire, share our controls documentation, and support a due-diligence review. We'd rather show you where we actually are than display a badge we haven't earned.

What happens if there's a security incident? +

If an incident affecting your data occurred, we would contain it, investigate, and notify you without undue delay (with what we know, what's affected, and what we're doing about it) and support any notifications you in turn need to make to your own clients or to a regulator. The specific notification window (72 hours) and the responsibilities on each side are set out in the DPA, so they're contractual, not just a promise on a page.

Scale & continuity
Can you handle enterprise volume, and what about turnaround and continuity? +

Yes. The panel is a standing resource rather than people assembled per project, so volume scales without retraining from scratch each time. We agree throughput and turnaround with you up front and hold capacity against a committed volume, with a service level set in writing for large or ongoing programmes.

Continuity is built into the method: because a sample of every batch is double-marked and every assessor marks to a common standard, no single marker is a point of failure. Work re-routes within the panel without changing how a score is reached. The reliability figure we publish each quarter is also the early warning if consistency ever drifts as volume grows.

Integrity of the marking
How is the marking kept honest? +

Blindness is enforced in the system, not by policy. An assessor never sees the AI's score, the vendor's report, another assessor's marks, or a real name, because the moment a marker is anchored to the answer they're supposed to be checking, the comparison is worthless.

One conversation in five is marked twice, by a second assessor working apart, and the inter-rater agreement is measured and published each quarter. That figure is what turns a mark from an opinion into a measurement, and it's reported in full, not just when it flatters us.

With Peer

Peer grades every call. The panel checks Peer.

Peer grades every one of your sales calls competency by competency, moment by moment. Blind Human Call Grading adds the standing panel: a blind human read of a sample of the same calls, criterion by criterion, with a reliability figure every cycle. Setup is included — the rubric, the anonymisation pipeline and marker training already exist for a Peer workspace.

Standard cycles from £6,500 a quarter alongside your Peer agreement; Enterprise cycles and a one-off Baseline for larger teams — sizes and prices above.

Said plainly: the check is performed by PeerLab’s own marking panel under the published blind method — stated openly, not badged as third-party.

Get started

Add it to your Peer agreement.

Human grading by a standing panel that never sees Peer's score. Tell us the team and the cadence and we'll size the cycles.

Read our Services Agreement and Data Processing Agreement, or email [email protected]. We'll send the full pack straight over.