How AI Grading Actually Scores a New Agent's Practice Call

A new hire finishes their fifth practice call of the morning, takes off the headset, and a score appears on the screen before they have even stood up. No supervisor was listening in. No one will listen to this recording for another two weeks, if ever. The score is already there, broken into parts, with two lines flagged for a coach to look at later.
That moment, not the finished dashboard report, is what "AI grading" actually looks like for a new agent. It is worth being precise about what is happening inside it, because the word "grading" implies a level of judgment the system does not fully have, and understanding that gap matters more during onboarding than at any other point in an agent's tenure.
AI grading of a practice call works by transcribing the audio, then checking that transcript against a scorecard of specific, mostly rule-based criteria (greeting, disclosures, required questions, script adherence), and returning a score with flagged moments within seconds of the call ending. It is reliably accurate on binary, observable items. It is far less reliable on subjective judgment calls like tone, empathy, or de-escalation, which is exactly the territory a nervous new agent is most likely to struggle in for reasons that have nothing to do with competence.
What "AI grading" actually means for a practice call
The phrase gets used loosely, so it is worth separating from adjacent ideas. AI grading is not sentiment analysis guessing how a call "felt," and it is not a generic AI listening for vague qualities like "professionalism." It is rubric matching: a defined scorecard, applied consistently, to a transcript.
This distinction matters specifically for onboarding because a rubric built for evaluating a five-year veteran's live customer calls and a rubric built for a first-week trainee's practice reps are answering different questions. One is asking "did this call meet the bar." The other should be asking "is this specific skill developing in the right direction." Most of what is written about AI call scoring online is aimed at the first question, evaluating production call quality across an established team, not the second.
How the scoring engine reads a call, step by step
The mechanics are consistent across most tools in this category, according to vendor and industry write-ups on the process (Aircall, JustCall):
- The call (or in onboarding, the practice roleplay) is recorded and converted to a transcript via speech-to-text.
- Each item on the QA scorecard is checked against that transcript by the model, looking for specific, observable language or structure.
- The system produces a score per item, an aggregate score, and time-stamped flags on the moments that triggered a pass or fail.
- Results land on a dashboard, usually within moments of the call ending, without a human reviewer in the loop at this stage.
For a new agent, step 3 is the part that changes onboarding the most. A trainee does not wait days for a supervisor to sample their calls. They see, almost immediately, which specific line in a scorecard tripped them up on rep four versus rep twelve.
| Criteria type | Example item | How reliably AI grades it |
|---|---|---|
| Compliance / binary | Stated required disclosure, verified caller identity, captured a callback number | High. These are pattern-matchable in the transcript. |
| Structural / behavioral | Avoided interrupting in the first 30 seconds, followed the call flow in order, used an approved opening | Moderate. Reliable only when the rubric item is written specifically enough to be checked against text. |
| Tone / judgment | Sounded empathetic, de-escalated a frustrated customer, matched the customer's energy | Lower. These are the categories AI scoring tools and QA analysts alike flag as the weakest fit for automated grading (JustCall). |
What the AI grades well, and where it gets a new agent's score wrong
This is the part L&D leads evaluating these tools tend to skip past, and it matters more for someone on day three than for a tenured agent. A new agent's practice call will often score poorly on tone criteria not because their instincts are bad, but because they are still consciously assembling the script in real time. Their voice is flatter, their pauses are longer, their phrasing is stiffer. To a model checking for "warmth markers" in a transcript, that reads as a deficiency. To a coach who has trained new hires before, it reads as week one.
The practical implication: a low behavioral or tone score on an early practice call is a weaker signal than the same score on call fifty. Treating both the same way, as if the rubric is equally diagnostic at every stage of ramp, is where an otherwise well-designed grading system starts giving L&D teams bad information. The fix is not to distrust the scoring engine. It is to weight compliance and structural items more heavily in week one, and hold the tone and judgment categories more loosely until an agent has enough reps for those scores to stabilize.
Building a rubric that is fair to someone on their first week
Rubric design is consistently cited as the step that determines whether AI grading is useful at all, more than the underlying model itself (Insight7). A vague item like "was professional" gives the system nothing concrete to check. A specific item like "did the agent avoid interrupting the customer in the first 30 seconds" gives it something a transcript can actually confirm or fail.
For onboarding practice calls specifically, a first-two-weeks scorecard typically needs to look different from the standing production QA form. Items worth including early:
- Opening compliance: name, department, and any required disclosure stated within the first 15 to 20 seconds.
- Information gathering: the specific fields the call is supposed to capture (account number, issue category, callback number) were actually requested.
- Script or flow adherence: the call followed the intended structure in roughly the right order, allowing for natural conversation.
- One behavioral item, not five: a single, clearly observable behavior (not interrupting, confirming understanding before moving on) rather than a long list of soft-skill judgments a new hire cannot reasonably be held to yet.
- Resolution or next-step confirmation: the call ended with a clear next step stated, whether that is a resolution, an escalation, or a callback time.
Keeping the early rubric narrow and concrete is what makes the AI's grading trustworthy at this stage. A crowded scorecard full of subjective items just produces noisy scores that a coach then has to explain away.
Why scoring every practice call, not sampling, changes onboarding
Traditional QA in most call centers samples a small fraction of calls. Industry coverage of this points to roughly 4 to 6 calls per agent per month as a common target, with a large share of centers falling short of even that (Aircall). For a tenured agent handling hundreds of calls a month, that sampling gap is a known limitation. For a brand-new agent doing twenty or thirty practice reps in their first week alone, sampling one or two of them is close to useless as a feedback mechanism.
Grading every practice call changes what onboarding feedback actually looks like. Instead of a trainer sitting in on a handful of sessions and hoping they caught a representative sample, every rep gets scored, and patterns show up across reps two, five, and nine rather than being inferred from a single spot check. That shift, full coverage instead of sampling, is really the core value proposition of automated grading in a training context specifically, more than in ongoing production QA where the agent's behavior is already established.
Platforms built around this idea tend to structure it the same way: a simulated call or roleplay scenario, graded automatically against a scorecard, with flags surfaced to a coach rather than buried in a report nobody opens. Eduqat's roleplay and grading tools follow that same structure for onboarding specifically, scoring every practice rep an agent runs rather than a sampled few, so a trainer sees the pattern building across a new hire's first calls rather than a single snapshot of one.
Where a human coach still has to step in
None of the above makes a human trainer optional. The AI's job is triage: it tells a coach where to look and cuts out the hours spent listening to calls that were fine. The coach's job is everything the model cannot do reliably yet, interpreting whether a rough tone score reflects nerves or a genuine gap, having the actual coaching conversation, and making the judgment call on when a new agent is ready to take a live call. Treating an automated score as a final verdict rather than a starting point for a coaching conversation is the most common way L&D teams get less value out of these tools than they should.
A reasonable operating rule for the first two weeks: let the AI grade every call, review compliance and structural scores directly, and route anything flagged on tone or judgment criteria to a human before drawing conclusions about the agent, not the tool.
Frequently Asked Questions
Is AI grading accurate for a brand-new agent's calls? It is generally accurate on rule-based, observable criteria like disclosures, required questions, and call structure. It is less reliable on subjective categories like tone or empathy, which is worth remembering since new agents often score lower there for reasons tied to inexperience rather than skill.
Can AI grading assess soft skills like empathy? It can flag patterns associated with empathy, such as acknowledgment phrases or pacing, but this remains one of the weaker areas for automated scoring according to industry analysis of AI call scoring accuracy. Most teams treat these scores as a prompt for human review rather than a standalone verdict.
How is AI grading different from a human QA reviewer scoring the same call? A human applies judgment and context but can only review a small sample of calls. AI applies the same rubric consistently across every call, but only within the limits of what a rubric item can specifically check for in a transcript.
Should a new agent's practice call scores count the same as a tenured agent's production call scores? Not directly. Compliance and structural items transfer well across both. Tone and judgment scores are less stable in the first few weeks and are better used as coaching prompts than as performance verdicts during onboarding.
How many practice calls should a new agent complete before scores are used for a real evaluation? There is no universal number, since it depends on call complexity and how much of the rubric applies. The more useful marker is whether compliance and structural scores have stabilized across several consecutive reps, not a fixed rep count.
Key Takeaways
- AI grading works by transcribing a call and checking it against a scorecard, not by judging how the call "felt." That distinction matters most for new agents, whose stiffness in week one can look like a deficiency to a model even when it is not.
- The scoring engine is reliably accurate on binary, rule-based items (disclosures, required questions, script steps) and considerably weaker on subjective categories like tone and empathy.
- A rubric built for a new agent's first practice calls should stay narrow: compliance, information gathering, one behavioral item, and resolution confirmation, rather than importing a full production QA form on day one.
- Grading every practice rep instead of sampling a few is the real shift AI brings to onboarding specifically, since traditional QA sampling was already thin for tenured agents and close to meaningless across a new hire's first dozen calls.
- Automated scores are a triage layer, not a verdict. The coaching conversation, and the judgment call on when someone is ready for a live call, still belongs to a human.
This article builds on our broader look at how onboarding itself is changing: AI in Call Center Training: What Actually Changes for Onboarding.