Rubricon Score

Every call scored against the rubric your QA team already uses.

Not a generic scorecard, and not a sentiment number. Rubricon Score reads the whole call, marks each line of your existing rubric, and quotes the moment in the transcript it marked on — so a score can be argued with.

Coverage

What changes when coverage goes from a sample to all of it.

Manual QA is bounded by scorer hours, so it samples. A sample tells you the average and hides the individual call. Reading every call changes what the output is for.

01

The rare failure becomes findable

The compliance miss, the wrong policy quoted, the abusive caller — these are exactly the calls a random sample is least likely to contain, and the ones that cost the most when they are missed.

02

Coaching becomes specific

An agent's whole month is scored, so the note is "you skipped the disclosure on nine of these calls, here they are" rather than a figure derived from four calls.

03

Scorer time moves upstream

Your QA team stops marking routine calls and spends the hours on calibration and on the disputes — the work that actually needs a person's judgement.

Calibration

A score nobody has checked is a number, not a measurement.

Before any score is shown to a supervisor, the model has to agree with your own scorers on calls they marked by hand. That set stays in place afterwards as the thing new versions are re-tested against.

01

Ground truth

Around 200 past calls, marked by your scorers on your rubric, covering every language the queue takes. Their judgement is the target; the model is never the reference.

02

Agreement

Per rubric line and per language, never pooled. A rubric can agree at 90% in English and 60% in Tamil, and a single number hides exactly that.

03

Disagreements read

Every mismatch is looked at. Some are model error; some show the rubric line was ambiguous and needed rewriting anyway.

04

Monthly reconciliation

A fresh hand-scored batch each month, checked against the live scores, so drift surfaces as a number instead of as a complaint.

The rubric is versioned. Change a line and the affected history is re-scored under the new version, so a trend never mixes two definitions of the same check.
What gets scored

Whatever is on your scorecard, plus what the audio itself says.

Rubric lines
Your checks, in your wording — greeting and identification, verification, problem capture, resolution, next step, close. Scored with the transcript line quoted as evidence, so a supervisor can overturn it and the overturn is recorded.
In-language evidence
The quoted line is in the language it was spoken in, and the supervisor who overturns the score reads that language. A pipeline that scores an English translation is grading a summary, not the call — and the mistranslation becomes invisible at exactly the moment it matters.
Compliance
Mandatory disclosures, consent, prohibited claims and language, and any script obligation specific to your regulator. Pass or fail per call, with the timestamp.
Conversation mechanics
Dead air, hold length, talk-over, interruption, monologue length, speaking rate. Read from the diarised audio rather than from the words.
Outcome
What the call was actually about and whether it ended resolved, escalated, abandoned or booked for follow-up — as structured fields, reconcilable against your CRM's own disposition code.
Reasons behind the volume
Repeat contact drivers clustered across the whole month's calls, which is a question a sample cannot answer and a disposition dropdown answers badly.
Overrides
A supervisor can change any score. The change, who made it and why are kept — and overrides that repeat on the same rubric line are what tell you the line, not the agent, is the problem.
Human and agent, one rubric

The same scorecard grades your people and our bot.

This is the part that is hard to assemble from two vendors. When the agent and the human team are measured on one versioned rubric, the comparison is real: you can see which call types the agent handles at or above your team's standard, and hand it those and no others.

It also means Rubricon is graded by its own audit layer, on your definition of a good call. If the agent regresses after a change, it shows up in your dashboard on your rubric — not in a status page of ours.

Starting

It runs on recordings you already have.

Nothing has to change in the call flow to find out whether the scoring is any good. That is deliberately the cheapest thing to test first.

Give us a hundred recordings your team has already scored, and your rubric as it stands. We score them, and you look at where we disagree with you. The disagreements are the useful output — they tell you whether the model is wrong, or whether two of your own scorers would have disagreed there too.

Send the rubric and a hundred calls.

No integration, no change to your telephony, and you keep the comparison whatever you decide afterwards.