Rubricon runs phone conversations end to end over the telephony you already have — in Hindi, English and the regional language the caller switches into. Every call it handles is transcribed and scored against a rubric you define, and so is every call your human team takes.
Aapka order Bhiwandi hub se nikal chuka hai — main tracking link abhi text kar deta hoon.
Aur agar main ghar pe nahin hua delivery ke time?
Voice vendors ship an agent and leave measurement to someone else. Quality vendors score the calls but have nothing to fix. Rubricon owns both ends, so a finding in the audit becomes a change in the agent and the next audit shows whether it worked.
The agent takes the call, or your team does. Same trunk, same queue.
Diarised, timestamped, with hold and talk-over marked.
Graded line by line against your rubric, with the evidence quoted.
A named failure — missed disclosure, wrong policy, dead air — not a number alone.
Agent prompt, tool or escalation rule for the bot; a coaching note for a person.
You can run either on its own. Most teams start with scoring, because it works on the calls they already have and needs nothing changed in the call flow.
Streaming speech recognition into a language model into streaming speech synthesis, so the agent starts answering before the caller has finished. It interrupts and can be interrupted, calls your systems mid-conversation, and hands to a human with the context attached.
Sampling five per cent of calls tells you about five per cent of calls. Rubricon Score reads all of them against the rubric your QA team already uses, quotes the line it scored on, and is calibrated against a set your own scorers marked by hand.
of calls scored — agent-handled and human-handled alike, on the same rubric.
calls hand-scored by your team at kickoff, as the ground truth the model is measured against.
target turn latency, measured per deployment and reported, not claimed as a headline.
rubric across both. Change it once and the bot and the humans are re-graded together.
Speech systems are overwhelmingly tuned on American English telephony audio. The calls our customers take sound nothing like it, and that gap is where most voice deployments in India come apart.
One sentence carries Hindi grammar, English nouns and an English number. Deciding the call is “Hindi” and running one model against it loses the half of every sentence that is in the other language.
Hindi from Patna and Hindi from Indore are different acoustic problems, and a national queue takes both within the same hour. Coverage is measured against your recordings rather than assumed from the language name.
A rubric applied to a Tamil call must quote the Tamil line it scored on, and the supervisor who overturns it reads Tamil. Scoring an English translation grades a summary, not the call.
Nothing is pointed at a live caller until the rubric agrees with your own scorers on calls you have already handled.
We take your existing scorecard as it is. Your QA team hand-scores a set of past calls; that set is what the model has to agree with.
Recordings you already hold get scored and reconciled against the hand-scored set. Disagreements are read, not averaged away.
The agent takes a single, bounded call type on your trunk, with a human escalation path from the first call.
Scope grows when the scores say it should. Each new call type repeats steps one to three rather than inheriting trust.
Voice is not a request-response workload. Audio arrives continuously, the model has to answer before the sentence ends, and the recording has to be stored, transcribed and scored afterwards. That shape decides the stack.
The fastest way to judge this is to have us score calls you have already scored yourself, and compare. No integration required to start.