Voice agents · Call quality

Voice agents that take the call, and an audit layer that scores every one of them.

Rubricon runs phone conversations end to end over the telephony you already have — in Hindi, English and the regional language the caller switches into. Every call it handles is transcribed and scored against a rubric you define, and so is every call your human team takes.

Your SIP trunk Hindi, English & regional Barge-in & warm transfer 100% of calls scored Your rubric, versioned
On call Outbound · Order follow-up 02:14
AgentAapka order Bhiwandi hub se nikal chuka hai — main tracking link abhi text kar deta hoon.
CallerAur agar main ghar pe nahin hua delivery ke time?
Tooldelivery.reschedule_options(awb) → 3 slots
AgentChecking delivery options…
Scored by Rubricon ScoreRubric v4
Identity & disclosure Pass
Resolution offered Pass
Consent captured Pass
Language matched caller Pass
Next step confirmed Pass
Hold time 18s
Rubric score92 / 100
The loop

Every call produces a score, and the score is what changes the next call.

Voice vendors ship an agent and leave measurement to someone else. Quality vendors score the calls but have nothing to fix. Rubricon owns both ends, so a finding in the audit becomes a change in the agent and the next audit shows whether it worked.

01

Call

The agent takes the call, or your team does. Same trunk, same queue.

02

Transcript

Diarised, timestamped, with hold and talk-over marked.

03

Score

Graded line by line against your rubric, with the evidence quoted.

04

Finding

A named failure — missed disclosure, wrong policy, dead air — not a number alone.

05

Change

Agent prompt, tool or escalation rule for the bot; a coaching note for a person.

Back to 01. The changed agent is graded on the same rubric version, so the effect of the change is measurable rather than asserted.
What you deploy

Two halves of one system.

You can run either on its own. Most teams start with scoring, because it works on the calls they already have and needs nothing changed in the call flow.

Rubricon Voice

Agents that hold a phone conversation

Streaming speech recognition into a language model into streaming speech synthesis, so the agent starts answering before the caller has finished. It interrupts and can be interrupted, calls your systems mid-conversation, and hands to a human with the context attached.

  • Inbound support, qualification and routing
  • Outbound follow-up, booking and reminders
  • Tool calls into CRM, order and ticketing systems
  • Warm transfer with the transcript already on the screen
Rubricon Score

A quality agent that grades every call

Sampling five per cent of calls tells you about five per cent of calls. Rubricon Score reads all of them against the rubric your QA team already uses, quotes the line it scored on, and is calibrated against a set your own scorers marked by hand.

  • Your rubric, versioned — not a generic scorecard
  • Hand-scored ground truth before any score is trusted
  • Compliance and disclosure checks on every call
  • Per-agent coaching notes, and drift tracked over time
100%

of calls scored — agent-handled and human-handled alike, on the same rubric.

~200

calls hand-scored by your team at kickoff, as the ground truth the model is measured against.

<1s

target turn latency, measured per deployment and reported, not claimed as a headline.

1

rubric across both. Change it once and the bot and the humans are re-graded together.

Indian calls

An Indian queue is not a US English queue with a different accent.

Speech systems are overwhelmingly tuned on American English telephony audio. The calls our customers take sound nothing like it, and that gap is where most voice deployments in India come apart.

01

Code-switching is the default

One sentence carries Hindi grammar, English nouns and an English number. Deciding the call is “Hindi” and running one model against it loses the half of every sentence that is in the other language.

02

One language, many accents

Hindi from Patna and Hindi from Indore are different acoustic problems, and a national queue takes both within the same hour. Coverage is measured against your recordings rather than assumed from the language name.

03

Scoring has to work in-language

A rubric applied to a Tamil call must quote the Tamil line it scored on, and the supervisor who overturns it reads Tamil. Scoring an English translation grades a summary, not the call.

Languages in scope are Hindi, English, Tamil, Telugu, Marathi, Bengali, Kannada, Malayalam, Gujarati and Punjabi. Which of them work well enough for your queue is answered by running a sample of your own audio before anything goes live — not by publishing a list.
How it goes in

Scoring first, then the agent on one call type.

Nothing is pointed at a live caller until the rubric agrees with your own scorers on calls you have already handled.

Step 01

Rubric and ground truth

We take your existing scorecard as it is. Your QA team hand-scores a set of past calls; that set is what the model has to agree with.

Step 02

Score the back catalogue

Recordings you already hold get scored and reconciled against the hand-scored set. Disagreements are read, not averaged away.

Step 03

One call type, live

The agent takes a single, bounded call type on your trunk, with a human escalation path from the first call.

Step 04

Widen on evidence

Scope grows when the scores say it should. Each new call type repeats steps one to three rather than inheriting trust.

Built on

A real-time audio pipeline, run as our own infrastructure.

Voice is not a request-response workload. Audio arrives continuously, the model has to answer before the sentence ends, and the recording has to be stored, transcribed and scored afterwards. That shape decides the stack.

GPU inference — ASR + TTS Streaming WebSocket media plane SIP / WebRTC telephony Containers, autoscaled per concurrent call Object storage for recordings Vector store for policy retrieval Postgres for rubrics, scores and audit trail Per-turn latency telemetry

Bring us a hundred recordings.

The fastest way to judge this is to have us score calls you have already scored yourself, and compare. No integration required to start.