The models that grade the models.
Every FormantAI agent is judged by a second model, trained on your QA team's labels, on every call — and that judge is what decides whether the next version ships.
A voice agent that is not measured on every call is not improving — it is drifting. We built the evaluator before we built the learner.
Grader. Simulator. Judge. Calibration.
The evaluation model
A model trained on your QA team's labels to score every call on grounding, flow, compliance and persona — 100% of calls, not a sample.
The customer model
Plays the customer through 1,240 scenarios: interruptions, anger, "already paid", wrong numbers — in Hindi, Hinglish and Tamil.
The release judge
Red-to-green and green-to-red diffs on every candidate. Any regression on a frozen line or a compliance rubric blocks the release.
Humans grade the grader
Your reviewers correct a sample every week; the corrections retrain the evaluation model. The loop that grades the loop.
Loops inside loops, each feeding the next.
The turn check produces the call grade. The call grades produce the week's training set. The week's release produces the portfolio's lessons. And human corrections to the grader retrain the grader itself — the loop that grades the loop.
Red to green, green to red.
Every candidate is replayed against the golden set and diffed against the live version. Improvements are counted. Regressions block the release.
Illustrative run. replace with a live eval run
Four rubrics. Every call.
No invented numbers.
Every amount, date and policy fact checked against the system of record. A made-up EMI is a failed call.
No loops, no repeats.
Did the agent ask branch-or-link twice? Miss "already paid"? Flow failures are clustered and named.
Frozen lines, exact.
AI disclosure, recording consent, grievance line, calling hours — verified verbatim on every call, never learned.
One voice, every language.
Tone, pace and phrasing consistent across Hindi, Hinglish and Tamil — parity-checked, not assumed.
Illustrative cadence. confirm with the eval team
The evaluator is itself a learner.
A grader that never changes will be gamed by the agent it grades. So your reviewers correct a sample every week, and those corrections retrain the evaluation model. Agreement with humans is tracked like any other metric — and reported to you.
Governance and audit