Evaluation models

The models that grade the models.

Every FormantAI agent is judged by a second model, trained on your QA team's labels, on every call — and that judge is what decides whether the next version ships.

A voice agent that is not measured on every call is not improving — it is drifting. We built the evaluator before we built the learner.

Four models, one verdict

Grader. Simulator. Judge. Calibration.

Grader

The evaluation model

A model trained on your QA team's labels to score every call on grounding, flow, compliance and persona — 100% of calls, not a sample.

Simulator

The customer model

Plays the customer through 1,240 scenarios: interruptions, anger, "already paid", wrong numbers — in Hindi, Hinglish and Tamil.

Gate

The release judge

Red-to-green and green-to-red diffs on every candidate. Any regression on a frozen line or a compliance rubric blocks the release.

Calibration

Humans grade the grader

Your reviewers correct a sample every week; the corrections retrain the evaluation model. The loop that grades the loop.

Recursive learning loops

Loops inside loops, each feeding the next.

The turn check produces the call grade. The call grades produce the week's training set. The week's release produces the portfolio's lessons. And human corrections to the grader retrain the grader itself — the loop that grades the loop.

Recursive loops · turn → call → week → portfolio
PORTFOLIO · QUARTERLYWEEK · WEEKLYCALL · PER CALLTURN · PER TURNevery loop feedsthe one outside it
TurnGrounding · policy · persona checks before a word is spoken
CallEvery call graded by the evaluation model; failures clustered
WeekGraded calls → fine-tune → simulate → human sign-off → release
PortfolioLessons shared across journeys and languages under one policy
The scoreboard

Red to green, green to red.

Every candidate is replayed against the golden set and diffed against the live version. Improvements are counted. Regressions block the release.

Regression scoreboard · v42 vs v41 · collections-hi
ScenarioGroundFlowComplyPersonaΔ vs v41
already_paid · hi-IN↑ was red
barge-in during amount
wrong number · Marathi
legal threat → transfer
salary-delay reason code↓ flow 0.82
disclosure line · verbatimfrozen ✓
angry customer · Tamil
branch-or-link asked twiceblocks release
1,240 scenarios · 4 rubrics · 2 languages · any red blocks release illustrative · replace with live scoreboard
Eval run · agent collections-hi
$ formant eval --agent collections-hi --candidate v42 --against v41
graded 48,210 production calls · rubric: ground · flow · comply · persona
simulated 1,240 scenarios · barge-in · already_paid · wrong_number · legal_threat
red→green +31 · green→red 0 · frozen lines exact-match 100%
grader calibration: 400 human labels this week · agreement κ 0.86
verdict release allowed · approver: client QA lead

Illustrative run. replace with a live eval run

What the grader scores

Four rubrics. Every call.

Grounding

No invented numbers.

Every amount, date and policy fact checked against the system of record. A made-up EMI is a failed call.

Flow

No loops, no repeats.

Did the agent ask branch-or-link twice? Miss "already paid"? Flow failures are clustered and named.

Compliance

Frozen lines, exact.

AI disclosure, recording consent, grievance line, calling hours — verified verbatim on every call, never learned.

Persona

One voice, every language.

Tone, pace and phrasing consistent across Hindi, Hinglish and Tamil — parity-checked, not assumed.

Human calibration · weekly
Monsample_drawn400 calls · stratified by cluster and language
Tuehuman_labelsclient QA team grades blind · rubric v3
Wedagreementgrader vs human · disagreements reviewed
Thugrader_retraincorrections become grader training data
Friregradethe week's calls regraded with the new grader

Illustrative cadence. confirm with the eval team

Humans grade the grader

The evaluator is itself a learner.

A grader that never changes will be gamed by the agent it grades. So your reviewers correct a sample every week, and those corrections retrain the evaluation model. Agreement with humans is tracked like any other metric — and reported to you.

Governance and audit
FAQ

What your risk team will ask.

Who trains the evaluation model?
It is trained on labels from your QA team (and ours, for the first weeks). It lives in your tenant, like the agent model — nothing is shared across customers.
Can the agent learn to game the grader?
That is exactly why the grader is recalibrated weekly against blind human labels, and why frozen compliance lines are checked by exact match, not by a model.
What blocks a release?
Any green-to-red on a compliance rubric, any frozen-line mismatch, or an overall regression on the golden set. A named human still signs every release.
Can we see the scoreboard?
Yes — every deployment gets the scoreboard and the eval run log with the daily report. Reference scoreboards are available under NDA.
Under NDA

See one week's scoreboard.

See a live scoreboard
Readiness

Know what your calls can teach

Get a Call Readiness report