Same seven business tools. Matched scenario fixtures. Inspect the conversation
and the resulting state.
60recorded calls
30matched pairs
50 EN / 10 ROEnglish + Romanian
13personal listening reviews
OUR PICK FOR THIS PILOT
We prefer Gemini’s conversation experience. The tool receipts keep us honest.
Gemini leads the corrected saved-state checks and Patrick’s partial listening scores.
GPT-Live had fewer prohibited tool attempts. There is no established overall latency
winner, and this pilot does not prove which model is best for every voice agent.
Saved-state checks ask whether the sandbox meets the scenario’s requirements. They are
stricter than “the conversation sounded good” and do not alone establish complete task
success.
Gemini GPT-Live + TerraOne square = one call · select a square to inspect it
Loading the versioned measurements…
Keep the grading history visible. The original saved-state score was
Gemini 24/30 and GPT 18/30. A documented correction treats an empty optional order
reference like an absent one, bringing GPT to 20/30. Caller deviations, exact lifecycle
requirements and consent still need separate interpretation.
Read the failure audit ↗
S027 / RECOVERY
A blocked attempt. Then a correct order.
Gemini first attempted an unconfirmed write. The sandbox blocked it. After
confirmation, it saved the correct order. The state passed; the earlier attempt
remains flagged.
A failed check can mean a missing action, a blocked attempt that later recovered, a
strict grader, or an AI caller that changed its own facts. The 16 corrected state
failures are not 16 proven model mistakes.
01
Attempt ≠ commit
An unconfirmed request can be rejected before it changes anything. We show the state
before and after each tool.
02
A fluent claim ≠ a saved record
“I saved it” needs a corresponding action and a successful business result. S048
makes the difference visible.
03
The tester can be wrong
Caller drift and grading errors remain in the record. We explain them without
quietly deleting difficult calls.
Loading the frozen sandbox receipts…
Tool calls and state changes come directly from the hash-verified frozen audit. Spoken
explanations use its provisional transcript review; they are not new human audio
adjudications. confirmed:true is an argument, not independent proof of
audible agreement. The sandbox’s older-order policy allows lookup, then a human
follow-up message; it does not allow direct mutation of an older order.
02 / TIME, MEASURED CAREFULLY
A fast acknowledgment isn’t the whole conversation.
We reanalyzed the original two-channel recordings with the same speech detector for both
systems. These are automatically detected speech gaps, not
human-verified turn boundaries or pure model processing latency.
Caller speech endsDetected gap →First detected agent speech… substantive answer / saved action not timed in this release
What the gap measures
End of a detected caller speech chunk → start of the next detected agent speech
chunk, before the caller speaks again. It can be “okay” or a full answer. Phone
transport and buffering are included.
What it leaves out
Overlapping speech and chunks without a paired response are counted separately.
Neither is silently given zero latency or called a failed answer. Speech detection
does not determine conversational intent.
Why three minutes isn’t latency
The caller, task complexity and hangup behavior all affect duration. The frozen
hangup fallback missed some “bye” endings, leaving silence until the 180-second cap.
Shorter calls can also end with unfinished tasks.
Show timing sensitivity and coverage
Silero uses 32 ms frames, probability thresholds .5/.35, a minimum speech length of 96
ms, and a primary silence merge of 400 ms. The silence hold confirms a boundary; it is
not added to the timestamp. Changing the merge rule to 250 or 700 ms is a sensitivity
check, not a second experiment. No gain or playback-speed changes were made.
INSIDE THE GPT CONFIGURATION
Delegation is real. Its exact contribution needs a controlled test.
Loading captured backend timings…
The captured Terra responses used medium reasoning. Our in-memory mock tools run
quickly; waiting around a save can include backend reasoning, follow-up responses and
spoken explanation. We did not test a faster backend or reasoning setting in this
frozen study, so we cannot attribute the full conversation difference to delegation.
OpenAI’s delegation guidance ↗
03 / ONE PERSON’S EXPERIENCE
Patrick preferred Gemini.
SUBJECTIVE · PARTIAL SAMPLE
“It felt way more fast paced … and a bit more natural.” — Patrick, GRAI Labs builder. His
saved scores are reported separately from automated checks. Voices were recognizable, and
some provider guesses were wrong.
This is one affiliated reviewer, with self-selected coverage. Six rated calls are
Romanian: three per model. The full study contains only ten Romanian calls and no other
non-English languages, so it cannot establish general multilingual superiority. Scores
reflect the tested voices, prompts and backend configurations.
Original recordings at their original volume and speed. The waveform shares one
full-scale reference across both channels. Colored regions are detected speech; click a
gap below to hear the surrounding exchange.
A highlighted gap is an automatic timing candidate. Listening reviews did not include
boundary annotations; human-verified latency remains unavailable.
05 / THE COST OF A CONVERSATION
What did a call cost?
Start with the answering model, then account for the rest of the system. These are
estimates from captured usage in the 60-call study, in USD. They are not final invoices.
Gemini’s estimate is partial. All 30 answering-model snapshots contain
tokens whose type was not captured. That usage could not be priced and is excluded. The
numbers below do not establish an exact saving or a complete cost winner.
Loading cost evidence…
Which minute are we counting?
Per call is the total captured cost divided by the number of calls.
Per recorded minute divides that same cost by the summed length of
the original recordings. It includes silence and unfinished calls, counts each
two-channel recording once, and is not a provider’s billable-minute rate.
The AI caller and both post-call transcription passes are evaluation expenses. A
customer calling your agent does not require an AI caller. Phone charges here reflect
this benchmark’s call path, not a quote for your deployment. LiveKit media/SIP,
hosting, storage, tax, currency conversion and uncaptured provider usage remain
outside the known subtotal. Calibration calls are separate.
What are the published model rates?
GPT-Live 1: $0.05 per session minute, billed per second, plus backend
model usage. Our tested Terra backend is included in the captured answering-model
figures above; the voice-session fee alone is not the full model cost.
OpenAI pricing ↗
Gemini 3.1 Flash Live: $0.005 per minute of input audio and $0.018 per minute of
output audio, as listed by Google. Billing is token-based: $3/$12 per million input/output audio
tokens, plus $0.75/$4.50 per million input/output text tokens. Those audio rates refer
to different streams; adding them does not give a flat price for every minute of a
phone call.
Google pricing ↗
Published rates checked 13 September 2026. Study estimates retain the original 12
September pricing table; the evidence has not been repriced.
60 AI-to-AI phone calls: 30 matched pairs, 50 English and 10 Romanian. Seven
calibration calls were excluded. The target order was balanced; half the pairs used a
Gemini caller and half a GPT caller.
Both targets received the same seven business tools and matched task fixtures, caller
agenda, seed and time budget. All 30 pairs passed the common tool-contract audit.
Provider-specific prompt placement differs. Captured configurations prove what the SDK
requested, not provider acknowledgment of each tool.
Targets: gemini-3.1-flash-live-preview with Puck;
gpt-live-1 with marin and a gpt-5.6-terra backend. Calls
used Twilio and LiveKit. These are tested system configurations, not isolated base
models.
The simulator and post-call transcription are evaluation overhead. Shorter recording
duration does not by itself establish lower cost per successful task. Full invoice
reconciliation, some modality accounting, media/SIP charges, hosting, storage, tax and
currency conversion remain unresolved.
Different from a provider-direct test
Speko reports its own scenarios, tool-action scoring and first-audio methodology. This
study uses a phone path, a different GPT backend, sandbox-state checks and automated
speech boundaries. Our values should not be treated as a replication of
Speko’s leaderboard.
What this release does not establish
A universal winner, causal model-only advantage, or statistically reliable ranking
from one run per arm per scenario.
Human-validated conversational latency, time to the substantive answer, or time from
confirmation to an audibly confirmed save. Those endpoints were not annotated.
A clean task-success rate from saved-state checks alone. Caller mistakes, strict
grading, spoken truth and consent matter separately.
A general language ranking from English and a small Romanian sample. Both transcript
passes use OpenAI ASR and can share errors.
All recordings and original grades remain available. Timing and personal reviews are a
dated, additive analysis; the frozen study has not been rerun or quietly rescored.
GRAI VOICE BENCH · OPEN SOURCE
Give your voice agent a business regression test.
Use the eight stateful order workflows, mock tools and offline replays. Test a new
prompt against an amendment, cancellation or lost tool response, then inspect the
resulting records.
Early developer toolkit. The public release supports the documented Twilio/LiveKit
setup; turnkey arbitrary-agent onboarding and browser/Telnyx simulation transports are
still future work.
git clone https://github.com/patrick25076/grai-voice-bench
cd grai-voice-bench
# Python 3.12/3.13; no keys or dependencies
python -m voicelab.simulations demo \
--language en --output runs/first-demo
# Open runs/first-demo/index.html
This first demo uses scripted tool traces. It makes no AI or phone calls.
Our practical contribution is the connection between a voice conversation and an
inspectable business state. Use it to catch regressions in your own agent, even when
every recording sounds convincing.
01 / DEFINE
Write the business contract.
Specify the allowed records, quantities, ownership and confirmation rules. Check
that both adapters receive equivalent tool schemas. “Sounds good” is not an
assertion.
02 / ISOLATE
Run tools in a sandbox.
Use synthetic orders and deterministic state. Enforce authorization and stock rules
in code. A model-supplied confirmation boolean alone is not a production consent
control.
03 / CHALLENGE
Make the workflow go wrong.
Change the address after saving. Cancel. Run out of stock. Lose a response after
committing. Retry with the same key and assert that the order was created only once.
04 / OBSERVE
Keep the whole chain of evidence.
Record separate audio channels, tool arguments, business results and state changes.
Count overlap and unanswered chunks separately; distinguish first sound from the
substantive answer.
05 / REVIEW
Evaluate the evaluator.
Check whether the AI caller actually followed the scenario. Hear ambiguous consent
and numbers yourself. Preserve failed runs and publish grading corrections with
their reasons.
06 / RELEASE
Start with supervised actions.
For real customer operations, validate a draft in the company’s system and have a
human approve it. Expand permissions only after representative tests, monitoring and
recovery paths are in place.
Useful building blocks, with clear limits.
Simulation and evaluation are established ideas. This toolkit contributes reusable
order scenarios, stateful mock tools, failure injection, replay and linked phone-call
evidence. It complements broader model benchmarks such as Speko. It is an early
foundation for developer testing, not a claim that one small pilot certifies
production readiness.