GRAI LABSTHE VOICE OF YOUR BUSINESS
FIELD STUDY 001 / SEPTEMBER 2026

Great voice.
But did it
save the order?

Gemini 3.1 Flash Live vs GPT-Live 1.
Two voice systems, real phone calls, and a sandbox that keeps receipts.

THE PHONE TEST01 / 01
GOOGLE

Gemini

3.1 Flash Live · Puck

VS
OPENAI

GPT-Live 1

marin · Terra backend

Same seven business tools. Matched scenario fixtures.
Inspect the conversation and the resulting state.

60recorded calls
30matched pairs
50 EN / 10 ROEnglish + Romanian
13personal listening reviews
OUR PICK FOR THIS PILOT

We prefer Gemini’s conversation experience.
The tool receipts keep us honest.

Gemini leads the corrected saved-state checks and Patrick’s partial listening scores. GPT-Live had fewer prohibited tool attempts. There is no established overall latency winner, and this pilot does not prove which model is best for every voice agent.

Read the limits of the comparison
01 / WHAT HAPPENED

A result for each question.

Saved-state checks ask whether the sandbox meets the scenario’s requirements. They are stricter than “the conversation sounded good” and do not alone establish complete task success.

Gemini GPT-Live + TerraOne square = one call · select a square to inspect it

Loading the versioned measurements…

Keep the grading history visible. The original saved-state score was Gemini 24/30 and GPT 18/30. A documented correction treats an empty optional order reference like an absent one, bringing GPT to 20/30. Caller deviations, exact lifecycle requirements and consent still need separate interpretation. Read the failure audit ↗
S027 / RECOVERY

A blocked attempt.
Then a correct order.

Gemini first attempted an unconfirmed write. The sandbox blocked it. After confirmation, it saved the correct order. The state passed; the earlier attempt remains flagged.

Inspect the trace ↗
S048 / UNSUPPORTED CLAIM

“Saved” needs
a saved record.

Gemini looked up the older order, then claimed a message was saved. No take_message action occurred. Fluent speech alone would miss this failure.

Inspect the trace ↗
OPEN THE RECEIPTS

What actually failed?

60 INSPECTABLE SANDBOXES

A failed check can mean a missing action, a blocked attempt that later recovered, a strict grader, or an AI caller that changed its own facts. The 16 corrected state failures are not 16 proven model mistakes.

01

Attempt ≠ commit

An unconfirmed request can be rejected before it changes anything. We show the state before and after each tool.

02

A fluent claim ≠ a saved record

“I saved it” needs a corresponding action and a successful business result. S048 makes the difference visible.

03

The tester can be wrong

Caller drift and grading errors remain in the record. We explain them without quietly deleting difficult calls.

Loading the frozen sandbox receipts…

Tool calls and state changes come directly from the hash-verified frozen audit. Spoken explanations use its provisional transcript review; they are not new human audio adjudications. confirmed:true is an argument, not independent proof of audible agreement. The sandbox’s older-order policy allows lookup, then a human follow-up message; it does not allow direct mutation of an older order.

02 / TIME, MEASURED CAREFULLY

A fast acknowledgment
isn’t the whole conversation.

We reanalyzed the original two-channel recordings with the same speech detector for both systems. These are automatically detected speech gaps, not human-verified turn boundaries or pure model processing latency.

Caller speech endsDetected gap →First detected agent speech… substantive answer / saved action
not timed in this release

What the gap measures

End of a detected caller speech chunk → start of the next detected agent speech chunk, before the caller speaks again. It can be “okay” or a full answer. Phone transport and buffering are included.

What it leaves out

Overlapping speech and chunks without a paired response are counted separately. Neither is silently given zero latency or called a failed answer. Speech detection does not determine conversational intent.

Why three minutes isn’t latency

The caller, task complexity and hangup behavior all affect duration. The frozen hangup fallback missed some “bye” endings, leaving silence until the 180-second cap. Shorter calls can also end with unfinished tasks.

Show timing sensitivity and coverage

Silero uses 32 ms frames, probability thresholds .5/.35, a minimum speech length of 96 ms, and a primary silence merge of 400 ms. The silence hold confirms a boundary; it is not added to the timestamp. Changing the merge rule to 250 or 700 ms is a sensitivity check, not a second experiment. No gain or playback-speed changes were made.

INSIDE THE GPT CONFIGURATION

Delegation is real. Its exact contribution needs a controlled test.

Loading captured backend timings…

The captured Terra responses used medium reasoning. Our in-memory mock tools run quickly; waiting around a save can include backend reasoning, follow-up responses and spoken explanation. We did not test a faster backend or reasoning setting in this frozen study, so we cannot attribute the full conversation difference to delegation. OpenAI’s delegation guidance ↗

03 / ONE PERSON’S EXPERIENCE

Patrick preferred Gemini.

SUBJECTIVE · PARTIAL SAMPLE

“It felt way more fast paced … and a bit more natural.” — Patrick, GRAI Labs builder. His saved scores are reported separately from automated checks. Voices were recognizable, and some provider guesses were wrong.

This is one affiliated reviewer, with self-selected coverage. Six rated calls are Romanian: three per model. The full study contains only ten Romanian calls and no other non-English languages, so it cannot establish general multilingual superiority. Scores reflect the tested voices, prompts and backend configurations.

Download the 13 saved reviews ↗
04 / LISTEN FOR YOURSELF

Every call has a trace.

Original recordings at their original volume and speed. The waveform shares one full-scale reference across both channels. Colored regions are detected speech; click a gap below to hear the surrounding exchange.

CALLERAGENT

A highlighted gap is an automatic timing candidate. Listening reviews did not include boundary annotations; human-verified latency remains unavailable.

05 / THE COST OF A CONVERSATION

What did a call cost?

Start with the answering model, then account for the rest of the system. These are estimates from captured usage in the 60-call study, in USD. They are not final invoices.

Gemini’s estimate is partial. All 30 answering-model snapshots contain tokens whose type was not captured. That usage could not be priced and is excluded. The numbers below do not establish an exact saving or a complete cost winner.

Loading cost evidence…

Which minute are we counting?

Per call is the total captured cost divided by the number of calls. Per recorded minute divides that same cost by the summed length of the original recordings. It includes silence and unfinished calls, counts each two-channel recording once, and is not a provider’s billable-minute rate.

The AI caller and both post-call transcription passes are evaluation expenses. A customer calling your agent does not require an AI caller. Phone charges here reflect this benchmark’s call path, not a quote for your deployment. LiveKit media/SIP, hosting, storage, tax, currency conversion and uncaptured provider usage remain outside the known subtotal. Calibration calls are separate.

What are the published model rates?

GPT-Live 1: $0.05 per session minute, billed per second, plus backend model usage. Our tested Terra backend is included in the captured answering-model figures above; the voice-session fee alone is not the full model cost. OpenAI pricing ↗

Gemini 3.1 Flash Live: $0.005 per minute of input audio and $0.018 per minute of output audio, as listed by Google. Billing is token-based: $3/$12 per million input/output audio tokens, plus $0.75/$4.50 per million input/output text tokens. Those audio rates refer to different streams; adding them does not give a flat price for every minute of a phone call. Google pricing ↗

Published rates checked 13 September 2026. Study estimates retain the original 12 September pricing table; the evidence has not been repriced.

06 / HOW TO READ AND REUSE THIS

A small study.
An open set of receipts.

The experiment

60 AI-to-AI phone calls: 30 matched pairs, 50 English and 10 Romanian. Seven calibration calls were excluded. The target order was balanced; half the pairs used a Gemini caller and half a GPT caller.

Both targets received the same seven business tools and matched task fixtures, caller agenda, seed and time budget. All 30 pairs passed the common tool-contract audit. Provider-specific prompt placement differs. Captured configurations prove what the SDK requested, not provider acknowledgment of each tool.

Targets: gemini-3.1-flash-live-preview with Puck; gpt-live-1 with marin and a gpt-5.6-terra backend. Calls used Twilio and LiveKit. These are tested system configurations, not isolated base models.

The cost

Explore costs per call and per recorded minute ↑

The simulator and post-call transcription are evaluation overhead. Shorter recording duration does not by itself establish lower cost per successful task. Full invoice reconciliation, some modality accounting, media/SIP charges, hosting, storage, tax and currency conversion remain unresolved.

Different from a provider-direct test

Speko reports its own scenarios, tool-action scoring and first-audio methodology. This study uses a phone path, a different GPT backend, sandbox-state checks and automated speech boundaries. Our values should not be treated as a replication of Speko’s leaderboard.

What this release does not establish

  • A universal winner, causal model-only advantage, or statistically reliable ranking from one run per arm per scenario.
  • Human-validated conversational latency, time to the substantive answer, or time from confirmation to an audibly confirmed save. Those endpoints were not annotated.
  • A clean task-success rate from saved-state checks alone. Caller mistakes, strict grading, spoken truth and consent matter separately.
  • A general language ranking from English and a small Romanian sample. Both transcript passes use OpenAI ASR and can share errors.

All recordings and original grades remain available. Timing and personal reviews are a dated, additive analysis; the frozen study has not been rerun or quietly rescored.

GRAI VOICE BENCH · OPEN SOURCE

Give your voice agent
a business regression test.

Use the eight stateful order workflows, mock tools and offline replays. Test a new prompt against an amendment, cancellation or lost tool response, then inspect the resulting records.

Early developer toolkit. The public release supports the documented Twilio/LiveKit setup; turnkey arbitrary-agent onboarding and browser/Telnyx simulation transports are still future work.

git clone https://github.com/patrick25076/grai-voice-bench
cd grai-voice-bench
# Python 3.12/3.13; no keys or dependencies
python -m voicelab.simulations demo \
  --language en --output runs/first-demo
# Open runs/first-demo/index.html

This first demo uses scripted tool traces. It makes no AI or phone calls.

Try the sandbox: first-run guide ↗

FROM AN IMPRESSIVE DEMO TO A RELIABLE WORKFLOW

Test the action.
Then trust the agent with it.

Our practical contribution is the connection between a voice conversation and an inspectable business state. Use it to catch regressions in your own agent, even when every recording sounds convincing.

  1. 01 / DEFINE

    Write the business contract.

    Specify the allowed records, quantities, ownership and confirmation rules. Check that both adapters receive equivalent tool schemas. “Sounds good” is not an assertion.

  2. 02 / ISOLATE

    Run tools in a sandbox.

    Use synthetic orders and deterministic state. Enforce authorization and stock rules in code. A model-supplied confirmation boolean alone is not a production consent control.

  3. 03 / CHALLENGE

    Make the workflow go wrong.

    Change the address after saving. Cancel. Run out of stock. Lose a response after committing. Retry with the same key and assert that the order was created only once.

  4. 04 / OBSERVE

    Keep the whole chain of evidence.

    Record separate audio channels, tool arguments, business results and state changes. Count overlap and unanswered chunks separately; distinguish first sound from the substantive answer.

  5. 05 / REVIEW

    Evaluate the evaluator.

    Check whether the AI caller actually followed the scenario. Hear ambiguous consent and numbers yourself. Preserve failed runs and publish grading corrections with their reasons.

  6. 06 / RELEASE

    Start with supervised actions.

    For real customer operations, validate a draft in the company’s system and have a human approve it. Expand permissions only after representative tests, monitoring and recovery paths are in place.

Useful building blocks, with clear limits.

Simulation and evaluation are established ideas. This toolkit contributes reusable order scenarios, stateful mock tools, failure injection, replay and linked phone-call evidence. It complements broader model benchmarks such as Speko. It is an early foundation for developer testing, not a claim that one small pilot certifies production readiness.

Inspect the code and try the offline sandbox ↗