# Cost accounting for the 60-call pilot

Added 13 September 2026. All amounts are USD; estimates retain the study's
12 September 2026 pricing table. No new calls or retrospective repricing.

## Sources and reproducibility

- [Original public study data](https://patrick25076.github.io/grai-voice-bench/data.json)
- Source SHA-256: `3c1ea61dc8d0531568f7c138c932dc6cbf38094fd73e4f8abd28418869c2e6d7`
- Frozen timing analysis SHA-256: `e217e0524ee278090b433d036e5d31fae4011c1edd2141a4df187d899cd9b283`
- [Cost evidence and coverage](https://grailabs.ai/benchmarks/gptvsgemini/cost-data.json)
- [Per-call CSV](https://grailabs.ai/benchmarks/gptvsgemini/costs.csv)
- [Original cost calculation code](https://github.com/patrick25076/grai-voice-bench/blob/study60-audit-v1/costs.py)

The site exporter (`apps/web/scripts/export-costs.py`) verifies the public source's hash,
matches each trial to the frozen analysis by sample, provider, language and audio hash,
and copies cost components without estimating missing modalities. Recording lengths come
from the frozen analysis. The download contains all 60 calls, including unsuccessful calls.
Seven calibration calls are excluded. Browser filters select 30 calls per provider overall,
25 English or 5 Romanian. The cost filter is independent of the results filter.

## Definitions

For the selected calls and each cost component:

- **Per call:** sum of captured costs / number of calls with that component captured.
- **Per recorded minute:** sum of captured costs / sum of full recording durations in minutes
  for those same calls. This is a ratio of sums, not an average of per-call rates.
- **Known study subtotal:** answering model + AI caller + carrier/recording + both post-call
  transcription passes. It is not an all-in deployment price.

Recording time includes silence, capped recordings and unfinished calls. Each two-channel
recording counts once. It is neither active speech duration nor model-billed session duration.
The renderer excludes missing cost rows from that component's denominator and reports captured
counts; it never converts missing amounts to zero. Invalid duration makes the per-minute rate
unavailable. All 60 released rows have a captured partial amount and a valid recording duration.

## Captured answering-model costs

| Configuration                  | Calls | Recording minutes | Known model total | Mean per call | Per recorded minute |
| ------------------------------ | ----: | ----------------: | ----------------: | ------------: | ------------------: |
| Gemini 3.1 Flash Live, partial |    30 |            52.982 |        $1.4749365 |   $0.04916455 |         $0.02783845 |
| GPT-Live 1 + gpt-5.6-terra     |    30 |         69.082333 |       $4.00446967 |   $0.13348232 |         $0.05796662 |

Every Gemini answering-model snapshot contains input/output tokens with unclassified modality.
Their price is unresolved and excluded, not zero. GPT's captured categories are complete, but
final usage and invoices remain unreconciled for both providers. Caller-model estimates can also
have missing categories; the JSON retains those coverage flags and unresolved items.

The GPT total includes $3.45166667 of voice-session estimates and $0.552803 of Terra backend
estimates. The latter uses the original $2 input, $0.20 cached-input, $2.50 cache-write and
$12 output rates per million tokens. Do not attribute the combined cost to the voice fee alone.

The original known study subtotal is $16.38389917 across both arms. AI callers and two offline
ASR passes are evaluation overhead. Carrier charges describe the benchmark call path. LiveKit
SIP/media, hosting, storage, tax, FX and uncaptured provider usage are unreconciled. No cost-per-
successful-task winner or exact savings percentage is claimed from these incomplete totals.

## Published rates, checked 13 September 2026

[OpenAI](https://developers.openai.com/api/docs/pricing#live-session-duration) lists GPT-Live 1
at $0.05 per session minute, billed per second, plus backend model usage and applicable tool
charges. A session minute is not necessarily a recording minute.

[Google](https://ai.google.dev/gemini-api/docs/pricing#gemini-3.1-flash-live-preview) lists
Gemini 3.1 Flash Live audio input at $3 per million tokens ($0.005 per minute of input audio)
and audio output at $12 per million tokens ($0.018 per minute of output audio). Text input/output
is $0.75/$4.50 per million tokens. These are separate streams and billing is token-based;
$0.005 + $0.018 is not a flat rate for every elapsed minute of a phone call.

Published rates provide context. They do not fill gaps in the study snapshots or supersede the
original pricing table used by the estimates. The figures are not invoice-reconciled prices.
