Audio token estimates
Convert audio duration into tokens and transcription cost. Enter a length or drop in an audio file — the file never leaves your browser.
Common lengths
Duration
10m 0s
600 seconds
Cheapest transcription
$0.0058
on Gemini 3.5 Flash-Lite (audio input)
Most audio tokens
24,000
on gpt-4o-transcribe
| Model | Billing | Audio tokens | Est. cost |
|---|---|---|---|
gpt-transcribe OpenAI | $0.0045/min | — | $0.045 |
gpt-live-transcribe OpenAI | $0.017/min | — | $0.170 |
gpt-4o-transcribeDeprecated OpenAI | $0.006/min | 24,000 | $0.060 |
gpt-4o-mini-transcribeDeprecated OpenAI | $0.003/min | 24,000 | $0.030 |
whisper-1Deprecated OpenAI | $0.006/min | — | $0.060 |
gpt-4o-transcribe-diarizeDeprecated OpenAI | $0.006/min | 24,000 | $0.060 |
Gemini 3.8 Live (audio input) | $3/1M audio tokens | 19,200 | $0.058 |
Gemini 3.5 Flash-Lite (audio input) | $0.3/1M audio tokens | 19,200 | $0.0058 |
Estimates use published list prices (verified Sep 29, 2026) and Google's documented audio tokenization rate (32 tokens/sec). OpenAI's older transcription models are deprecated and retire Feb 26, 2027. Actual bills depend on your plan, region, and provider rounding.
About 1,920 tokens on Gemini's audio tokenization (32 tokens per second, per Google's audio documentation). OpenAI's transcription models are billed by audio minute instead of tokens — $0.0045 per minute for the current gpt-transcribe model.
Transcription models like whisper-1 and gpt-4o-transcribe charge for audio duration because the work scales with recording length, not with how many words are spoken. Gemini's native audio models bill per audio token instead.
Around $0.03 with Gemini 3.5 Flash-Lite audio input, $0.27 with gpt-transcribe, $0.18 with gpt-4o-mini-transcribe, or $0.36 with whisper-1 and gpt-4o-transcribe at standard rates. Note that whisper-1 and the gpt-4o transcription models are deprecated and shut down on Feb 26, 2027.
No. If you pick a file, the browser only reads its duration metadata. The audio never leaves your device.
Use gpt-transcribe at $0.0045 per audio minute — it is OpenAI's current recommended model for recorded speech. whisper-1 and the gpt-4o transcription models (including the diarization variant) were deprecated on Aug 26, 2026 and will be removed from the API on Feb 26, 2027.
32 audio tokens per second for input audio, according to Google's audio documentation (updated Sep 2026). That works out to 1,920 tokens per minute.