Audio token estimates

Audio Token Calculator

Convert audio duration into tokens and transcription cost. Enter a length or drop in an audio file — the file never leaves your browser.

Audio length

Common lengths

Duration

10m 0s

600 seconds

Cheapest transcription

$0.0058

on Gemini 3.5 Flash-Lite (audio input)

Most audio tokens

24,000

on gpt-4o-transcribe

ModelBillingAudio tokensEst. cost

gpt-transcribe

OpenAI

$0.0045/min—$0.045

gpt-live-transcribe

OpenAI

$0.017/min—$0.170

gpt-4o-transcribeDeprecated

OpenAI

$0.006/min24,000$0.060

gpt-4o-mini-transcribeDeprecated

OpenAI

$0.003/min24,000$0.030

whisper-1Deprecated

OpenAI

$0.006/min—$0.060

gpt-4o-transcribe-diarizeDeprecated

OpenAI

$0.006/min24,000$0.060

Gemini 3.8 Live (audio input)

Google

$3/1M audio tokens19,200$0.058

Gemini 3.5 Flash-Lite (audio input)

Google

$0.3/1M audio tokens19,200$0.0058

Estimates use published list prices (verified Sep 29, 2026) and Google's documented audio tokenization rate (32 tokens/sec). OpenAI's older transcription models are deprecated and retire Feb 26, 2027. Actual bills depend on your plan, region, and provider rounding.

Audio tokens FAQ

How many tokens is 1 minute of audio?

About 1,920 tokens on Gemini's audio tokenization (32 tokens per second, per Google's audio documentation). OpenAI's transcription models are billed by audio minute instead of tokens — $0.0045 per minute for the current gpt-transcribe model.

Why is audio transcription billed per minute instead of tokens?

Transcription models like whisper-1 and gpt-4o-transcribe charge for audio duration because the work scales with recording length, not with how many words are spoken. Gemini's native audio models bill per audio token instead.

How much does it cost to transcribe 1 hour of audio?

Around $0.03 with Gemini 3.5 Flash-Lite audio input, $0.27 with gpt-transcribe, $0.18 with gpt-4o-mini-transcribe, or $0.36 with whisper-1 and gpt-4o-transcribe at standard rates. Note that whisper-1 and the gpt-4o transcription models are deprecated and shut down on Feb 26, 2027.

Does this calculator upload my audio?

No. If you pick a file, the browser only reads its duration metadata. The audio never leaves your device.

Which OpenAI transcription model should I use?

Use gpt-transcribe at $0.0045 per audio minute — it is OpenAI's current recommended model for recorded speech. whisper-1 and the gpt-4o transcription models (including the diarization variant) were deprecated on Aug 26, 2026 and will be removed from the API on Feb 26, 2027.

How many audio tokens per second does Gemini use?

32 audio tokens per second for input audio, according to Google's audio documentation (updated Sep 2026). That works out to 1,920 tokens per minute.