Audio token estimates
Convert audio duration into tokens and transcription cost. Enter a length or drop in an audio file — the file never leaves your browser.
Common lengths
Duration
10m 0s
600 seconds
Cheapest transcription
$0.0045
on Gemini 3.5 Flash-Lite (audio input)
Most audio tokens
24,000
on gpt-4o-transcribe
| Model | Billing | Audio tokens | Est. cost |
|---|---|---|---|
gpt-4o-transcribe OpenAI | $0.006/min | 24,000 | $0.060 |
gpt-4o-mini-transcribe OpenAI | $0.003/min | 24,000 | $0.030 |
whisper-1 OpenAI | $0.006/min | — | $0.060 |
gpt-4o-transcribe-diarize OpenAI | $0.006/min | 24,000 | $0.060 |
Gemini 3.8 Live (audio input) | $3/1M audio tokens | 15,000 | $0.045 |
Gemini 3.5 Flash-Lite (audio input) | $0.3/1M audio tokens | 15,000 | $0.0045 |
Estimates use published list prices and standard audio tokenization rates (25 tokens/sec for Gemini). Actual bills depend on your plan, region, and provider rounding.
About 1,500 tokens on Gemini's audio tokenization (25 tokens per second). OpenAI's transcription models are billed by audio minute instead of tokens, at roughly $0.006 per minute.
Transcription models like whisper-1 and gpt-4o-transcribe charge for audio duration because the work scales with recording length, not with how many words are spoken. Gemini's native audio models bill per audio token instead.
Around $0.03 with Gemini 3.5 Flash-Lite audio input, $0.18 with gpt-4o-mini-transcribe, or $0.36 with whisper-1 and gpt-4o-transcribe at standard rates.
No. If you pick a file, the browser only reads its duration metadata. The audio never leaves your device.
Both cost $0.006 per audio minute. gpt-4o-transcribe is the newer model with better accuracy and streaming support; whisper-1 is the classic baseline. gpt-4o-mini-transcribe halves the price at $0.003 per minute.
25 audio tokens per second for input audio, according to Google's pricing documentation. That works out to 1,500 tokens per minute.