Transcription cost per minute: a dated price reference
Transcription cost per minute at nine vendors: developer APIs from $0.0025, consumer plans from $0.005, human work from $1.99. Prices dated 15 September 2026.
Transcription cost per minute runs from $0.0025 to about $0.17 when a machine does the work, or €0.20 for minutes bought outside a plan, and $1.75 to $1.99 when a person does. Developer APIs publish a real per-minute rate. Consumer services publish a monthly fee, so you divide by the included minutes to get one. A model running in your own browser has no per-minute price at all. Below is the published figure for every vendor whose pricing page we could load on 15 September 2026. Prices move often, so read this as dated rather than permanent.
The table, dated 15 September 2026
| Vendor and tier | Published price | Per minute | What that covers |
|---|---|---|---|
| AssemblyAI Universal-2, async | $0.15/hr | $0.0025 | API access, billed per audio hour submitted |
| OpenAI gpt-4o-mini-transcribe | “$0.003 / minute” estimate | $0.003 | API access, actually billed on tokens |
| AssemblyAI Universal-3.5 Pro, async | $0.21/hr | $0.0035 | API access, billed per audio hour submitted |
| Deepgram Nova-3 monolingual, pre-recorded | $0.0043/min | $0.0043 | API access, pay-as-you-go tier |
| Deepgram Whisper Large, pre-recorded | $0.0048/min | $0.0048 | API access, same price on both tiers |
| Rev Essentials, annual | $25.49 per seat/month, 5,000 min | $0.0051 | App, editor, captions, one seat |
| OpenAI Whisper API | $0.006 / minute | $0.006 | API access |
| AWS Transcribe batch, US East (N. Virginia) | $0.006 per minute | $0.006 | API access, one-second increments |
| Otter Pro, annual | $8.33 per user/month, 1,200 min | $0.0069 | App, live meeting capture, 10 file imports a month |
| AWS Transcribe streaming, US East (N. Virginia) | $0.01 per minute | $0.01 | API access, live audio |
| Otter Pro, monthly | $16.99 per user/month, 1,200 min | $0.0142 | Same plan, billed monthly |
| Descript Hobbyist, annual | $16 per person/month, 10 media hours | $0.0267 | Full video and audio editor |
| Happy Scribe Pro, annual | €19/month, 600 min | €0.0317 | App, subtitle editor, translation |
| Happy Scribe Basic, annual | €8.50/month, 120 min | €0.0708 | Same app, smaller allowance |
| Sonix Core | $25/month, 5 hrs | $0.0833 | App, one seat, transcription and translation |
| Sonix pay-as-you-go | $10/hr | $0.1667 | Same app, no subscription |
| Happy Scribe extra credits | €0.20/min | €0.20 | Minutes beyond a plan allowance |
| Happy Scribe human proofreading | “From €1.75/min” | €1.75 | A person checks the machine output |
| Rev human transcription | $1.99 /min. | $1.99 | A person transcribes the audio |
| A model in your own browser | none | $0 | One 200 MB download, then electricity |
Two currencies appear because Happy Scribe publishes in euros and the rest in dollars. No conversion has been applied.
How to read a transcription cost per minute
A per-minute price sounds like the simplest unit in software, and it mostly is, but three things underneath it change the arithmetic.
The first is what gets counted. AssemblyAI bills pre-recorded work “per audio hour submitted” and charges a multichannel file per channel, so a two-track interview recorded on separate microphones is two hours of billing for one hour of conversation. AssemblyAI also bills streaming on WebSocket session duration, open to close rather than audio sent, and says idle time counts, so a connection left open through a coffee break still bills.
The second is rounding. AWS documents that Amazon Transcribe usage is “billed in one-second increments, with no minimum applied” for standard batch and streaming. That is unusually explicit. Its generative call summarisation add-on is the exception, billed in one-second increments “with a minimum per request charge of 15 seconds”. Sonix says its pay-as-you-go work is prorated to the nearest second. Most of the others publish no rounding rule at all, so a very short clip may cost more than the raw arithmetic suggests.
The third is that some per-minute prices are not the billing unit. OpenAI bills gpt-4o-transcribe at $2.50 per million input tokens and $10.00 per million output tokens, and prints $0.006 a minute next to it as an estimate. The estimate assumes ordinary speech. Dense, fast talking produces more tokens per minute of audio than a slow monologue does, so the bill moves with what is being said.
Deepgram adds a fourth wrinkle worth knowing about. Several of its streaming lines show a “Current price” alongside a “Regular price”, for example Nova-3 monolingual streaming at “Current price $0.0048/min Regular price $0.0077/min”. Promotional pricing is a normal thing to publish, and it is also a good reason to check the page yourself rather than trust a table like this one six months from now.
The developer API tier is where the floor sits
Four of the vendors here sell speech recognition as an API and nothing else. AssemblyAI lists the lowest published rate on this page, $0.15 an hour for Universal-2 and $0.21 an hour for Universal-3.5 Pro, which is $0.0025 and $0.0035 a minute. Deepgram’s pre-recorded Nova-3 monolingual is $0.0043 a minute on pay-as-you-go and $0.0036 on its Growth tier. OpenAI charges $0.006 a minute for Whisper and estimates $0.003 for gpt-4o-mini-transcribe. Amazon Transcribe is $0.006 a minute for batch and $0.01 for streaming in US East (N. Virginia).
That is a narrow band. Every published async rate in this group falls between a quarter of a cent and six tenths of a cent a minute. Where a vendor prices the same model both ways, streaming usually costs more, because it is holding a connection open for you: Deepgram, AWS and AssemblyAI’s Universal-3.5 Pro all price it above their batch rate.
These prices come with no interface. You write code, you move the audio files yourself, and the audio goes to someone else’s servers. Free credits soften the start: AssemblyAI gives “$50 in free credits” at signup, and AWS gives new accounts 60 minutes a month for twelve months.
Consumer plans hide their per-minute rate inside a monthly fee
None of the subscription services publish a per-minute price for their included minutes. You get one by dividing the monthly fee by the allowance, which is the arithmetic in the third column of the table.
Rev Essentials is $25.49 per seat per month on the annual plan with 5,000 AI transcription and caption minutes per user, so $0.0051 a minute. Otter Pro is $8.33 per user per month annually, or $16.99 monthly, for 1,200 minutes: $0.0069 and $0.0142. Descript Hobbyist is $16 a person a month on the annual plan, $24 month to month, for 10 media hours, so $0.0267 a minute on the annual rate. Happy Scribe Pro is €19 a month on the annual plan for 600 minutes, €0.0317 a minute. Sonix Core is $25 a month for 5 hours, $0.0833 a minute, and its pay-as-you-go rate of $10 an hour is $0.1667.
Every one of those numbers assumes you use the entire allowance. Use half of it and the rate doubles. That is the main reason a subscription and an API can look a hundred times apart on a price list and ten times apart on your actual invoice.
Allowances also carry rules that minutes alone do not describe. Otter’s Basic plan is 300 minutes a month but caps a conversation at 30 minutes and allows three file imports for the life of the account, and Pro allows ten imports a month against its 1,200 minutes. If your work is files you already have rather than meetings Otter sits in, the imports run out long before the minutes do. There is more on caps like that in why free transcription sites stop at 30 minutes.
Human transcription is priced two orders of magnitude higher
Rev publishes human transcription at $1.99 a minute and human captions starting at the same figure, with subscriber discounts of 10% on the annual Essentials plan, 3% on monthly Essentials, 15% on annual Pro and 5% on monthly Pro. Global subtitles are priced by language, from $6.49 a minute for Spanish and Hindi up to $15.99 for Japanese and Korean.
Happy Scribe publishes human proofreading from €1.75 a minute, falling to €1.66 on its Business plan.
Round numbers make the gap plain. An hour of audio is about $119 at Rev’s human rate and about 15 cents through AssemblyAI’s cheapest async API. Those are different products bought for different reasons, and no table can flatten that into one ranking.
Local processing has no per-minute price to publish
A speech model running on your own machine has no meter. The whole per-minute frame stops applying, because there is no server issuing invoices and no allowance to divide.
Two costs remain. One is a download: OpenAI’s open-source Whisper base English model is about 200 MB, fetched once and then cached by the browser. The other is electricity. On the desktop we test on, an hour of audio takes roughly 40 minutes of compute at about 60 W of extra draw, which is 0.04 kWh, or around a cent at European domestic rates.
The trade-offs are real and belong next to that zero. It needs desktop Chrome or Edge with a working WebGPU adapter, so Firefox, Safari and phones are out for now. It is English only. It uses the base model, which is the small end of the size range, and names, technical vocabulary, heavy accents and noisy rooms are where it slips, so read the output before you rely on it. It runs at about 1.5 times real time on a desktop with a graphics card in our test, and roughly real time on a thin laptop, with the tab open the whole while.
If that fits the work you have, FreeTranscribe does it in the browser with no account, no upload and no minute cap. For a per-hour view of the three big consumer apps, our cost per hour comparison works the same prices from the other end.
Frequently asked questions
Why do the API prices look so much lower than the app prices? You are buying less. The API returns a transcript in response to a file you send with code you wrote. The apps wrap an editor, storage, search, speaker labels, exports and support around the same recognition step, and the distance between $0.003 and $0.08 a minute is what that wrapper costs.
Does a lower per-minute rate mean a lower bill? Not on its own. Look at what gets counted first. Per-channel billing on a multi-microphone recording, session-duration billing on a streaming connection, and unused subscription minutes all move the real number further than a tenth of a cent of list price does.
Are annual and monthly prices really this far apart? At Otter, Rev, Descript, Happy Scribe and Sonix the annual commitment is cheaper per month than paying month to month, in some cases by close to half. The table uses whichever figure each vendor publishes for that plan and says which one it is.
How long will these figures be right? Not long. Deepgram is showing promotional prices next to regular ones on several lines today, and free allowances at these services have changed repeatedly in recent years. Check the vendor page before you budget against any row here.
Sources, checked 15 September 2026
- Deepgram pricing, per-minute rates for streaming and pre-recorded models: https://deepgram.com/pricing
- AssemblyAI pricing, per-hour async and streaming rates and billing terms: https://www.assemblyai.com/pricing
- OpenAI API pricing, transcription models and per-minute estimates: https://developers.openai.com/api/docs/pricing
- Amazon Transcribe pricing, per-minute rates and one-second billing increments: https://aws.amazon.com/transcribe/pricing/
- Rev pricing, human per-minute rates and subscription plans: https://www.rev.com/pricing
- Otter.ai pricing, plan prices, minute allowances and import limits: https://otter.ai/pricing
- Descript pricing, plan prices and media hours: https://www.descript.com/pricing
- Happy Scribe pricing, plan prices in euros, credit rate and human proofreading: https://www.happyscribe.com/pricing
- Sonix pricing, subscription and pay-as-you-go hourly rates: https://sonix.ai/pricing
Could not be checked on the day: Google Cloud Speech-to-Text pricing and Azure Speech service pricing, whose pages returned no usable figures to an automated fetch, and Speechmatics, whose page showed a number without a unit next to it. They are left out rather than guessed at.
If you spot a row that has moved, tell us and we will date a new version of the table.