Comparisons

Transcription cost per minute: a dated price reference

Transcription cost per minute at nine vendors: developer APIs from $0.0025, consumer plans from $0.005, human work from $1.99. Prices dated 15 September 2026.

Transcription cost per minute runs from $0.0025 to about $0.17 when a machine does the work, or €0.20 for minutes bought outside a plan, and $1.75 to $1.99 when a person does. Developer APIs publish a real per-minute rate. Consumer services publish a monthly fee, so you divide by the included minutes to get one. A model running in your own browser has no per-minute price at all. Below is the published figure for every vendor whose pricing page we could load on 15 September 2026. Prices move often, so read this as dated rather than permanent.

The table, dated 15 September 2026

Vendor and tier Published price Per minute What that covers
AssemblyAI Universal-2, async $0.15/hr $0.0025 API access, billed per audio hour submitted
OpenAI gpt-4o-mini-transcribe “$0.003 / minute” estimate $0.003 API access, actually billed on tokens
AssemblyAI Universal-3.5 Pro, async $0.21/hr $0.0035 API access, billed per audio hour submitted
Deepgram Nova-3 monolingual, pre-recorded $0.0043/min $0.0043 API access, pay-as-you-go tier
Deepgram Whisper Large, pre-recorded $0.0048/min $0.0048 API access, same price on both tiers
Rev Essentials, annual $25.49 per seat/month, 5,000 min $0.0051 App, editor, captions, one seat
OpenAI Whisper API $0.006 / minute $0.006 API access
AWS Transcribe batch, US East (N. Virginia) $0.006 per minute $0.006 API access, one-second increments
Otter Pro, annual $8.33 per user/month, 1,200 min $0.0069 App, live meeting capture, 10 file imports a month
AWS Transcribe streaming, US East (N. Virginia) $0.01 per minute $0.01 API access, live audio
Otter Pro, monthly $16.99 per user/month, 1,200 min $0.0142 Same plan, billed monthly
Descript Hobbyist, annual $16 per person/month, 10 media hours $0.0267 Full video and audio editor
Happy Scribe Pro, annual €19/month, 600 min €0.0317 App, subtitle editor, translation
Happy Scribe Basic, annual €8.50/month, 120 min €0.0708 Same app, smaller allowance
Sonix Core $25/month, 5 hrs $0.0833 App, one seat, transcription and translation
Sonix pay-as-you-go $10/hr $0.1667 Same app, no subscription
Happy Scribe extra credits €0.20/min €0.20 Minutes beyond a plan allowance
Happy Scribe human proofreading “From €1.75/min” €1.75 A person checks the machine output
Rev human transcription $1.99 /min. $1.99 A person transcribes the audio
A model in your own browser none $0 One 200 MB download, then electricity

Two currencies appear because Happy Scribe publishes in euros and the rest in dollars. No conversion has been applied.

How to read a transcription cost per minute

A per-minute price sounds like the simplest unit in software, and it mostly is, but three things underneath it change the arithmetic.

The first is what gets counted. AssemblyAI bills pre-recorded work “per audio hour submitted” and charges a multichannel file per channel, so a two-track interview recorded on separate microphones is two hours of billing for one hour of conversation. AssemblyAI also bills streaming on WebSocket session duration, open to close rather than audio sent, and says idle time counts, so a connection left open through a coffee break still bills.

The second is rounding. AWS documents that Amazon Transcribe usage is “billed in one-second increments, with no minimum applied” for standard batch and streaming. That is unusually explicit. Its generative call summarisation add-on is the exception, billed in one-second increments “with a minimum per request charge of 15 seconds”. Sonix says its pay-as-you-go work is prorated to the nearest second. Most of the others publish no rounding rule at all, so a very short clip may cost more than the raw arithmetic suggests.

The third is that some per-minute prices are not the billing unit. OpenAI bills gpt-4o-transcribe at $2.50 per million input tokens and $10.00 per million output tokens, and prints $0.006 a minute next to it as an estimate. The estimate assumes ordinary speech. Dense, fast talking produces more tokens per minute of audio than a slow monologue does, so the bill moves with what is being said.

Deepgram adds a fourth wrinkle worth knowing about. Several of its streaming lines show a “Current price” alongside a “Regular price”, for example Nova-3 monolingual streaming at “Current price $0.0048/min Regular price $0.0077/min”. Promotional pricing is a normal thing to publish, and it is also a good reason to check the page yourself rather than trust a table like this one six months from now.

The developer API tier is where the floor sits

Four of the vendors here sell speech recognition as an API and nothing else. AssemblyAI lists the lowest published rate on this page, $0.15 an hour for Universal-2 and $0.21 an hour for Universal-3.5 Pro, which is $0.0025 and $0.0035 a minute. Deepgram’s pre-recorded Nova-3 monolingual is $0.0043 a minute on pay-as-you-go and $0.0036 on its Growth tier. OpenAI charges $0.006 a minute for Whisper and estimates $0.003 for gpt-4o-mini-transcribe. Amazon Transcribe is $0.006 a minute for batch and $0.01 for streaming in US East (N. Virginia).

That is a narrow band. Every published async rate in this group falls between a quarter of a cent and six tenths of a cent a minute. Where a vendor prices the same model both ways, streaming usually costs more, because it is holding a connection open for you: Deepgram, AWS and AssemblyAI’s Universal-3.5 Pro all price it above their batch rate.

These prices come with no interface. You write code, you move the audio files yourself, and the audio goes to someone else’s servers. Free credits soften the start: AssemblyAI gives “$50 in free credits” at signup, and AWS gives new accounts 60 minutes a month for twelve months.

Consumer plans hide their per-minute rate inside a monthly fee

None of the subscription services publish a per-minute price for their included minutes. You get one by dividing the monthly fee by the allowance, which is the arithmetic in the third column of the table.

Rev Essentials is $25.49 per seat per month on the annual plan with 5,000 AI transcription and caption minutes per user, so $0.0051 a minute. Otter Pro is $8.33 per user per month annually, or $16.99 monthly, for 1,200 minutes: $0.0069 and $0.0142. Descript Hobbyist is $16 a person a month on the annual plan, $24 month to month, for 10 media hours, so $0.0267 a minute on the annual rate. Happy Scribe Pro is €19 a month on the annual plan for 600 minutes, €0.0317 a minute. Sonix Core is $25 a month for 5 hours, $0.0833 a minute, and its pay-as-you-go rate of $10 an hour is $0.1667.

Every one of those numbers assumes you use the entire allowance. Use half of it and the rate doubles. That is the main reason a subscription and an API can look a hundred times apart on a price list and ten times apart on your actual invoice.

Allowances also carry rules that minutes alone do not describe. Otter’s Basic plan is 300 minutes a month but caps a conversation at 30 minutes and allows three file imports for the life of the account, and Pro allows ten imports a month against its 1,200 minutes. If your work is files you already have rather than meetings Otter sits in, the imports run out long before the minutes do. There is more on caps like that in why free transcription sites stop at 30 minutes.

Human transcription is priced two orders of magnitude higher

Rev publishes human transcription at $1.99 a minute and human captions starting at the same figure, with subscriber discounts of 10% on the annual Essentials plan, 3% on monthly Essentials, 15% on annual Pro and 5% on monthly Pro. Global subtitles are priced by language, from $6.49 a minute for Spanish and Hindi up to $15.99 for Japanese and Korean.

Happy Scribe publishes human proofreading from €1.75 a minute, falling to €1.66 on its Business plan.

Round numbers make the gap plain. An hour of audio is about $119 at Rev’s human rate and about 15 cents through AssemblyAI’s cheapest async API. Those are different products bought for different reasons, and no table can flatten that into one ranking.

Local processing has no per-minute price to publish

A speech model running on your own machine has no meter. The whole per-minute frame stops applying, because there is no server issuing invoices and no allowance to divide.

Two costs remain. One is a download: OpenAI’s open-source Whisper base English model is about 200 MB, fetched once and then cached by the browser. The other is electricity. On the desktop we test on, an hour of audio takes roughly 40 minutes of compute at about 60 W of extra draw, which is 0.04 kWh, or around a cent at European domestic rates.

The trade-offs are real and belong next to that zero. It needs desktop Chrome or Edge with a working WebGPU adapter, so Firefox, Safari and phones are out for now. It is English only. It uses the base model, which is the small end of the size range, and names, technical vocabulary, heavy accents and noisy rooms are where it slips, so read the output before you rely on it. It runs at about 1.5 times real time on a desktop with a graphics card in our test, and roughly real time on a thin laptop, with the tab open the whole while.

If that fits the work you have, FreeTranscribe does it in the browser with no account, no upload and no minute cap. For a per-hour view of the three big consumer apps, our cost per hour comparison works the same prices from the other end.

Frequently asked questions

Why do the API prices look so much lower than the app prices? You are buying less. The API returns a transcript in response to a file you send with code you wrote. The apps wrap an editor, storage, search, speaker labels, exports and support around the same recognition step, and the distance between $0.003 and $0.08 a minute is what that wrapper costs.

Does a lower per-minute rate mean a lower bill? Not on its own. Look at what gets counted first. Per-channel billing on a multi-microphone recording, session-duration billing on a streaming connection, and unused subscription minutes all move the real number further than a tenth of a cent of list price does.

Are annual and monthly prices really this far apart? At Otter, Rev, Descript, Happy Scribe and Sonix the annual commitment is cheaper per month than paying month to month, in some cases by close to half. The table uses whichever figure each vendor publishes for that plan and says which one it is.

How long will these figures be right? Not long. Deepgram is showing promotional prices next to regular ones on several lines today, and free allowances at these services have changed repeatedly in recent years. Check the vendor page before you budget against any row here.

Sources, checked 15 September 2026

Could not be checked on the day: Google Cloud Speech-to-Text pricing and Azure Speech service pricing, whose pages returned no usable figures to an automated fetch, and Speechmatics, whose page showed a number without a unit next to it. They are left out rather than guessed at.

If you spot a row that has moved, tell us and we will date a new version of the table.

pricingper minuteapissubscriptionslocal transcription
FreeTranscribe

Written by the people who build FreeTranscribe. We test every claim on our own files and date every price. About the site.

Transcribe a file now. Free, in your browser.
Open the transcriber