A
AssemblyAI
Guest
On August 17, 2026, Wispr raised $280 million at a $2 billion valuation. Two days later, The New York Times Magazine ran a review of its dictation app under the headline "Everyone's Using This A.I. Dictation App That I Want to Murder With a Hammer."
Buried in the funding coverage was the more interesting detail: Wispr acknowledged error rates above 30% in hard conditions — noise, accents, music — and previewed a new model to bring that down. So the reviewer's frustration and the company's own diagnosis landed in the same week.
Here's the thing. Most of what users experience as "this dictation app is bad" traces back to an architecture decision someone made months earlier, before a single word was transcribed. Dictation looks like a solved problem right up until you build one, and then you discover that the obvious choices are mostly wrong.
This post is about that decision. Four architectures, what each one actually costs, and where each one breaks.
It's free, it ships with the browser, and you can have a demo working in twenty minutes. For a prototype, use it. For a product, the constraints compound fast.
The honest summary: the Web Speech API is a great way to find out whether your users want dictation, and a bad way to serve them once they do.
This is the mistake I see most often, and it's an understandable one. Dictation is real-time-ish, streaming APIs are for real-time things, so developers reach for a WebSocket.
But look at what streaming gives you: partial hypotheses that update as the speaker talks, and end-of-turn detection to figure out when they've stopped. Both are genuinely valuable — for live captions, for a voice agent that has to start thinking before you finish your sentence. Universal-3.5 Pro Realtime exists precisely for those workloads.
Dictation needs neither. Your user is holding down a key. The keypress is the end-of-turn signal — you have a perfect, zero-latency answer to the question streaming spends real engineering effort estimating. And nobody wants to watch text flicker and rewrite itself inside the Slack message they're composing. They want the finished sentence.
So you'd be paying for two features and using zero of them. Literally paying: streaming is billed on session duration — how long the WebSocket is open — not on how much audio you sent. A dictation app that holds a socket open while the user stares at the ceiling composing their thought is billing the ceiling-staring.
Keep streaming for the workloads that react mid-utterance. For push-to-talk, it's the expensive way to build the cheap thing.
Pre-recorded transcription fixes the billing problem — you pay for audio duration, and Universal-3.5 Pro runs $0.21/hr. It's the right call for meetings, podcasts, and call recordings.
The mismatch is the interaction model. Async APIs are job-submission systems: upload the file, get an ID, poll or wait on a webhook. That design is correct when the audio is forty minutes long and the user isn't waiting. It's the wrong shape when the audio is eleven seconds long and someone is staring at a blinking cursor.
You can make it work. You just end up building a request-response API out of a job queue, which is a strange thing to do on purpose.
For push-to-talk dictation, use a synchronous request. One utterance in, one response out, no connection to manage and no job to poll.
That's what the Dictation API is — a single POST, released September 15, 2026:
A few constraints worth knowing before you build against it. Audio is capped at 120 seconds per call, which is the right bound for an utterance and the wrong one for a lecture. It accepts WAV or raw 16-bit PCM only — the endpoint decodes audio as it arrives, and compressed formats can't be decoded incrementally, so MP3 and M4A come back as a 415. And the config part has to arrive before the first audio byte, because the server starts transcribing while the rest of the body is still in flight.
That last constraint is also the feature, which brings us to latency.
Four things, in rough order of how much they buy you.
Upload while the user is still talking. Because the server transcribes what it has as the rest arrives, a client that streams audio during the recording isn't waiting on the upload afterward. What the user waits for after they stop speaking is the last fragment of audio, not the whole clip. Raw PCM is easiest here — there's no container header to finalize, so frames go on the wire as they come off the microphone:
Warm the connection at key-down. DNS, TCP, and TLS in front of a two-second utterance is a meaningful share of the total wait. There's a
Call it the moment the user reaches for the record button. Two gotchas: it only helps if the transcription request goes through the same HTTP client and the same host, and idle connections get evicted after a few seconds — so warm shortly before the request, not at application startup. (We wrote up the full pre-warming technique in a step-by-step build guide if you want the details.)
Getting your free API key takes about a minute if you want to time this yourself: assemblyai.com/dashboard/signup.
It returns two things instead of one.
A transcription API gives you the words that were spoken. A dictation API gives you the words that were spoken and the text your user actually meant to write:
That gap is where most dictation products go wrong, because there are really three separate jobs stacked on top of each other and they fail in completely different ways:
Layer 3 is not solved, and I'd argue it shouldn't live in the dictation layer at all. But layers 1 and 2 get conflated constantly — including by reviewers, who experience all three failing at once and file it as "the transcription is bad."
Three separate knobs control this, and mixing them up is a common source of frustration:
The first two steer what the model hears. The third reshapes what it already heard. No amount of
One detail worth noting: the rewrite is best-effort. If it fails, you still get a 200 with the transcription intact and
This one deserves more attention than it gets.
If you're piping a raw transcript into an LLM cleanup prompt — and if you rolled your own layer 2, you are — you've built an injection vector into your own product. Not hypothetically. People say "ignore what I just said" out loud while dictating, constantly. They say "translate this into French." They say "actually, scratch that, write it as a list." Those are ordinary speech, and they're also instructions.
A naive cleanup pipeline will cheerfully execute them, and your user will watch their email get translated into French because they thought out loud.
The Dictation API passes the transcript to the model as fenced data with instructions not to act on anything inside it, so dictated commands get rewritten as speech rather than carried out. If you build the rewrite layer yourself, this is your problem to solve, and it's worth solving before launch rather than after a support ticket.
Dictation pricing splits into two very different models, and the comparison is less close than you'd think.
Consumer dictation apps run subscriptions — Wispr Flow's Pro tier is $15/month, or $12/month billed annually, for unlimited words. The Dictation API is $0.62/hr of audio, all-inclusive, with recognition and cleanup in the same call.
So how much dictation before the subscription wins? About 24 hours of speech a month. At a typical dictation rate of roughly 220 words per minute, that's somewhere north of 14,000 words every working day — not written, spoken into a microphone, every day, all month.
Realistic heavy use looks different. A developer dictating 3,000 words a day across commit messages, Slack, and code comments is speaking for about 14 minutes a day, or five hours a month. That's about $3.10. Moderate use lands closer to a dollar.
The point isn't that subscriptions are a rip-off. It's that the subscription was never priced against the transcription — it's priced against the app, the syncing, the polish, and the fact that someone else maintains it. Which is a perfectly reasonable thing to pay for, right up until you want dictation inside your product, where the per-hour number is the one that matters.
If you'd rather read working code than a spec, Blurt is an open-source macOS dictation app built on this API — MIT licensed, native AppKit and SwiftUI, no Electron. Hold a key, talk, and the text lands in whatever app you're typing in. You bring your own API key, and the whole audio path is a single POST you can read end to end in the repo.
It's new, and it's a reference implementation rather than a polished product. That's rather the point — it exists so you can see exactly how the pieces fit before you commit to an architecture.
The reviewer's complaint was real. Her transcripts needed as much editing as she'd saved, and she reasonably concluded the technology wasn't ready.
But she was describing three different failures and calling them one. Some of it was recognition, which is measurable and has gotten dramatically better. Some of it was the cleanup layer, which most apps were building themselves out of a transcription API and a prompt, which is exactly where quality drifts and why users started noticing regressions. And some of it — the part about speech not being writing, about the app being unable to reorganize a thought — is a real, unsolved problem that no transcription model is going to fix.
That last one is worth sitting with. We've gotten good at hearing people accurately and decent at making their speech read like prose. Turning a spoken ramble into a well-structured argument is a different problem, and I don't think it belongs in the dictation layer at all. It belongs wherever the writing happens.
Which means the interesting work in dictation over the next year probably isn't in the model. It's in figuring out where the editor lives.
Buried in the funding coverage was the more interesting detail: Wispr acknowledged error rates above 30% in hard conditions — noise, accents, music — and previewed a new model to bring that down. So the reviewer's frustration and the company's own diagnosis landed in the same week.
Here's the thing. Most of what users experience as "this dictation app is bad" traces back to an architecture decision someone made months earlier, before a single word was transcribed. Dictation looks like a solved problem right up until you build one, and then you discover that the obvious choices are mostly wrong.
This post is about that decision. Four architectures, what each one actually costs, and where each one breaks.
The four options, side by side
| Browser Web Speech API | Streaming (WebSocket) | Async / pre-recorded | Purpose-built dictation endpoint | |
|---|---|---|---|---|
| Shape | Browser event API | Persistent socket, partial + final hypotheses | Submit job, poll or webhook | One HTTP request, one response |
| Billing | Free | Session duration (socket open time) | Audio duration | Audio duration |
| Latency to final text | Varies by browser | Hundreds of ms after speech ends | Seconds to minutes | Typically under 1s |
| Good for | Prototypes, hobby projects | Live captions, voice agents, anything reacting mid-utterance | Meetings, podcasts, long recordings | Push-to-talk dictation |
| Breaks on | Browser lock-in, no vocabulary control, no server-side record | Paying for silence while the user thinks | Job-submission round trip | Utterances over 120 seconds |
What are the limitations of the browser Web Speech API for production?
It's free, it ships with the browser, and you can have a demo working in twenty minutes. For a prototype, use it. For a product, the constraints compound fast.
- You don't control the model. In Chrome, audio goes to Google's servers and comes back as text. You can't pin a version, so recognition quality changes underneath you with no changelog and no rollback.
- You can't teach it your vocabulary. There's no reliable way to bias recognition toward your users' proper nouns, your API's method names, or the drug names in a clinical workflow. This is the single biggest driver of "the dictation feels dumb" complaints, and the browser API gives you no lever at all.
- Browser support is genuinely uneven. Implementations differ across Chrome, Safari, and Firefox in ways that go beyond a polyfill — different event semantics, different continuous-mode behavior, different failure modes on network loss.
- You never see the audio. The recognition happens outside your application, so you can't log it, re-run it against a better model later, redact it, or debug a specific bad transcript. When a user reports "it got my name wrong," you have nothing to look at.
- It's a browser API. The moment your roadmap includes a desktop app, a mobile app, or a CLI, you're rewriting the whole path anyway.
The honest summary: the Web Speech API is a great way to find out whether your users want dictation, and a bad way to serve them once they do.
Why streaming is usually the wrong default
This is the mistake I see most often, and it's an understandable one. Dictation is real-time-ish, streaming APIs are for real-time things, so developers reach for a WebSocket.
But look at what streaming gives you: partial hypotheses that update as the speaker talks, and end-of-turn detection to figure out when they've stopped. Both are genuinely valuable — for live captions, for a voice agent that has to start thinking before you finish your sentence. Universal-3.5 Pro Realtime exists precisely for those workloads.
Dictation needs neither. Your user is holding down a key. The keypress is the end-of-turn signal — you have a perfect, zero-latency answer to the question streaming spends real engineering effort estimating. And nobody wants to watch text flicker and rewrite itself inside the Slack message they're composing. They want the finished sentence.
So you'd be paying for two features and using zero of them. Literally paying: streaming is billed on session duration — how long the WebSocket is open — not on how much audio you sent. A dictation app that holds a socket open while the user stares at the ceiling composing their thought is billing the ceiling-staring.
Keep streaming for the workloads that react mid-utterance. For push-to-talk, it's the expensive way to build the cheap thing.
Why async isn't quite it either
Pre-recorded transcription fixes the billing problem — you pay for audio duration, and Universal-3.5 Pro runs $0.21/hr. It's the right call for meetings, podcasts, and call recordings.
The mismatch is the interaction model. Async APIs are job-submission systems: upload the file, get an ID, poll or wait on a webhook. That design is correct when the audio is forty minutes long and the user isn't waiting. It's the wrong shape when the audio is eleven seconds long and someone is staring at a blinking cursor.
You can make it work. You just end up building a request-response API out of a job queue, which is a strange thing to do on purpose.
So: sync, streaming, or async?
For push-to-talk dictation, use a synchronous request. One utterance in, one response out, no connection to manage and no job to poll.
That's what the Dictation API is — a single POST, released September 15, 2026:
Code:
curl -X POST https://dictation.assemblyai.com/v1/transcribe/live \
-H 'Authorization: <YOUR_API_KEY>' \
-F 'config={};type=application/json' \
-F '[email protected];type=audio/wav'
A few constraints worth knowing before you build against it. Audio is capped at 120 seconds per call, which is the right bound for an utterance and the wrong one for a lecture. It accepts WAV or raw 16-bit PCM only — the endpoint decodes audio as it arrives, and compressed formats can't be decoded incrementally, so MP3 and M4A come back as a 415. And the config part has to arrive before the first audio byte, because the server starts transcribing while the rest of the body is still in flight.
That last constraint is also the feature, which brings us to latency.
What's the lowest-latency path from keypress to finished text?
Four things, in rough order of how much they buy you.
Upload while the user is still talking. Because the server transcribes what it has as the rest arrives, a client that streams audio during the recording isn't waiting on the upload afterward. What the user waits for after they stop speaking is the last fragment of audio, not the whole clip. Raw PCM is easiest here — there's no container header to finalize, so frames go on the wire as they come off the microphone:
Code:
import assemblyai as aai
aai.settings.api_key = "<YOUR_API_KEY>"
config = aai.DictationConfig(sample_rate=16000, channels=1)
def record():
"""Your microphone loop, yielding 16-bit PCM bytes."""
while recording:
yield stream.read(4096)
result = aai.DictationTranscriber().transcribe_live(record(), config)
print(result.text) # verbatim transcript
print(result.llm_response) # cleaned-up text, ready to send
Warm the connection at key-down. DNS, TCP, and TLS in front of a two-second utterance is a meaningful share of the total wait. There's a
GET /warm endpoint that does nothing except leave an open connection in your HTTP client's pool:
Code:
curl https://dictation.assemblyai.com/warm
# {"warm": "toasty"}
Call it the moment the user reaches for the record button. Two gotchas: it only helps if the transcription request goes through the same HTTP client and the same host, and idle connections get evicted after a few seconds — so warm shortly before the request, not at application startup. (We wrote up the full pre-warming technique in a step-by-step build guide if you want the details.)
- Set your client timeout high, not low. Typical short clips come back in under a second, but set the HTTP timeout to 90 seconds anyway. A tight timeout turns a slow response into a failed one, and a failed dictation is much worse than a slow one.
- Don't retry a chunked body. A streamed upload can't be replayed. If you want retries, keep the audio in memory.
Getting your free API key takes about a minute if you want to time this yourself: assemblyai.com/dashboard/signup.
What makes a dictation API different from a regular speech-to-text API?
It returns two things instead of one.
A transcription API gives you the words that were spoken. A dictation API gives you the words that were spoken and the text your user actually meant to write:
Code:
text: "um so can we uh move the the meeting to thursday
i think friday works better actually"
llm_response: "Can we move the meeting to Friday? That works better."
That gap is where most dictation products go wrong, because there are really three separate jobs stacked on top of each other and they fail in completely different ways:
- Hearing the words correctly. Pure recognition. The Dictation API runs on Universal-3.5 Pro, which scores 3.87% word error rate on short-form audio.
- Turning speech into text a person would have typed. Filler removal, self-correction resolution, punctuation, capitalization. Speech has disfluencies that writing doesn't, and stripping them is a separate problem from hearing them.
- Structural editing. Reordering paragraphs, tightening an argument, cutting the redundant clause.
Layer 3 is not solved, and I'd argue it shouldn't live in the dictation layer at all. But layers 1 and 2 get conflated constantly — including by reviewers, who experience all three failing at once and file it as "the transcription is bad."
Three separate knobs control this, and mixing them up is a common source of frustration:
| Parameter | Acts on | Use it for |
|---|---|---|
stt_prompt | Recognition | Describing the situation: "A doctor dictating a patient visit note." |
keyterms_prompt | Recognition | Pinning exact spellings: names, drug names, your API's method names |
llm_instruction | The rewrite, afterward | Reshaping output: "Turn this into a bulleted list of action items." |
The first two steer what the model hears. The third reshapes what it already heard. No amount of
llm_instruction will fix a misheard word — a fluent wrong answer is worse than a messy right one, and we went deep on that failure mode in a separate post on dictation cleanup.One detail worth noting: the rewrite is best-effort. If it fails, you still get a 200 with the transcription intact and
llm_response set to null. Don't treat that as a failed request — fall back to text and move on.The thing nobody budgets for: dictation is a prompt-injection surface
This one deserves more attention than it gets.
If you're piping a raw transcript into an LLM cleanup prompt — and if you rolled your own layer 2, you are — you've built an injection vector into your own product. Not hypothetically. People say "ignore what I just said" out loud while dictating, constantly. They say "translate this into French." They say "actually, scratch that, write it as a list." Those are ordinary speech, and they're also instructions.
A naive cleanup pipeline will cheerfully execute them, and your user will watch their email get translated into French because they thought out loud.
The Dictation API passes the transcript to the model as fenced data with instructions not to act on anything inside it, so dictated commands get rewritten as speech rather than carried out. If you build the rewrite layer yourself, this is your problem to solve, and it's worth solving before launch rather than after a support ticket.
What it actually costs
Dictation pricing splits into two very different models, and the comparison is less close than you'd think.
Consumer dictation apps run subscriptions — Wispr Flow's Pro tier is $15/month, or $12/month billed annually, for unlimited words. The Dictation API is $0.62/hr of audio, all-inclusive, with recognition and cleanup in the same call.
So how much dictation before the subscription wins? About 24 hours of speech a month. At a typical dictation rate of roughly 220 words per minute, that's somewhere north of 14,000 words every working day — not written, spoken into a microphone, every day, all month.
Realistic heavy use looks different. A developer dictating 3,000 words a day across commit messages, Slack, and code comments is speaking for about 14 minutes a day, or five hours a month. That's about $3.10. Moderate use lands closer to a dollar.
The point isn't that subscriptions are a rip-off. It's that the subscription was never priced against the transcription — it's priced against the app, the syncing, the polish, and the fact that someone else maintains it. Which is a perfectly reasonable thing to pay for, right up until you want dictation inside your product, where the per-hour number is the one that matters.
A reference implementation you can read
If you'd rather read working code than a spec, Blurt is an open-source macOS dictation app built on this API — MIT licensed, native AppKit and SwiftUI, no Electron. Hold a key, talk, and the text lands in whatever app you're typing in. You bring your own API key, and the whole audio path is a single POST you can read end to end in the repo.
It's new, and it's a reference implementation rather than a polished product. That's rather the point — it exists so you can see exactly how the pieces fit before you commit to an architecture.
Back to the hammer
The reviewer's complaint was real. Her transcripts needed as much editing as she'd saved, and she reasonably concluded the technology wasn't ready.
But she was describing three different failures and calling them one. Some of it was recognition, which is measurable and has gotten dramatically better. Some of it was the cleanup layer, which most apps were building themselves out of a transcription API and a prompt, which is exactly where quality drifts and why users started noticing regressions. And some of it — the part about speech not being writing, about the app being unable to reorganize a thought — is a real, unsolved problem that no transcription model is going to fix.
That last one is worth sitting with. We've gotten good at hearing people accurately and decent at making their speech read like prose. Turning a spoken ramble into a well-structured argument is a different problem, and I don't think it belongs in the dictation layer at all. It belongs wherever the writing happens.
Which means the interesting work in dictation over the next year probably isn't in the model. It's in figuring out where the editor lives.