
Diarization Is the Part That Answers Who Said That
Speaker diarization splits a recording into turns and tags each one, usually as Speaker A and Speaker B. It is a separate job from writing down the words, and most providers price it separately too.
The cost is small. AssemblyAI lists $0.21 per hour for its Universal-3.5 Pro async model and $0.02 per hour to add diarization, so labelled audio runs $0.23 per hour as of Aug 2026.
Pay for it whenever attribution changes the meaning. Interviews, panel discussions, user research and anything you plan to quote all need turn labels, while a solo voice memo never does.
Transcription and Diarization Are Two Separate Jobs Sold as One Feature
- ● Transcription - what was said
- ● Diarization - who took each turn
- ● Identification - which name
A transcription model turns sound into words. A diarization model ignores the words entirely and looks at voice characteristics to decide where one speaker stops and another starts.
The two run over the same audio and their outputs get stitched together. That is why a transcript can be word perfect and still credit the wrong person for a sentence.
A third step sits on top of both. Attaching a real name to Speaker B is speaker identification, and it needs either your manual edit or a stored voice profile the tool has been given in advance.
Knowing the three layers helps when something goes wrong. Bad words point at the transcription model, bad turn boundaries point at diarization, and wrong names point at whatever mapped the labels.
Vendors rarely draw the line this clearly in their marketing. A page that promises accurate meeting notes is describing all three layers at once, and the accuracy figure it quotes almost always refers to the words alone.
Ask about turn accuracy separately when you evaluate a tool. Word error rate and speaker error rate are different measurements, and a product can look excellent on one while struggling on the other.
Priced by the Hour It Is the Cheapest Line on the Invoice
- ● Async model - $0.21 per hour
- ● Diarization add on - $0.02 per hour
- ● Streaming add on - $0.12 per hour
Pay as you go providers publish the add on cost openly. The numbers are small enough that switching it off saves almost nothing.
| What you are buying | AssemblyAI price as of Aug 2026 | What it gives you |
|---|---|---|
| Universal-2 async transcription | $0.15 per hour | Words only, no turn labels |
| Universal-3.5 Pro async transcription | $0.21 per hour | Words only, higher accuracy tier |
| Async diarization add on | $0.02 per hour | Speaker A and B turn labels |
| Async experimental diarization | $0.065 per hour | Newer labelling approach |
| Streaming diarization add on | $0.12 per hour | Labels applied live during the call |
| Universal-Streaming English | $0.15 per hour | Real time words, add on priced separately |
Figures come from the published AssemblyAI pricing page as of Aug 2026, and rates for this category move often. Confirm current pricing on the official site before you budget a large batch.
Two patterns stand out. Labelling recorded audio costs about a tenth of what the transcription itself costs, while doing it live costs six times the recorded rate because the model cannot look ahead.
That live premium is the real decision. If you can wait until the call ends, you pay $0.02 rather than $0.12 for the same information.
The Failure Modes Are Predictable Once You Know the Mechanism
- ● Crosstalk blends two voices
- ● Similar voices merge into one
- ● Room mics blur distant speakers
Diarization works from voice characteristics, so anything that blurs those characteristics blurs the labels. The failures are not random and you can design around them.
Crosstalk is the worst case. Two people talking at once produce one blended signal, and the model has to split something that was never separate.
Similar voices are the second problem. Two colleagues with the same accent, pitch and speaking pace can collapse into a single label for a whole section of a meeting.
Distance matters as much as similarity. One microphone in the middle of a boardroom picks up reflections rather than clean voices, and phone bridge audio arrives already compressed. Our comparison of AI against human transcription services covers what that compression removes.
Speaker Labels Are Not the Same as Speaker Names
A raw diarized transcript gives you letters, not people. Someone still has to decide that Speaker B is the candidate rather than the interviewer, and that step is where most attribution errors enter.
Meeting apps paper over this with voice profiles. Otter lists speaker identification on its free Basic plan, and adds taggable speakers and team vocabulary on the paid tiers as of Aug 2026.
Profiles help in recurring meetings and do little for one off recordings. A new interviewee has no stored voice, so their turns arrive as an anonymous label whatever you are paying.
The practical habit is simple. Name the speakers once at the top of the file, in the first minute where each person introduces themselves, and every later reference inherits it.
Meeting Apps Hide the Same Feature Inside a Seat Price
Consumer meeting tools do not sell diarization by the hour. They fold it into a monthly seat and compete on recording limits instead, which makes the comparison a question of volume.
| Plan as of Aug 2026 | Price | Volume limit | Attribution included |
|---|---|---|---|
| Otter Basic | Free | 300 transcription minutes a month, 30 minutes per conversation | Speaker identification |
| Otter Pro | $16.99 per user monthly, $8.33 annual | 1,200 recording minutes a month, 90 minutes per meeting | Speaker identification |
| Otter Business | $30 per user monthly, $19.99 annual | Unlimited meetings, 4 hours per meeting | Taggable speakers and team vocabulary |
| Rev Free | Free | 45 AI transcription minutes a month | Automated labels |
| Rev Essentials | $25.49 monthly, $29.99 annual billing | 5,000 AI minutes per seat monthly | Automated labels |
| Rev human transcription | $1.99 per minute | Per file, no monthly cap | Human attribution, 99 percent accuracy |
Prices come from the published Otter and Rev pricing pages as of Aug 2026. Confirm current pricing on the official site, since both vendors change plan limits more often than they change headline prices.
Read the per meeting cap before the monthly one. A plan with 1,200 minutes a month and a 90 minute ceiling per meeting cuts a three hour workshop in half, which is exactly the recording where labels matter most.
The human line at the bottom sets a useful reference. At $1.99 per minute a one hour interview costs about $119, so automated labelling only has to save an hour of your cleanup time to pay for itself many times over.
Separate Microphones Beat Any Model You Can Buy
The cheapest accuracy improvement happens before the recording starts. Give each speaker their own microphone and the attribution problem largely disappears.
With separate tracks there is nothing to infer. Track one is one person and track two is another, so the software labels turns by which file the sound arrived in rather than by voice similarity.
Most remote platforms can export per participant audio, and that setting is worth finding once. It also fixes crosstalk, since two people talking over each other land in two clean files instead of one blended waveform.
In person recordings are the hard case. A single laptop microphone at the end of a table hears the far speaker as a quiet reflection, and no amount of processing turns that back into a distinct voice.
A cheap pair of lapel microphones changes the result more than a better model does. Spend there first, and treat diarization as the tidy up rather than the rescue.
What a Labelled Transcript Still Needs From You
Spot check the boundaries rather than reading every line. Jump to the moments where the conversation gets fast and see whether the labels survive the overlap.
Check the speaker count first of all. A two person interview that produced three labels means the model split one voice, and merging them is a two minute fix that saves a wrong quote.
Then verify anything you plan to publish. A quote attributed to the wrong person is the one error in a transcript that carries real consequences, and it is invisible unless you listen back.
For high stakes recordings the fallback is human work. Our roundup of the AI transcription tools worth using shows which ones expose speaker editing properly.
Which Diarization Setup Fits Your Recordings
One person recording voice notes: Skip it entirely. There is one speaker, the labels add nothing, and the add on is a line of noise in the output.
Two person interviews and podcasts: Automated diarization is reliable here and worth the cents it costs. Use separate microphones where you can, since two clean tracks make the split trivial.
Recurring team meetings: A meeting app with stored voice profiles beats raw API labels. Otter Business at $19.99 per user annually removes the naming step for the people who attend every week.
Large panels and workshops: Expect degraded labels and plan for cleanup time. More voices in one room means more merges, and a per meeting cap of 4 hours becomes the constraint to check first.
Legal, medical or journalistic use: Automated attribution is a draft, not a record. Rev human transcription at $1.99 per minute exists for exactly the cases where a misattributed sentence causes a problem.
The Question to Ask Before Turning It On
Diarization answers one narrow question about your audio, and it answers it cheaply. Everything else in the transcript comes from a different model with different failure modes.
Decide whether attribution changes what the transcript is for. If the answer is yes, pay the small add on and budget a few minutes to check the turns where people talked over each other.
If the answer is no, leave it off and keep the output clean. A single stream of text is easier to read than one broken into labels nobody needs.
Labels are only as good as the audio underneath them, and the same overlap that confuses the speaker split also garbles the words. Our breakdown of why AI transcription fails on accents and crosstalk walks through the recording conditions that cause both problems at once.
FAQ
What is the difference between diarization and speaker identification?
Diarization splits audio into turns and labels them Speaker A, Speaker B and so on without knowing who anyone is. Naming those speakers is a separate step, done by you or by an app that recognises a saved voice profile.
How much does speaker diarization cost?
Sold on its own it is cheap. AssemblyAI charges an extra $0.02 per hour on top of $0.21 per hour for its top async model as of Aug 2026, so an hour of labelled audio costs $0.23. Meeting apps bundle it into a seat price instead.
Why does diarization mix up speakers who talk over each other?
Crosstalk is the hardest case, because two overlapping voices give the model one blended signal to split. Similar voices, a single microphone in a large room and phone bridge audio all push accuracy down for the same reason.
Can diarization handle more than two speakers?
Most systems do, and many accept a hint about how many speakers to expect. Accuracy usually falls as the count rises, so a six person meeting produces more label errors than a two person interview of the same length.
Is an automatically labelled transcript good enough to quote from?
Only if the labels are correct, and that is worth spot checking. Rev prices human transcription at $1.99 per minute with 99 percent or better accuracy as of Aug 2026, which is the fallback when attribution has to hold up under scrutiny.
Sources
- Amazon Transcribe docs: Partitioning speakers (diarization) — checked 2026-09-27
- Otter.ai pricing — checked 2026-09-27
- Rev: subscription plans and per-minute pricing — checked 2026-09-27
Some links may be affiliate links. We may earn a commission at no extra cost to you.
This article was written with AI assistance. It is researched and fact-checked, not based on personal hands-on testing unless explicitly stated.
Comments
Post a Comment