Blog

TL;DR: To build a live speaker-attributed transcript, run pyannoteAI streaming diarization (with Live-1) alongside a streaming ASR and align their outputs on a shared clock. Use a reconciler that assigns each word to the active speaker when the speaker count is large or unknown. Use one ASR stream per speaker when the count is small and overlapping speech is common.
pyannoteAI now offers real-time speaker diarization, powered by our streaming model live-1. It unlocks applications where you need to know who is speaking as the conversation happens, not minutes later. Think live subtitling and dubbing, real-time meeting assistants, multi-party voice agents, and more.
Diarization tells you who is speaking and when. But for most of these applications, you also need to know what was said. That requires transcription. Put diarization and transcription together, and you get a speaker-attributed transcript: the thing almost every real-world voice-based product needs.
Combining the two live is harder than doing it with batch processing on a pre-recorded audio file.
In a batch deployment setting, you have the complete output of both systems before you start. You align everything in one pass.
In a live setting, words and speaker segments arrive piece by piece and can be revised. You have to attribute each word with incomplete information, under a latency budget, and correct earlier decisions as new data comes in.
This post focuses on exactly that: how to combine diarization and transcription in real time to get correctly attributed transcripts. It shares provider-agnostic best practices: the same principles apply whichever transcription engine you pair with pyannoteAI diarization.
Two companion tutorials show end-to-end code examples:
Framing the context
How real-time diarization works
Streaming diarization consumes audio and emits speaker events. pyannoteAI sends a JSON event when a speaker starts speaking and another when a speaker stops speaking. Each event carries a speaker label (for example, SPEAKER_00) and a timestamp. See documentation for more details.
How real-time transcription works
Streaming transcription consumes the same audio and emits timestamped words, each with a start and an end.
Two mechanisms matter when you combine transcription with diarization.
Turn detection. Many ASRs detect when a speaker has finished talking: the end of a turn. One turn is not the same as a diarization segment. A single turn can span several diarization segments, for example, when a speaker pauses mid-thought, then starts again to finish their sentence.
Partials vs. finals. Most streaming ASRs emit both partial and final results. Partials arrive fast but can be revised. Finals are slower but are more accurate. Once a final is emitted, the text is immutable. What triggers a final varies by provider: a fixed time cadence, every N tokens, or at the end of a detected turn.
Combining diarization and transcription
The objective
Diarization provides timestamped speaker segments (who, when). Transcription provides timestamped words (what). To produce a speaker-attributed transcript, you have to reconcile these two outputs. The rest of this post explains how.

General best practices
Get the audio format right for each service: pyannoteAI diarization and your ASR provider may have different audio requirements. pyannoteAI streaming expects 16 kHz mono float32 PCM in 100 ms chunks. Your ASR may want 16-bit PCM, a different sample rate, or a different chunk size. Confirm both before you start. Capture the microphone once, then convert the stream independently for each destination.
Let diarization handle speaker attribution: Many ASR providers ship their own diarization or speaker-label feature. Bundled pipelines can be enough for simple cases, but pyannoteAI diarization is more accurate. As soon as speaker accuracy matters, use the ASR only for words, and take every speaker decision from diarization.

Keep a shared internal clock. Diarization and transcription are separate processes, so their timestamps need a common reference. You have two options.
Wait for both connections to be ready before you send any audio, so they share a start instant.
Or log the moment each stream starts, compute the offset between them, and apply it to every timestamp you receive later.
Reconciliation options
There are two main approaches to obtaining a speaker-attributed transcript.
Approach 1: build a reconciler that aligns words to speaker segments
In this approach, diarization and transcription run in parallel on the same audio. One stream gives you timestamped words, the other timestamped speaker segments. A reconciler then assigns each word to a speaker in real time.
The reconciler is a set of rules designed to assign words to speakers. Here is a typical rule set, which you can adapt:
Words are assigned to the speaker whose diarization segment is active at the word's time.
When a word overlaps more than one speaker segment, it is assigned to the speaker whose segment started first.
When a word matches no diarization segment, it is assigned to the speaker who stopped talking most recently.
The transcript is continuously re-evaluated as new information arrives: a new word, an ASR revision, a new speaker segment, a segment ending.
When to use this approach? A single ASR stream keeps costs flat, no matter how many people join. You also keep full control of the attribution logic. The trade-off: that logic relies on heuristics, and heuristics struggle when people talk at once, since one mixed stream can't separate simultaneous voices. Errors also cluster around the moments speakers change.
Reach for Approach 1 when the speaker count is large, variable, or unknown, and when cost matters: multi-party meetings, conferences, webinars and panels, or podcasts with a rotating cast.
This approach is demonstrated in detail with code in 🔗 Real-time diarized transcription with OpenAI.
Approach 2: split the input into one audio stream per speaker
In this approach, you don’t reconcile words after the fact. Instead, you use diarization to route each speaker to their own transcription stream.
Pick a maximum number of speakers,
M: a hard cap on how many parallel transcription streams you'll ever open.Run diarization on the mixed input. Open
Mchild streams, and assign one to each speaker the diarization discovers.In each child stream, preserve the original timeline. Frames where that stream's speaker is talking pass through unchanged. Every other frame becomes silent.
Send each speaker-specific stream to the ASR on its own.
When to use this approach? In the multi-stream approach, each stream carries one speaker, which means there's no reconciler to build or maintain. Overlapping speech is handled cleanly because each speaker is isolated, and the ASR transcribes crosstalk one voice at a time. The trade-off is cost and rigidity. You cap the speaker count at M in advance; any speaker beyond M is dropped, and you pay for M streams for the full session, even silent ones. Cost scales with M, not with how much anyone talks.
Reach for Approach 2 when the speaker count is small and fixed, when accuracy matters more than cost, or when people interrupt each other often: one-to-one interviews, doctor–patient consultation, and sales and discovery calls.
This approach is demonstrated in detail with code in 🔗 Real-time diarized transcription with AssemblyAI.
Where to go next
Both approaches are worth understanding before you commit. As a rule of thumb, reach for the reconciler when the speaker count is open-ended and cost matters, and for one-stream-per-speaker when the speaker count is small and overlapping speech is common.
Ready to build it? The two companion tutorials implement each approach end-to-end:
🔗 Real-time diarized transcription with OpenAI: a full reconciler.
🔗 Real-time diarized transcription with AssemblyAI: one stream per speaker, using partials, finals, and turn detection.
FAQ:
How do you combine speaker diarization with real-time transcription? Run both on the same audio, align their timestamps, then either assign each word to the active speaker segment (reconciler) or route each speaker to their own ASR stream.
Should I use my ASR provider's built-in diarization? It can work for simple cases. When speaker accuracy matters, use the ASR for words only and take every speaker decision from dedicated diarization.
How do you handle overlapping speech in live transcription? Splitting the input into one stream per speaker isolates each voice, so the ASR transcribes crosstalk one speaker at a time.
