
Customer story
About Notta
Notta is one of the most widely adopted AI meeting productivity platforms in Japan, used by business professionals and enterprise teams across consulting, sales, HR, and engineering. The platform has 15 million users globally, serves 5,000 enterprise customers, and has processed millions of hours of transcription, with a large share coming from the Japanese market. It supports transcription in 58 languages, along with real-time translation, one-click AI summaries, and collaboration features.
Notta also builds Notta Memo, a 28-gram, credit-card-sized AI voice recorder with four MEMS microphones and one bone conduction microphone, built for in-person and hybrid meetings.
Growth has largely been bottom-up: one person adopts Notta, and usage spreads through the organization from there. At Yachiyo Engineering, a single employee's purchase eventually led to a deployment covering more than 500 of the company's 1,000 employees.
None of the value of Notta works without knowing who is speaking. Meeting summaries, action items, and decisions are only as useful as the speaker context behind them, and that context has to hold up in real meeting rooms, not just clean recordings.
TJ Dai
CTO of Notta
Where speaker attribution gets hard?
Notta's hardest cases aren't the meetings with a mic per participant. They're the ones without any platform-level per-participant audio at all.
Hybrid meetings on a single shared device. A typical hard scenario: several participants in one conference room, all captured through a single Notta Memo device, while others join remotely. Unlike Zoom or Teams, there's no separate audio stream per person. pyannoteAI has to distinguish speakers directly from that shared audio.
Japanese conversation patterns. Japanese meetings carry their own diarization challenges: frequent short backchannels, or aizuchi, such as "はい," "うん," and "なるほど," appear far more often than in English conversation and can read as false speaker changes if a system isn't built to handle them. Structured turn-taking with short utterances also demands tighter precision on where one speaker's turn ends and the next one begins.
Audio quality variance in hybrid setups. In-room and remote participants rarely produce comparable audio quality, and that gap has to be accounted for rather than treated as noise.

Where pyannoteAI fits
pyannoteAI runs as a post-processing step after transcription, adding speaker attribution to completed transcripts. That's especially important for multi-speaker recordings and Notta Memo scenarios where no per-participant stream exists to fall back on.
The integration isn't part of Notta's real-time pipeline today. It's focused on improving speaker attribution after the transcription stage, which is where the biggest usability gap was: without speaker context, it's much harder for a user to reconstruct who said what, and downstream features like summaries, action items, and decisions become far less reliable.
What changed
Before integrating pyannoteAI, speaker attribution wasn't consistently available across every transcription scenario. With it, Notta has extended speaker-aware transcription across a broader set of languages, recordings, and meeting formats, particularly for non-Japanese and multi-speaker cases where attribution used to be limited or missing entirely.
The clearest impact shows up downstream. Meeting summaries, action items, decisions, and follow-ups all become more useful once they're tied to the right participant, and that matters most in multi-person meetings and Notta Memo recordings, where users expect the output to reflect who actually said what.
"Speaker diarization is not just an isolated feature for us. It is an important foundation for downstream AI experiences. Our AI summaries, action items, meeting analytics, and the Notta Memo experience all become more useful when speaker context is accurate and reliable."
TJ Dai
CTO of Notta
Notta Memo, specifically
Notta Memo captures multi-speaker, in-person audio with no platform-level speaker separation built in. That makes pyannoteAI's diarization essential rather than incremental. It's what allows Notta Memo to hold up in real-world meeting rooms instead of only in ideal recording conditions.
What's next
Notta is deepening the integration on a few fronts:
Post-processing quality. Continued improvement to speaker attribution for completed recordings, with particular focus on multi-speaker meetings and Notta Memo scenarios.
Broader language and meeting coverage. Extending diarization quality across more languages and meeting formats, with continued attention to Japanese business meetings and hybrid setups.
Real-time diarization. An area Notta may explore going forward. Once real-time speaker attribution moves into the live transcription pipeline, latency and attribution delay become the metrics that matter.
Speaker continuity. Exploring ways to track speaker identity across recordings, with privacy safeguards and user controls built in from the start.
Discover more stories



