
Customer story
About Bluejay
Bluejay is an end-to-end testing and observability platform for conversational agents (both voice AI and chatbots). The YC-backed company serves teams at Google, DoorDash, Presto, and many others who build and deploy AI agents at scale.
Bluejay’s platform has two core components:
Simulation testing to generate synthetic customer conversations and stress-test AI agents.
Live observability to monitor production calls and evaluate agent behavior at scale.
The company processes millions of conversations annually across industries, including healthcare, finance, food delivery, and big tech.
Their approach goes beyond basic transcription. Bluejay ingests audio, transcripts, tool calls, traces, and custom metadata, then runs two types of evaluations:
Deterministic evaluations: Latency, interruptions, handoff behavior
LLM-based evaluations: Customer satisfaction, problem resolution, toxicity, compliance
Because each industry has different requirements, Bluejay's platform allows customers to define the evaluations that matter most for their specific context.
Initial situation
Bluejay's mission is to build the trust layer between humans and AI agents, and between AI agents and other AI agents.
To build that trust, they ingest far more than transcripts. They combine audio, transcripts, tool calls, traces, and custom metadata. Additionally, they conduct deterministic evaluations, including latency and interruption detection, as well as LLM-based evaluations, such as CSAT, problem resolution, and compliance.
But there was a bottleneck.
Before implementing pyannoteAI, speaker diarization was unreliable, especially on single-channel recordings with multiple speakers.
This became critical when:
An AI agent handed off to a human agent.
IVR voices changed mid-call.
Multiple participants spoke over each other.
Wrong speaker attribution meant broken metrics. Suppose the system cannot tell who interrupted whom; latency and turn-taking metrics collapse. When that happens, the value of an observability platform goes to zero.
Bluejay was facing customer complaints and manual debugging cycles caused by inaccurate diarization. They needed reliable speaker tracking at a production scale.
Rohan Vasishth
CEO and co-founder @ Bluejay
Objectives
Bluejay’s goals were clear:
Achieve high-fidelity conversation understanding across complex audio scenarios.
Make deterministic evaluations, especially latency and interruptions, trustworthy.
Reduce customer complaints tied to speaker misattribution.
Scale without increasing manual QA overhead.
They were not looking for incremental improvement. They needed a diarization that simply worked in real-world calls.

Solution
After testing multiple diarization engines, Bluejay selected pyannoteAI’s diarization model.
“It was the first solution that consistently separated speakers correctly in difficult, multi-speaker, single-channel calls.”
Rohan Vasishth
CEO and co-founder @ Bluejay
Integration was direct and fast. From initial testing to production rollout, it took less than four days.
How does pyannoteAI fit into their pipeline?
Their workflow today:
Receive raw audio and optional metadata.
Run diarization with pyannoteAI.
Run transcription through various transcription providers, depending on privacy requirements.
Feed the diarized transcript, audio, tool calls, and traces into deterministic and LLM-based evaluation modules.
pyannoteAI plays a central role in deterministic evaluations, especially:
Interruption detection to determine who spoke over whom.
Latency measurement to calculate reaction times between the agent and the customer.
Bluejay relies on pyannoteAI timestamps over standard STT timestamps for these metrics. More accurate speaker segmentation leads directly to more reliable evaluation outputs.
Results
Diarization quality
Before pyannoteAI, diarization quality was inconsistent and often unusable for advanced metrics.
After implementation:
Significant improvement in speaker separation accuracy, especially for hybrid AI-human calls.
Direct reduction in customer complaints related to speaker identity errors.
Stable performance while scaling to tens of millions of conversations annually.
Evaluation reliability
Accurate diarization improved downstream evaluations:
Interruption detection became reliable for large enterprise customers where this metric is business-critical.
Latency measurements became more precise due to accurate speaker timestamps.
Deterministic metrics aligned better with real conversation behavior.
Because LLM-based evaluations depend on structured conversation data, better diarization also reduced false positives and false negatives in automated scoring.

Business impact
Bluejay’s customers saw measurable operational gains:
A healthcare client reduced agent deployment cycles from roughly two weeks to two days by using Bluejay’s automated testing.
Fewer issues surfaced late in manual UAT after simulation testing.
Increased confidence in automated evaluation results across production environments.
These outcomes depend on accurate conversation representation. pyannoteAI is a core technical dependency behind that reliability.
How accurate metadata changes everything?
Technical Foundation
PyannoteAI provides the conversation metadata that makes Bluejay's evaluations possible. Without accurate speaker identification and timestamps:
Interruption detection would be unreliable
Latency measurements would be inaccurate
Turn-taking analysis would fail
The entire evaluation pipeline would be compromised
Integration Simplicity
The API integration took 3-4 days from testing to production. This rapid deployment allowed Bluejay to focus on building evaluation logic rather than debugging diarization infrastructure.
Scalability
PyannoteAI handles Bluejay's volume (millions of conversations per year) without degradation in accuracy. This consistent performance is critical for a platform that promises automated, reliable agent testing.
What's next
Bluejay is now evaluating deeper consolidation. Their long-term goal is to work with a unified provider delivering accurate diarization, robust multilingual transcription, and privacy-safe audio processing at scale.
For Bluejay, understanding what was said is not enough. Understanding who said it, when, and in what context is what makes their observability platform credible. pyannoteAI sits at the center of that layer.
Discover more stories




