Customer story

Bluejay x pyannoteAI

Bluejay x pyannoteAI

An end-to-end testing and observability for conversational agents

An end-to-end testing and observability for conversational agents

Observability starts with “who spoke when”: How Bluejay powers reliable Voice Agent evaluation?

Observability starts with “who spoke when”: How Bluejay powers reliable Voice Agent evaluation?

About

About

Bluejay

Bluejay

About Bluejay

Bluejay is an end-to-end testing and observability platform for conversational agents (both voice AI and chatbots). The YC-backed company serves teams at Google, DoorDash, Presto, and many others who build and deploy AI agents at scale.

Bluejay’s platform has two core components:

  • Simulation testing to generate synthetic customer conversations and stress-test AI agents.

  • Live observability to monitor production calls and evaluate agent behavior at scale.

The company processes millions of conversations annually across industries, including healthcare, finance, food delivery, and big tech.

Their approach goes beyond basic transcription. Bluejay ingests audio, transcripts, tool calls, traces, and custom metadata, then runs two types of evaluations:

  • Deterministic evaluations: Latency, interruptions, handoff behavior

  • LLM-based evaluations: Customer satisfaction, problem resolution, toxicity, compliance

Because each industry has different requirements, Bluejay's platform allows customers to define the evaluations that matter most for their specific context.

Initial situation

Bluejay's mission is to build the trust layer between humans and AI agents, and between AI agents and other AI agents.

To build that trust, they ingest far more than transcripts. They combine audio, transcripts, tool calls, traces, and custom metadata. Additionally, they conduct deterministic evaluations, including latency and interruption detection, as well as LLM-based evaluations, such as CSAT, problem resolution, and compliance.

But there was a bottleneck.

Before implementing pyannoteAI, speaker diarization was unreliable, especially on single-channel recordings with multiple speakers.

This became critical when:

  • An AI agent handed off to a human agent.

  • IVR voices changed mid-call.

  • Multiple participants spoke over each other.

Wrong speaker attribution meant broken metrics. Suppose the system cannot tell who interrupted whom; latency and turn-taking metrics collapse. When that happens, the value of an observability platform goes to zero.

Bluejay was facing customer complaints and manual debugging cycles caused by inaccurate diarization. They needed reliable speaker tracking at a production scale.

“Risk of incorrect evaluations would make the value of our product go to zero if mis-assigned. If speaker attribution is wrong, the value of our evaluations goes to zero.”

“Risk of incorrect evaluations would make the value of our product go to zero if mis-assigned. If speaker attribution is wrong, the value of our evaluations goes to zero.”

Rohan Vasishth

CEO and co-founder @ Bluejay

Objectives

Bluejay’s goals were clear:

  • Achieve high-fidelity conversation understanding across complex audio scenarios.

  • Make deterministic evaluations, especially latency and interruptions, trustworthy.

  • Reduce customer complaints tied to speaker misattribution.

  • Scale without increasing manual QA overhead.

They were not looking for incremental improvement. They needed a diarization that simply worked in real-world calls.

Solution

After testing multiple diarization engines, Bluejay selected pyannoteAI’s diarization model.

“It was the first solution that consistently separated speakers correctly in difficult, multi-speaker, single-channel calls.”

Rohan Vasishth

CEO and co-founder @ Bluejay

Integration was direct and fast. From initial testing to production rollout, it took less than four days.

How does pyannoteAI fit into their pipeline?

Their workflow today:

  1. Receive raw audio and optional metadata.

  2. Run diarization with pyannoteAI.

  3. Run transcription through various transcription providers, depending on privacy requirements.

  4. Feed the diarized transcript, audio, tool calls, and traces into deterministic and LLM-based evaluation modules.

pyannoteAI plays a central role in deterministic evaluations, especially:

  • Interruption detection to determine who spoke over whom.

  • Latency measurement to calculate reaction times between the agent and the customer.

Bluejay relies on pyannoteAI timestamps over standard STT timestamps for these metrics. More accurate speaker segmentation leads directly to more reliable evaluation outputs.

Results

Diarization quality

Before pyannoteAI, diarization quality was inconsistent and often unusable for advanced metrics.

After implementation:

  • Significant improvement in speaker separation accuracy, especially for hybrid AI-human calls.

  • Direct reduction in customer complaints related to speaker identity errors.

  • Stable performance while scaling to tens of millions of conversations annually.

Evaluation reliability

Accurate diarization improved downstream evaluations:

  • Interruption detection became reliable for large enterprise customers where this metric is business-critical.

  • Latency measurements became more precise due to accurate speaker timestamps.

  • Deterministic metrics aligned better with real conversation behavior.

Because LLM-based evaluations depend on structured conversation data, better diarization also reduced false positives and false negatives in automated scoring.

Business impact

Bluejay’s customers saw measurable operational gains:

  • A healthcare client reduced agent deployment cycles from roughly two weeks to two days by using Bluejay’s automated testing.

  • Fewer issues surfaced late in manual UAT after simulation testing.

  • Increased confidence in automated evaluation results across production environments.

These outcomes depend on accurate conversation representation. pyannoteAI is a core technical dependency behind that reliability.

How accurate metadata changes everything?

Technical Foundation

PyannoteAI provides the conversation metadata that makes Bluejay's evaluations possible. Without accurate speaker identification and timestamps:

  • Interruption detection would be unreliable

  • Latency measurements would be inaccurate

  • Turn-taking analysis would fail

  • The entire evaluation pipeline would be compromised

Integration Simplicity

The API integration took 3-4 days from testing to production. This rapid deployment allowed Bluejay to focus on building evaluation logic rather than debugging diarization infrastructure.

Scalability

PyannoteAI handles Bluejay's volume (millions of conversations per year) without degradation in accuracy. This consistent performance is critical for a platform that promises automated, reliable agent testing.

What's next

Bluejay is now evaluating deeper consolidation. Their long-term goal is to work with a unified provider delivering accurate diarization, robust multilingual transcription, and privacy-safe audio processing at scale.

For Bluejay, understanding what was said is not enough. Understanding who said it, when, and in what context is what makes their observability platform credible. pyannoteAI sits at the center of that layer.

Speaker Intelligence Platform for developers

Detect, segment, label and separate speakers in any language.

Speaker Intelligence Platform for developers

Detect, segment, label and separate speakers in any language.

Make the most of conversational speech
with AI

Detect, segment, label and separate speakers in any language.