Precision-3 now available on API and On-Premise

Precision-3 now available on API and On-Premise

Our latest diarization model, Precision-3, is available now on the API and for on-premises deployments.

Added

  • Two new input parameters: vadSensitivity and crosstalkSensitivity. They adjust voice activity and overlap detection independently of one another.

  • Three new output scores: speakerProbability, speechProbability, and crosstalkProbability. Each answers one question: is this speaker active, is speech present, is there crosstalk.

  • Three on-premise operating modes: speed, balance, and accuracy. Each sets a different trade-off between processing speed and diarization quality.

Changed

  • Average diarization error rate drops 10.4% against Precision-2, from 16.02 to 14.35 DER across 15 benchmark datasets. Over half of that improvement (53%) comes from more accurate speaker attribution: the model assigns speech to the wrong speaker 18.7% less often.

  • On-premise, vadSensitivity replaces vad_aggressiveness and inverts the scale: high sensitivity now means more speech detected, where high aggressiveness previously meant the opposite.

Deprecated

  • The frame-level confidence output is deprecated in favor of speakerProbability, speechProbability, and crosstalkProbability. turnLevelConfidence is unchanged.

Benchmark results:

Pricing is unchanged. Existing integrations keep running Precision-2 until you set model: "precision-3" or wait for the default switch on October 1. Precision-2 is deprecated on October 15.


๐Ÿ“’ Read more about Precision-3 in our blogpost
๐Ÿ”Ž Explore the technical tutorials to deploy Input parameters and Output scores
โ›ณ๏ธ Follow the migration guide from Precision-2 to Precision-3
๐Ÿ“น Check your guide video tutorial

Speaker Intelligence Platform for developers

Detect, segment, label and separate speakers in any language.

Speaker Intelligence Platform for developers

Detect, segment, label and separate speakers in any language.

Make the most of conversational speech
with AI

Detect, segment, label and separate speakers in any language.