Qualcomm® AI HubAI Hub

Pyannote-Speaker-Diarization

Open‑source speaker diarization model identifying "who spoke when" in audio recordings.

Pyannote Speaker Diarization is an open‑source speaker diarization model that identifies "who spoke when" in an audio recording. It detects per‑frame speaker activity and encodes speaker identity, enabling accurate multi‑speaker attribution in meetings, interviews, and other multi‑party conversations.

Not supported

This model is currently not supported on any Automotive chipset.

To see performance metrics for this model on other chipsets, click the button below.

View for other chipsets

Technical Details

Input resolution (embedding):32 x 498 x 80 (batch x frames x mel-bins)
Input resolution (segmentation):32 x 1 x 80000 (batch x channel x 5s at 16kHz)
Model checkpoint:pyannote/speaker-diarization-3.1
Number of speakers (segmentation output):7 (powerset encoding for up to 3 speakers)
Quantization:Segmentation float, Embedding W8A16 via AI Hub

Applicable Scenarios

  • Meeting Transcription
  • Smart Home
  • Accessibility

License

Model:MIT

Tags

  • foundation

Supported Automotive Devices

  • SA8295P ADP

Supported Automotive Chipsets

  • Qualcomm® SA8295P

Related Models

See all models

Looking for more? See models created by industry leaders.

Discover Model Makers