Qualcomm® AI HubAI Hub

Pyannote-Speaker-Diarization

Open‑source speaker diarization model identifying "who spoke when" in audio recordings.

Pyannote Speaker Diarization is an open‑source speaker diarization model that identifies "who spoke when" in an audio recording. It detects per‑frame speaker activity and encodes speaker identity, enabling accurate multi‑speaker attribution in meetings, interviews, and other multi‑party conversations.

Not supported

This model is currently not supported on any Mobile chipset.

To see performance metrics for this model on other chipsets, click the button below.

View for other chipsets

Technical Details

Input resolution (embedding):32 x 498 x 80 (batch x frames x mel-bins)
Input resolution (segmentation):32 x 1 x 80000 (batch x channel x 5s at 16kHz)
Model checkpoint:pyannote/speaker-diarization-3.1
Number of speakers (segmentation output):7 (powerset encoding for up to 3 speakers)
Quantization:Segmentation float, Embedding W8A16 via AI Hub

Applicable Scenarios

  • Meeting Transcription
  • Smart Home
  • Accessibility

Supported Mobile Form Factors

  • Phone
  • Tablet

License

Model:MIT

Tags

  • foundation

Supported Mobile Devices

  • Samsung Galaxy S21
  • Samsung Galaxy S21 Ultra
  • Samsung Galaxy S22 5G
  • Samsung Galaxy S22 Ultra 5G
  • Samsung Galaxy S22+ 5G
  • Samsung Galaxy S23
  • Samsung Galaxy S23 Ultra
  • Samsung Galaxy S23+
  • Samsung Galaxy S24
  • Samsung Galaxy S24 Ultra
  • Samsung Galaxy S24+
  • Samsung Galaxy S25
  • Samsung Galaxy S25 Ultra
  • Samsung Galaxy S25+
  • Samsung Galaxy S26
  • Samsung Galaxy S26 Ultra
  • Samsung Galaxy S26+
  • Samsung Galaxy Tab S8
  • Xiaomi 12

Supported Mobile Chipsets

  • Snapdragon® 8 Elite For Galaxy Mobile
  • Snapdragon® 8 Elite Gen 5 For Galaxy Mobile
  • Snapdragon® 8 Gen 1 Mobile
  • Snapdragon® 8 Gen 2 Mobile
  • Snapdragon® 8 Gen 3 Mobile
  • Snapdragon® 888 Mobile

Related Models

See all models

Looking for more? See models created by industry leaders.

Discover Model Makers