Qualcomm® AI HubAI Hub

Zipformer

Transformer‑based automatic speech recognition (ASR) model for English and Chinese language.

Zipformer streaming ASR (Automatic Speech Recognition) model is a state‑of‑the‑art system designed for transcribing spoken language into written text streamingly. This model is based on the transformer architecture and has been optimized for edge inference by replacing linear layers with convolutional (conv) layers. It exhibits robust performance in realistic, noisy environments, making it highly reliable for real‑world applications. Specifically, it excels in long‑form transcription, capable of accurately transcribing audios. Time to the first token is the encoder's latency, while time to each additional token is joiner's latency, where we assume a max decoded length specified below.

Not supported

This model is currently not supported on any IoT chipset.

To see performance metrics for this model on other chipsets, click the button below.

View for other chipsets

Technical Details

Input resolution:80x71 (0.71 seconds audio)
Max decoded sequence length:200 tokens
Model checkpoint:pfluo/k2fsa-zipformer-chinese-english-mixed
Model size (decoder) (float):13.2 MB
Model size (encoder) (float):242 MB
Model size (joiner) (float):12.2 MB
Number of parameters (decoder):3.47M
Number of parameters (encoder):63.2M
Number of parameters (joiner):3.21M

Applicable Scenarios

  • Smart Home
  • Accessibility

License

Tags

  • foundation

Supported IoT Devices

  • Dragonwing IQ-8275 EVK
  • Dragonwing IQ-9075 EVK
  • Dragonwing IQ-X5121
  • Dragonwing IQ-X7181
  • Dragonwing Q-8750
  • QCS8550 (Proxy)

Supported IoT Chipsets

  • Qualcomm® Dragonwing™ IQ-X7181
  • Qualcomm® Dragonwing™ QCS8275
  • Qualcomm® Dragonwing™ QCS8550 (Proxy)
  • Qualcomm® Dragonwing™ Q-8750
  • Qualcomm® Dragonwing™ IQ-9075

Related Models

See all models

Looking for more? See models created by industry leaders.

Discover Model Makers