Wav2Vec2-Conformer-Large-960h
English speech recognition with Wav2Vec2‑Conformer encoder and CTC decoder.
Facebook Wav2Vec2‑Conformer‑Large is an English ASR model fine‑tuned on 960 hours of LibriSpeech. It uses a Conformer encoder with relative‑position attention (24 layers, 1024 hidden dim, 16 heads) and a CTC head for decoding. The model accepts raw 16kHz audio and outputs transcribed text via CTC greedy decoding.
Not supported
This model is currently not supported on any IoT chipset.
To see performance metrics for this model on other chipsets, click the button below.
View for other chipsetsTechnical Details
Input resolution:1x160000 (10s at 16kHz)
Model checkpoint:facebook/wav2vec2-conformer-rel-pos-large-960h-ft
Model size (float):~2.4 GB
Number of parameters:~600M
Applicable Scenarios
- Smart Home
- Accessibility
- Transcription
License
Model:APACHE-2.0
Tags
- real-time
Supported IoT Devices
- Arduino VENTUNO Q
- Dragonwing IQ-9075 EVK
- Dragonwing IQ-X5121
- Dragonwing IQ-X7181
- Dragonwing Q-6690 MTP
- Dragonwing Q-7790
- Dragonwing Q-8750
- QCS8550 (Proxy)
Supported IoT Chipsets
- Qualcomm® Dragonwing™ Q-6690
- Qualcomm® QCS5121
- Qualcomm® Dragonwing™ IQ-X7181
- Qualcomm® Dragonwing™ Q-7790
- Qualcomm® Dragonwing™ IQ-8275
- Qualcomm® Dragonwing™ QCS8550 (Proxy)
- Qualcomm® Dragonwing™ Q-8750
- Qualcomm® Dragonwing™ IQ-9075
Related Models
See all modelsLooking for more? See models created by industry leaders.
Discover Model Makers










