Wav2Vec2-Conformer-Large-960h
English speech recognition with Wav2Vec2‑Conformer encoder and CTC decoder.
Facebook Wav2Vec2‑Conformer‑Large is an English ASR model fine‑tuned on 960 hours of LibriSpeech. It uses a Conformer encoder with relative‑position attention (24 layers, 1024 hidden dim, 16 heads) and a CTC head for decoding. The model accepts raw 16kHz audio and outputs transcribed text via CTC greedy decoding.
Not supported
This model is currently not supported on any Automotive chipset.
To see performance metrics for this model on other chipsets, click the button below.
View for other chipsetsTechnical Details
Input resolution:1x160000 (10s at 16kHz)
Model checkpoint:facebook/wav2vec2-conformer-rel-pos-large-960h-ft
Model size (float):~2.4 GB
Number of parameters:~600M
Applicable Scenarios
- Smart Home
- Accessibility
- Transcription
License
Model:APACHE-2.0
Tags
- real-time
Supported Automotive Devices
- SA8295P ADP
Supported Automotive Chipsets
- Qualcomm® SA8295P
Related Models
See all modelsLooking for more? See models created by industry leaders.
Discover Model Makers










