Qualcomm® AI HubAI Hub

Profile Job Results

Jobs
j5qe631o5
Results Ready
Name
whisper_small_v2_HfWhisperDecoder
Target Device
  • Snapdragon X Elite CRD
  • Windows 11
  • Snapdragon® X Elite | SC8380XP
Creator
ai-hub-support@qti.qualcomm.com
Target Model
Input Specs
input_ids: int32[1, 1]
attention_mask: float32[1, 1, 1, 200]
k_cache_self_0_in: float32[12, 1, 64, 199]
v_cache_self_0_in: float32[12, 1, 199, 64]
k_cache_self_1_in: float32[12, 1, 64, 199]
v_cache_self_1_in: float32[12, 1, 199, 64]
k_cache_self_2_in: float32[12, 1, 64, 199]
v_cache_self_2_in: float32[12, 1, 199, 64]
k_cache_self_3_in: float32[12, 1, 64, 199]
v_cache_self_3_in: float32[12, 1, 199, 64]
k_cache_self_4_in: float32[12, 1, 64, 199]
v_cache_self_4_in: float32[12, 1, 199, 64]
k_cache_self_5_in: float32[12, 1, 64, 199]
v_cache_self_5_in: float32[12, 1, 199, 64]
k_cache_self_6_in: float32[12, 1, 64, 199]
v_cache_self_6_in: float32[12, 1, 199, 64]
k_cache_self_7_in: float32[12, 1, 64, 199]
v_cache_self_7_in: float32[12, 1, 199, 64]
k_cache_self_8_in: float32[12, 1, 64, 199]
v_cache_self_8_in: float32[12, 1, 199, 64]
k_cache_self_9_in: float32[12, 1, 64, 199]
v_cache_self_9_in: float32[12, 1, 199, 64]
k_cache_self_10_in: float32[12, 1, 64, 199]
v_cache_self_10_in: float32[12, 1, 199, 64]
k_cache_self_11_in: float32[12, 1, 64, 199]
v_cache_self_11_in: float32[12, 1, 199, 64]
k_cache_cross_0: float32[12, 1, 64, 1500]
v_cache_cross_0: float32[12, 1, 1500, 64]
k_cache_cross_1: float32[12, 1, 64, 1500]
v_cache_cross_1: float32[12, 1, 1500, 64]
k_cache_cross_2: float32[12, 1, 64, 1500]
v_cache_cross_2: float32[12, 1, 1500, 64]
k_cache_cross_3: float32[12, 1, 64, 1500]
v_cache_cross_3: float32[12, 1, 1500, 64]
k_cache_cross_4: float32[12, 1, 64, 1500]
v_cache_cross_4: float32[12, 1, 1500, 64]
k_cache_cross_5: float32[12, 1, 64, 1500]
v_cache_cross_5: float32[12, 1, 1500, 64]
k_cache_cross_6: float32[12, 1, 64, 1500]
v_cache_cross_6: float32[12, 1, 1500, 64]
k_cache_cross_7: float32[12, 1, 64, 1500]
v_cache_cross_7: float32[12, 1, 1500, 64]
k_cache_cross_8: float32[12, 1, 64, 1500]
v_cache_cross_8: float32[12, 1, 1500, 64]
k_cache_cross_9: float32[12, 1, 64, 1500]
v_cache_cross_9: float32[12, 1, 1500, 64]
k_cache_cross_10: float32[12, 1, 64, 1500]
v_cache_cross_10: float32[12, 1, 1500, 64]
k_cache_cross_11: float32[12, 1, 64, 1500]
v_cache_cross_11: float32[12, 1, 1500, 64]
position_ids: int32[1]
Completion Time
4/8/2025, 9:09:15 AM
Versions
  • ONNX Runtime: 1.21.0+ (commit a46d212)
  • QAIRT: v2.32.0.250228225014_116386
  • Windows: Windows 11 (26100)
  • Build ID: APSS.WP_HA.1.0.c90-07350-SC8380XPSRSFNWZA-2
  • AI Hub: aihub-2025.03.24.0
Estimated Inference Time
22.8 ms
Estimated Peak Memory Usage
227 MB
Compute Units
NPU
2301
StageTimeMemory
First App Load
1.78 min2 GB
Subsequent App Load
682 ms367 MB
Inference
22.8 ms227 MB
ONNX RuntimeValue
execution_modeSEQUENTIAL
intra_op_num_threads0
inter_op_num_threads0
enable_memory_patternfalse
enable_cpu_memory_arenafalse
graph_optimization_levelENABLE_ALL
QNN Execution ProviderValue
htp_performance_mode"burst"
htp_graph_finalization_optimization_mode"3"
enable_htp_fp16_precision"1"

Sign up to run this model on a hosted Qualcomm® device!

Run on device