Qualcomm® AI HubAI Hub

OWL-ViT

Open‑Vocabulary Object Detection with Vision Transformers.

OWL‑ViT (Open‑World Localization with Vision Transformers) is an open‑vocabulary object detector that uses a CLIP‑based ViT‑B/32 backbone. Given an image and one or more free‑form text queries, the model predicts bounding boxes and confidence scores for each query.

Not supported

This model is currently not supported on any IoT chipset.

To see performance metrics for this model on other chipsets, click the button below.

View for other chipsets

Technical Details

Input image resolution:768x768
Model checkpoint:google/owlvit-base-patch32
Model size (float):613 MB
Number of parameters:153M

Applicable Scenarios

  • Open-Vocabulary Detection
  • Zero-Shot Object Detection
  • Factory Automation

License

Tags

  • foundation

Supported IoT Devices

  • Arduino VENTUNO Q
  • Dragonwing IQ-8275 EVK
  • Dragonwing IQ-9075 EVK
  • Dragonwing IQ-X5121
  • Dragonwing IQ-X7181
  • Dragonwing Q-6690 MTP
  • Dragonwing Q-7790
  • Dragonwing Q-8750
  • Dragonwing RB3 Gen 2 Vision Kit
  • QCS8550 (Proxy)

Supported IoT Chipsets

  • Qualcomm® Dragonwing™ Q-6690
  • Qualcomm® QCS5121
  • Qualcomm® Dragonwing™ QCS6490
  • Qualcomm® Dragonwing™ IQ-X7181
  • Qualcomm® Dragonwing™ Q-7790
  • Qualcomm® Dragonwing™ IQ-8275
  • Qualcomm® Dragonwing™ QCS8550 (Proxy)
  • Qualcomm® Dragonwing™ Q-8750
  • Qualcomm® Dragonwing™ IQ-9075

Related Models

See all models

Looking for more? See models created by industry leaders.

Discover Model Makers