Qualcomm® AI HubAI Hub

OWL-ViT

Open‑Vocabulary Object Detection with Vision Transformers.

OWL‑ViT (Open‑World Localization with Vision Transformers) is an open‑vocabulary object detector that uses a CLIP‑based ViT‑B/32 backbone. Given an image and one or more free‑form text queries, the model predicts bounding boxes and confidence scores for each query.

Not supported

This model is currently not supported on any Compute chipset.

To see performance metrics for this model on other chipsets, click the button below.

View for other chipsets

Technical Details

Input image resolution:768x768
Model checkpoint:google/owlvit-base-patch32
Model size (float):613 MB
Number of parameters:153M

Applicable Scenarios

  • Open-Vocabulary Detection
  • Zero-Shot Object Detection
  • Factory Automation

License

Tags

  • foundation

Supported Compute Devices

  • Snapdragon X Elite CRD
  • Snapdragon X Plus 8-Core CRD
  • Snapdragon X2 Elite CRD

Supported Compute Chipsets

  • Snapdragon® X Elite
  • Snapdragon® X Plus 8-Core
  • Snapdragon® X2 Elite

Related Models

See all models

Looking for more? See models created by industry leaders.

Discover Model Makers