Qualcomm® AI HubAI Hub

OWL-ViT

Open‑Vocabulary Object Detection with Vision Transformers.

OWL‑ViT (Open‑World Localization with Vision Transformers) is an open‑vocabulary object detector that uses a CLIP‑based ViT‑B/32 backbone. Given an image and one or more free‑form text queries, the model predicts bounding boxes and confidence scores for each query.

Not supported

This model is currently not supported on any Mobile chipset.

To see performance metrics for this model on other chipsets, click the button below.

View for other chipsets

Technical Details

Input image resolution:768x768
Model checkpoint:google/owlvit-base-patch32
Model size (float):613 MB
Number of parameters:153M

Applicable Scenarios

  • Open-Vocabulary Detection
  • Zero-Shot Object Detection
  • Factory Automation

Supported Mobile Form Factors

  • Phone
  • Tablet

License

Tags

  • foundation

Supported Mobile Devices

  • Samsung Galaxy S21
  • Samsung Galaxy S21 Ultra
  • Samsung Galaxy S22 5G
  • Samsung Galaxy S22 Ultra 5G
  • Samsung Galaxy S22+ 5G
  • Samsung Galaxy S23
  • Samsung Galaxy S23 Ultra
  • Samsung Galaxy S23+
  • Samsung Galaxy S24
  • Samsung Galaxy S24 Ultra
  • Samsung Galaxy S24+
  • Samsung Galaxy S25
  • Samsung Galaxy S25 Ultra
  • Samsung Galaxy S25+
  • Samsung Galaxy S26
  • Samsung Galaxy S26 Ultra
  • Samsung Galaxy S26+
  • Samsung Galaxy Tab S8
  • Snapdragon 7 Gen 4 QRD
  • Snapdragon 8 Elite Gen 5 QRD
  • Snapdragon 8 Elite QRD
  • Xiaomi 12

Supported Mobile Chipsets

  • Snapdragon® 7 Gen 4 Mobile
  • Snapdragon® 8 Elite Mobile
  • Snapdragon® 8 Elite Gen 5 Mobile
  • Snapdragon® 8 Gen 1 Mobile
  • Snapdragon® 8 Gen 2 Mobile
  • Snapdragon® 8 Gen 3 Mobile
  • Snapdragon® 888 Mobile

Related Models

See all models

Looking for more? See models created by industry leaders.

Discover Model Makers