Qualcomm® AI HubAI Hub

SigLIP2

Zero‑shot image‑text similarity and classification using SigLIP2.

SigLIP2 (Sigmoid Loss for Language‑Image Pre‑training 2) is a vision‑language model from Google that computes cosine‑similarity scores between images and text prompts. It can be used for zero‑shot image classification, image search, and content moderation without any task‑specific fine‑tuning.

Not supported

This model is currently not supported on any IoT chipset.

To see performance metrics for this model on other chipsets, click the button below.

View for other chipsets

Technical Details

Image input resolution:224x224
Model checkpoint:google/siglip2-base-patch16-224
Model size (image_encoder) (float):352 MB
Model size (image_encoder) (w8a16):92.8 MB
Model size (text_encoder) (float):1.05 GB
Model size (text_encoder) (w8a16):461 MB
Text sequence length:64

Applicable Scenarios

  • Image Search
  • Content Moderation
  • Zero-Shot Classification

License

Tags

  • foundation

Supported IoT Devices

  • Arduino VENTUNO Q
  • Dragonwing IQ-8275 EVK
  • Dragonwing IQ-9075 EVK
  • Dragonwing IQ-X5121
  • Dragonwing IQ-X7181
  • Dragonwing Q-6690 MTP
  • Dragonwing Q-7790
  • Dragonwing Q-8750
  • Dragonwing RB3 Gen 2 Vision Kit
  • QCS8550 (Proxy)

Supported IoT Chipsets

  • Qualcomm® Dragonwing™ Q-6690
  • Qualcomm® QCS5121
  • Qualcomm® Dragonwing™ QCS6490
  • Qualcomm® Dragonwing™ IQ-X7181
  • Qualcomm® Dragonwing™ Q-7790
  • Qualcomm® Dragonwing™ IQ-8275
  • Qualcomm® Dragonwing™ QCS8550 (Proxy)
  • Qualcomm® Dragonwing™ Q-8750
  • Qualcomm® Dragonwing™ IQ-9075

Related Models

See all models

Looking for more? See models created by industry leaders.

Discover Model Makers