SigLIP2
Zero‑shot image‑text similarity and classification using SigLIP2.
SigLIP2 (Sigmoid Loss for Language‑Image Pre‑training 2) is a vision‑language model from Google that computes cosine‑similarity scores between images and text prompts. It can be used for zero‑shot image classification, image search, and content moderation without any task‑specific fine‑tuning.
Not supported
This model is currently not supported on any Compute chipset.
To see performance metrics for this model on other chipsets, click the button below.
View for other chipsetsTechnical Details
Image input resolution:224x224
Model checkpoint:google/siglip2-base-patch16-224
Model size (image_encoder) (float):352 MB
Model size (image_encoder) (w8a16):92.8 MB
Model size (text_encoder) (float):1.05 GB
Model size (text_encoder) (w8a16):461 MB
Text sequence length:64
Applicable Scenarios
- Image Search
- Content Moderation
- Zero-Shot Classification
License
Model:APACHE-2.0
Tags
- foundation
Supported Compute Devices
- Snapdragon X Elite CRD
- Snapdragon X Plus 8-Core CRD
- Snapdragon X2 Elite CRD
Supported Compute Chipsets
- Snapdragon® X Elite
- Snapdragon® X Plus 8-Core
- Snapdragon® X2 Elite
Related Models
See all modelsLooking for more? See models created by industry leaders.
Discover Model Makers









