Qualcomm® AI HubAI Hub

Gemma-4-E2B-it

Lightweight multimodal model from Google DeepMind handling text and image input.

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on small models) and generating text output. This release includes open‑weights models in both pre‑trained and instruction‑tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

Not supported

This model is currently not supported on any Compute chipset.

To see performance metrics for this model on other chipsets, click the button below.

View for other chipsets

Technical Details

Input sequence length for Prompt Processor:128
Minimum QNN SDK version required:2.45.0
Model architecture:Transformer with SigLIP-style ViT Vision Encoder, Per-Layer Embeddings (PLE), interleaved sliding-window/global attention with Grouped Query Attention (GQA), and shared KV layers.
Number of parameters:2B (effective)
Precision:w4a16
Response Rate:Rate of response generation after the first response token.
Supported languages:Multilingual (trained on 140+ languages)
TTFT:Time To First Token is the time it takes to generate the first response token. This is expressed as a range because it varies based on the length of the prompt.
Use:Initiate conversation with prompt-processor and then token generator for subsequent iterations.

Applicable Scenarios

  • Dialogue
  • Content Generation

License

Tags

  • llm
  • vlm
  • generative-ai

Supported Compute Devices

  • Snapdragon X Elite CRD
  • Snapdragon X2 Elite CRD

Supported Compute Chipsets

  • Snapdragon® X Elite
  • Snapdragon® X2 Elite

Related Models

See all models

Looking for more? See models created by industry leaders.

Discover Model Makers