Qualcomm® AI HubAI Hub

Gemma-4-E2B-it

Lightweight multimodal model from Google DeepMind handling text and image input.

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on small models) and generating text output. This release includes open‑weights models in both pre‑trained and instruction‑tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

Not supported

This model is currently not supported on any Automotive chipset.

To see performance metrics for this model on other chipsets, click the button below.

View for other chipsets

Technical Details

Model architecture:Transformer with SigLIP-style ViT Vision Encoder, Per-Layer Embeddings (PLE), interleaved sliding-window/global attention with Grouped Query Attention (GQA), and shared KV layers.
Number of parameters:2B (effective)
Response Rate:Rate of response generation after the first response token.
Supported languages:Multilingual (trained on 140+ languages)
TTFT:Time To First Token is the time it takes to generate the first response token. This is expressed as a range because it varies based on the length of the prompt.

Applicable Scenarios

  • Dialogue
  • Content Generation

License

Tags

  • llm
  • vlm
  • generative-ai

Supported Automotive Chipsets

    Related Models

    See all models

    Looking for more? See models created by industry leaders.

    Discover Model Makers