Qualcomm® AI HubAI Hub

Intern3-5-VL-2B

Multimodal 2B vision‑language model for text and image understanding.

InternVL3.5 is a vision‑language model from OpenGVLab capable of understanding both text and images for multimodal reasoning tasks such as visual question answering and image captioning.

Not supported

This model is currently not supported on any Compute chipset.

To see performance metrics for this model on other chipsets, click the button below.

View for other chipsets

Quick Start

1

Install Windows CLI App

2

Run the Model

Paste into CLI and run the following code.

For application and server integration, see Docs.

Technical Details

Model architecture:Multimodal Transformer with InternViT Vision Encoder, Qwen-based Large Language Model (LLM) backend, Grouped Query Attention (GQA), and SwiGLU activation. It follows a ViT → MLP → LLM design for efficient vision-language understanding and reasoning.
Response Rate:Rate of response generation after the first response token.
Supported languages:No exact mention
TTFT:Time To First Token is the time it takes to generate the first response token. This is expressed as a range because it varies based on the length of the prompt.

Applicable Scenarios

  • Dialogue
  • Content Generation

License

Tags

  • llm
  • vlm
  • generative-ai

Supported Compute Devices

  • Snapdragon X Elite CRD
  • Snapdragon X Plus 8-Core CRD
  • Snapdragon X2 Elite CRD

Supported Compute Chipsets

  • Snapdragon® X Elite
  • Snapdragon® X Plus 8-Core
  • Snapdragon® X2 Elite

Related Models

See all models

Looking for more? See models created by industry leaders.

Discover Model Makers