Qualcomm® AI HubAI Hub

Qwen3-VL-4B-Instruct

Multimodal 4B vision‑language model with enhanced visual reasoning capabilities.

Qwen3‑VL is a vision‑language model from Alibaba Cloud capable of understanding both text and images for multimodal reasoning tasks such as visual question answering and image captioning.

Not supported

This model is currently not supported on any Mobile chipset.

To see performance metrics for this model on other chipsets, click the button below.

View for other chipsets

Quick Start

1

Install Windows CLI App

2

Run the Model

Paste into CLI and run the following code.

For application and server integration, see Docs.

Technical Details

Model architecture:Transformer with ViT Vision Encoder, Grouped Query Attention (GQA), and SwiGLU activation.
Response Rate:Rate of response generation after the first response token.
Supported languages:100+ languages and dialects
TTFT:Time To First Token is the time it takes to generate the first response token. This is expressed as a range because it varies based on the length of the prompt.

Applicable Scenarios

  • Dialogue
  • Content Generation

License

Tags

  • llm
  • vlm
  • generative-ai

Supported Mobile Devices

  • Samsung Galaxy S21
  • Samsung Galaxy S21 Ultra
  • Samsung Galaxy S22 5G
  • Samsung Galaxy S22 Ultra 5G
  • Samsung Galaxy S22+ 5G
  • Samsung Galaxy S23
  • Samsung Galaxy S23 Ultra
  • Samsung Galaxy S23+
  • Samsung Galaxy S24
  • Samsung Galaxy S24 Ultra
  • Samsung Galaxy S24+
  • Samsung Galaxy S25
  • Samsung Galaxy S25 Ultra
  • Samsung Galaxy S25+
  • Samsung Galaxy S26
  • Samsung Galaxy S26 Ultra
  • Samsung Galaxy S26+
  • Samsung Galaxy Tab S8
  • Xiaomi 12

Supported Mobile Chipsets

  • Snapdragon® 8 Elite For Galaxy Mobile
  • Snapdragon® 8 Elite Gen 5 For Galaxy Mobile
  • Snapdragon® 8 Gen 1 Mobile
  • Snapdragon® 8 Gen 2 Mobile
  • Snapdragon® 8 Gen 3 Mobile
  • Snapdragon® 888 Mobile

Related Models

See all models

Looking for more? See models created by industry leaders.

Discover Model Makers