Intern3-5-VL-2B
Multimodal 2B vision‑language model for text and image understanding.
InternVL3.5 is a vision‑language model from OpenGVLab capable of understanding both text and images for multimodal reasoning tasks such as visual question answering and image captioning.
Not supported
This model is currently not supported on any Automotive chipset.
To see performance metrics for this model on other chipsets, click the button below.
View for other chipsetsQuick Start
1
Install Windows CLI App
2
Run the Model
Paste into CLI and run the following code.
For application and server integration, see Docs.
Technical Details
Model architecture:Multimodal Transformer with InternViT Vision Encoder, Qwen-based Large Language Model (LLM) backend, Grouped Query Attention (GQA), and SwiGLU activation. It follows a ViT → MLP → LLM design for efficient vision-language understanding and reasoning.
Response Rate:Rate of response generation after the first response token.
Supported languages:No exact mention
TTFT:Time To First Token is the time it takes to generate the first response token. This is expressed as a range because it varies based on the length of the prompt.
Applicable Scenarios
- Dialogue
- Content Generation
License
Model:APACHE-2.0
Terms of Use:Qualcomm® Generative AI usage and limitations
Tags
- llm
- vlm
- generative-ai
Supported Automotive Devices
- SA8255P ADP
- SA8650P ADP
- SA8775P ADP
Supported Automotive Chipsets
- Qualcomm® SA8255P
- Qualcomm® SA8650P
- Qualcomm® SA8775P
Related Models
See all modelsLooking for more? See models created by industry leaders.
Discover Model Makers










