Qwen3 VL 235B A22B Instruct
Qwen's 235B-parameter MoE vision-language model for multimodal text and image understanding.
Pricing
per 1M tokensInput$0.21 /1M
Output$1.9 /1M
Cache read$0.1 /1M
Model Description
Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. The Instruct model targets general vision-language use (VQA, document parsing, chart/table...
Get the full context.
Sign up to read complete case studies, access detailed metrics, and unlock all use cases.
Get the full context.
Sign up to read complete case studies, access detailed metrics, and unlock all use cases.