QVQ by Qwen

QVQ by Qwen

by Qwen team at Alibaba Cloud

Open-weight image-text reasoning model, now legacy.

Reviewed September 2026by CBAI Editorial Team
5.5/10

Score

2.8 out of 5 · our score
Compare

Our verdict

QVQ by Qwen is an open-weight, 72B-parameter visual reasoning research preview built on Qwen2‑VL‑72B. It accepts image and text inputs and returns text-only outputs, with a strict single-round limit and no video support. The model posted strong research benchmarks (e.g., MMMU 70.3, MathVista-mini 71.4) but comes with vendor-documented failure modes, including language mixing, recursive reasoning loops, loss of image focus during long chains, and a need for enhanced safety. There is no subscription plan; weights are free under the custom Qwen licence, but full-precision inference demands roughly 145 GB+ VRAM and multi-GPU servers. Hosted visual reasoning access today is via Alibaba Cloud Model Studio on newer models (QVQ-Max, Qwen3-VL) billed per token. As of 2026, QVQ-72B-Preview is clearly legacy—useful as a research baseline or for controlled experiments, not for production conversational assistants or video tasks.

Overview

QVQ by Qwen is an open-weight visual reasoning research preview from the Qwen team at Alibaba Cloud. Built on Qwen2-VL-72B, it performs step-by-step reasoning over images with text input/output only. It supports single-round dialogue only and does not accept video input. Released Dec 25, 2024, QVQ-72B-Preview has been superseded by QVQ-Max (Mar 28, 2025) and the Qwen3-VL Instruct/Thinking family (Sep 23, 2025).

Score breakdown

Overall score

5.5/10
Output quality7.0/10
Ease of use5.5/10
Value for money5.5/10
Features & tools5.5/10
API & integrations4.5/10
Support & docs5.5/10

Scores are editorial assessments by the Compare Best AI team on a 0–10 scale.

Expert review

C

CBAI Editorial Team

Compare Best AI · Editorial Team

## Overview

How we tested

Days tested

7 days

Tasks evaluated

  • ·Solve multi-step visual math problems from MathVista-mini
  • ·Answer diagram-based science questions with chain-of-thought prompts
  • ·Stress-test long reasoning chains for focus drift and loops
  • ·Quantized (8-bit) vs bf16 inference latency/quality comparison

Method

Compared against QVQ-Max and Qwen3-VL (Instruct/Thinking) using identical inputs

Reviewer

CBAI Editorial Team

Plans & pricing

Most Popular

Self-host

$0

Local or self-hosted use

  • No public self-serve SaaS price table on the official page

Pricing may vary by region. Always verify on the vendor's website.

Feature comparison

FeatureQVQ by QwenQVQ-MaxQwen3-VL (Instruct/Thinking)
Status
Released 2024-12-25
Released 2025-03-28 or later
Current as of 2026-08
Access
Available via Alibaba Cloud Model Studio
I/O
Single-round dialogue only
No video input supported
Included Partial / add-on Not included

Is it right for you?

Good fit for

AI researchers

Study long-chain visual reasoning behavior, limitations, and failure modes in a large open-weight VL model.

Benchmarkers

Reproduce and compare MMMU, MathVista, MathVision, and OlympiadBench results on a known research baseline.

GPU-rich R&D teams

Prototype self-hosted multimodal reasoning with control over inference stack (Transformers/vLLM) on multi-GPU servers.

Model engineers

Experiment with quantization (AWQ/GPTQ/MLX 8-bit) and serving performance trade-offs.

Less suited for

Production conversational apps

Single-turn only, research-preview stability, and safety caveats make it unsuitable for production chat.

Video understanding

The model accepts images and text only; video input is not supported.

Teams without datacenter GPUs

Full-precision inference needs ~145 GB+ VRAM and multi-GPU hardware; hosted access is per-token on other models.

User reviews

2.8

Editorial score

5
33%
4
19%
3
14%
2
9%
1
4%

Distribution is estimated from our editorial score. Verified user reviews coming soon.

Use cases

Typical ways teams rely on this tool — from everyday tasks to specialized workflows.

  • Visual math reasoning
  • Scientific diagram question answering
  • Benchmark reproduction (MMMU, MathVista, MathVision)
  • Olympiad-style problem analysis
  • Long-chain visual reasoning studies

Integrations

Reported connectors

Apps and services commonly connected out of the box or via official connectors.

  • Hugging Face Transformers
  • vLLM
  • ModelScope

Details

Category

Developer

Price

  • Free to self-host / open weights
  • hosted usage billed by the cloud vendor if used

Free version

No

Best for

  • Academic benchmark replication and analysis
  • Olympiad-style visual math/science problem solving
  • Safety and hallucination behavior studies
  • Quantization and inference optimization experiments

Frequently asked questions

QVQ by Qwen

QVQ by Qwen

Developer

Ready to get started?

Visit the QVQ by Qwen website to explore plans and pricing.

Was this page helpful?

Compare Best AI may earn a commission when you click links on this page. This does not influence our editorial scores or recommendations. Advertiser disclosure