Qwen 1.5

Qwen 1.5

by Alibaba Cloud (Qwen team)

Retired, open-weights text-only LLM series.

Reviewed September 2026by CBAI Editorial Team
5.5/10

Score

2.8 out of 5 · our score
Compare

Our verdict

Qwen 1.5 is an open-weights, text-only LLM family from Alibaba Cloud that landed early 2024, spanning 0.5B–110B plus an MoE A2.7B. It delivers solid multilingual, code, and math performance for its generation and is easy to run locally thanks to first-class support in Transformers, vLLM, SGLang, llama.cpp, Ollama, and LM Studio, alongside official GPTQ/AWQ/GGUF quantizations. The big caveat in 2026 is lifecycle: Alibaba retired all hosted Qwen 1.5 endpoints in Aug 2025, and Qwen itself calls 1.5 the beta of Qwen2. Practically, that means there’s no vendor SLA or security updates, and only self-hosting remains. Licensing also demands care: dense models use the custom Tongyi Qianwen license (not OSI-approved), with Apache-2.0 only for the MoE A2.7B variant. If you need a managed API or modern features like longer contexts or multimodality, look to qwen-plus or newer Qwen2/2.5/3 models; otherwise, Qwen 1.5 still works for cost-free, local text generation and fine-tuning.

Overview

Qwen 1.5 is an open-weights, decoder-only text LLM series from the Qwen team at Alibaba Cloud. Released Feb 4, 2024 in nine sizes (0.5B–110B) plus a Qwen1.5-MoE-A2.7B variant, it offers a uniform 32,768-token context and strong multilingual, coding, and math performance for its era. It integrates cleanly with Hugging Face Transformers, vLLM, SGLang, llama.cpp, Ollama, and LM Studio, with vendor-provided quantized builds (GPTQ/AWQ/GGUF). All Alibaba Cloud hosted Qwen1.5 chat e…

Score breakdown

Overall score

5.5/10
Output quality7.0/10
Ease of use6.5/10
Value for money8.0/10
Features & tools7.0/10
API & integrations1.0/10
Support & docs2.0/10

Scores are editorial assessments by the Compare Best AI team on a 0–10 scale.

Expert review

C

CBAI Editorial Team

Compare Best AI · Editorial Team

## Overview

How we tested

Days tested

7 days

Tasks evaluated

  • ·Local inference via Transformers and vLLM across 7B/14B
  • ·Quantized GGUF runs with llama.cpp and Ollama
  • ·Instruction-tuning a 7B variant using Axolotl
  • ·RAG evaluation with multilingual passages

Method

Compared against Qwen2.5-72B-Instruct and qwen-plus using identical inputs

Reviewer

CBAI Editorial Team

Plans & pricing

Most Popular

Self-host

$0

Local or self-hosted use

  • No public self-serve SaaS price table on the official page

Pricing may vary by region. Always verify on the vendor's website.

Feature comparison

FeatureQwen 1.5Qwen2.5-72B-Instructqwen-plus (hosted)
Distribution
Open weights available
Access
First-party hosted API
Modalities
Vision/audio support
Context
32k-token window
Licensing
OSI-approved license
Ecosystem
Official quantized releases (GPTQ/AWQ/GGUF)
Runtime
Works with Ollama/llama.cpp
Support
Vendor SLA/support
Included Partial / add-on Not included

Is it right for you?

Good fit for

Local-first developers

Run small to medium models on a single GPU or CPU via GGUF without vendor lock-in.

Research and benchmarking

Compare sizes from 0.5B to 110B with a uniform 32k context across tasks and languages.

Fine-tuners

Train instruction or domain variants using Axolotl/LLaMA-Factory and deploy via vLLM or Ollama.

Multilingual apps

Build text-only assistants spanning 12 languages without relying on external APIs.

Less suited for

Teams needing a managed API

Alibaba retired all Qwen1.5 hosted endpoints on Aug 20, 2025; only self-hosting remains.

Vision or audio projects

Qwen 1.5 is text-only; use separate Qwen-VL or audio models instead.

Strict open-source licensing

Dense models use the custom Tongyi Qianwen license, not OSI-approved.

User reviews

2.8

Editorial score

5
33%
4
19%
3
14%
2
9%
1
4%

Distribution is estimated from our editorial score. Verified user reviews coming soon.

Use cases

Typical ways teams rely on this tool — from everyday tasks to specialized workflows.

  • Self-hosted chatbots
  • Code generation and assistance
  • Multilingual text generation
  • RAG and document QA

Integrations

Reported connectors

Apps and services commonly connected out of the box or via official connectors.

  • Hugging Face Transformers
  • vLLM
  • Ollama

Details

Category

Developer

Price

  • Free to self-host / open weights
  • hosted usage billed by the cloud vendor if used

Free version

No

Best for

  • Local inference on consumer GPUs/CPUs (smaller sizes)
  • Fine-tuning via Axolotl or LLaMA-Factory
  • Multilingual chat and code assistants
  • RAG pipelines with vLLM/Ollama

Frequently asked questions

Qwen 1.5

Qwen 1.5

Developer

Ready to get started?

Visit the Qwen 1.5 website to explore plans and pricing.

Was this page helpful?

Compare Best AI may earn a commission when you click links on this page. This does not influence our editorial scores or recommendations. Advertiser disclosure