Model
Model explorer

Qwen3.8-27B

OPEN
Alibaba · Qwen3.8 family · released Aug 5, 2026

The small open-weight sibling of Qwen3.8-Max, and the one that actually shipped under a permissive licence: Apache 2.0 with no revenue gate, against Qwen3.8-Max's separate-agreement requirement past US$50M of model-as-a-service revenue. RELEASE DATE CONFLICT: BenchLM records August 5, 2026 as an exact-day confirmation against the Hugging Face card, while Capital & Compute reports the weights landing August 14 after Alibaba pointed at the week of August 10; the earlier, primary-source-linked date is used here. A genuinely strong small model — it leads Qwen3.7-Plus (397B/17B) on SWE-bench Pro (61.7 vs 57.6), DeepSWE 1.1 (42.2 vs 14.2), CoWorkBench and JobBench, and posts 84.3 on OSWorld-Verified — while sitting behind it on GPQA Diamond and HLE. Vision rows use the without-code-interpreter setting where Qwen reports both. No first-party per-token API rate is published for this checkpoint, so no pricing row is filed.

ReasoningCodingVisionFunction callingTool useAgentic
2506.1
Elo · rank #30
Parameters
27B
Active params
Undisclosed
Context
262K tokens
Architecture
27B dense causal LM with a vision encoder; 64 layers in a repeating 16 x (3 x (Gated DeltaNet → FFN) → 1 x (Gated Attention → FFN)) hybrid linear/full-attention stack, multi-token prediction, 262,144-token native context extensible to 1M
License
Apache 2.0
Languages
API price (in/out)
No hosted API
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
Agents' Last ExamAgents42.9%#10
best: GPT-6 Astra · 59.3%
AndroidWorldAgents81.9%#2
best: Qwen3.8-Flash-Next · 84.5%
CharXivVision83.7%#9
best: Qwen3.8-Flash-Next · 90.6%
DeepSWECoding42.2%#14
best: Muse Spark 1.3 · 75.4%
GPQA DiamondReasoning89.2%#35
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning30.8%#50
best: Claude Opus 5 · 64.7%
IFBenchReasoning79.5%#10
best: MiniMax M3 · 83.0%
JobBenchAgents33.4%#7
best: Claude Opus 5 · 65.7%
LiveCodeBenchCoding90.3%#6
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MathVisionVision90.0%#4
best: Seed 2.1 Pro · 92.6%
NL2RepoCoding42.3%#4
best: GLM-5.3 · 58.0%
OSWorld-VerifiedAgents84.3%#2
best: Claude Fable 5 · 85.0%
best: Claude Opus 5 · 59.4%
SWE-bench ProCoding61.7%#14
best: Claude Fable 5.1 · 81.2%
Terminal-Bench 2.0Coding73.0%#23
best: Gemini 3.8 Flash · 89.4%
WebArena-VerifiedAgents64.8%#1
best: this model · 64.8%
Run it locally
VRAM @ Q4
VRAM @ FP16
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
Fine-tune it
Permissive
QLoRA19.0 GB1× RTX 3090 24GB
LoRA58.2 GB1× A100 80GB
Full fine-tune434.0 GB4× H200 141GB
QLoRA SFT on ~10k samples ≈ $13.91 (1× RTX 3090 24GB)
API price weights · each benchmark row carries its own source badge (see methodology)