Model
Model explorer
openPangu-2.0-Flash (Non-Thinking)
OPENHuawei · openPangu family · released Jun 12, 2026
Default (no-extended-reasoning) inference mode of the same openPangu-2.0-Flash weights; this is the mode a user gets without opting into Thinking mode.
ReasoningCodingVisionFunction callingTool useAgentic
1603.9
Elo · rank #117
Parameters
92B
Active params
6B (MoE)
Context
512K tokens
Architecture
MoE with MLA + DSA/SWA hybrid attention (1:2 ratio), 4-stream mHC residual, multi-token prediction heads; trained on Huawei Ascend NPUs (no NVIDIA hardware)
License
OpenPangu Model License Agreement v2.0
Languages
—
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
best: GPT-5.2 · 100.0%
best: GPT-6 Astra · 96.0%
best: MiniMax M3 · 83.0%
best: Gemma 4 26B A4B · 98.5%
best: DeepSeek-V4-Pro (Think Max) · 93.5%
best: Trinity-Large-Thinking · 91.9%
best: Claude Opus 5 · 96.0%
Run it locally
VRAM @ Q4
—
VRAM @ FP16
—
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
—
Fine-tune it
Conditional / customQLoRA62.6 GB1× A100 80GB
LoRA196.0 GB2× H200 141GB
Full fine-tune1476.6 GB8× B200 192GB
QLoRA SFT on ~10k samples ≈ $4.22 (1× A100 80GB)
openPangu family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)