CODE

Qwen 3.8 Max Review: 2.4T Flagship, Token Plan from $6/mo (Tested)

⭐ 4.4/5 💰 Token Plan from 39 CNY/mo (~$6) · API: ~$2 in / $6 out per 1M tokens
qwen 3.8 maxalibabatoken planopen sourceai image generation3d generationgame generationagentic coding
Tool
Qwen3.8-Max
Pricing
Token Plan from 39 CNY/mo (~$6) · API: ~$2 in / $6 out per 1M tokens
✅ Pros
  • 2.4T parameters, 1M context, and true multimodal input (image / text / video)
  • International API priced at 40% (input) and 24% (output) of Opus 5
  • Token Plan from ~$6/month with night rates as low as 0.2x; one key for 150+ models
  • Frontend Code Arena score of 1,668 — #4 overall, just 1 point behind Claude Opus 5 (High)
  • Generated playable Angry Birds, Plants vs Zombies, Gold Miner, a macOS simulator and a 3D projector room
  • Weights go open source next week (plus a 27B variant) — first Max-level Qwen to be opened
❌ Cons
  • 3D scenes still lack depth: flat lighting, shallow glass panels, imperfect reference fidelity
  • Behind on some benchmarks: SWE-bench Pro 67.7 vs Fable 5's 80.0; VideoMME not top
  • Token Plan credits roll on 7-day and 5-hour windows — heavy usage needs the Pro tier
  • Weights not released yet; self-hosting cost can only be verified after next week's drop
  • Full capabilities run through API/subscription, not a standalone consumer app

TL;DR

On August 3, Alibaba officially launched Qwen3.8-Max: 2.4 trillion total parameters with 95B active, a 1M-token context window, and native image/text/video input. It shipped together with the upgraded Token Plan subscription — personal plans from 39 CNY/month (~$6) and night-time rates as low as 0.2x. Internationally, API input/output prices come in at 40% and 24% of Opus 5 respectively.

For this Qwen 3.8 Max review I spent two days hammering it: a 3D earth dashboard rebuilt from a reference image, a cloud-doodling image edit, an Angry Birds clone, Plants vs Zombies, Gold Miner, a macOS simulator, and a 3D projector room. Everything ran. The verdict: this is not spec-sheet inflation — it ships working products, especially in visual coding and game generation. 3D polish is the one place it still falls short.

Qwen 3.8 Max official promo: 2.4 trillion parameters

Quick Facts

MetricData
Release DateAugust 3, 2026 (official)
Parameters2.4T total / 95B active (sparse MoE + hybrid attention)
Context Window1,000,000 tokens (max input 991,808 / max output 131,072)
Input ModalitiesImage / Text / Video
Domestic API Price12 CNY in / 36 CNY out per 1M tokens · 1.5 CNY cached input
International API Price~$2 in / ~$6 out per 1M tokens (40% / 24% of Opus 5)
Token PlanPersonal from 39 CNY/mo · Team from 150 CNY/seat/mo
Frontend Code Arena1,668 points — #4 overall
PaperBench93.0 (+28.2 vs previous generation)
OSWorld-Verified86.1 — #1 among mainstream models
Open SourceWeights next week (Hugging Face + ModelScope), plus 27B variant

Why This Launch Matters

Model launches have been coming so fast that I’ve gone numb. Qwen3.8-Max is different for three concrete reasons:

1. It’s the first Max-level Qwen that will be open-sourced. Qwen has always kept its open models and its Max flagship on separate tracks. This time Alibaba confirmed the weights hit Hugging Face and ModelScope next week, alongside a Qwen3.8-27B. That puts a 2.4T-class base model within reach of anyone who wants to self-host or fine-tune.

2. The price gap is an order of magnitude, not a discount. Domestically it’s 12 CNY in / 36 CNY out per million tokens; internationally ~$2 / $6 — 40% and 24% of Opus 5’s international pricing. The upgraded Token Plan starts at 39 CNY/month, and between 22:00–08:00 the promo rate stacks to an effective 0.2x. For anyone running agents all day, that changes the math entirely.

Bottom line: the same complex agent task that costs a fortune on Opus-5-class APIs runs at a quarter of the price here — with the gap in quality shrinking.

If you haven’t read my full test of the previous-generation preview, start with our Qwen 3.8 review. The models I keep comparing against — Claude Opus 5 and Kimi K3 — each have their own deep-dive review too.

Token Plan Breakdown

Token Plan is Alibaba Cloud Bailian’s subscription tier, upgraded alongside the Qwen3.8 launch. The core idea: one API key covers 150+ models — including Qwen3.8-Max and the HappyHorse-1.1 video model — with a unified Credits balance across text, image and video work.

PlanPriceCredits (7 days)Credits (5 hours)Agent Concurrency
Lite39 CNY/mo (~$6)2,5007001–2
Standard139 CNY/mo10,0003,0003–4
Pro499 CNY/mo40,00012,0006–8
Team150 CNY/seat/moShared poolShared poolTeam-scaled

⚠️ Two gotchas before you buy: 1) Credits roll on dual 7-day / 5-hour windows — the Lite tier’s 700 credits per 5 hours can run dry mid-task; 2) the 0.2x night promo targets the personal preview offer, so check the fine print on Team plans and post-promo billing.

My advice: start with Lite and hammer it for a week. With day-time credits at 0.1x during the promo, 39 CNY effectively buys several times its nominal volume.

Test 1: Frontend & Visual Coding

Let’s start with the leaderboard. On Arena AI’s Frontend Code chart, Qwen3.8-Max scores 1,668 — ranked #4 overall, one point behind Claude Opus 5 (High) and just below Kimi K3 (Max) and Claude Opus 5 (Max). For a model that just went official, that’s a serious frontend pedigree.

Arena AI Frontend Code leaderboard: Qwen3.8-Max ranks #4 with a score of 1668

Then I ran my favorite frontend test: one prompt + one reference image, rebuild a 3D earth dashboard. It produced a working page on the first pass — layout, structure and features all present: a 3D globe, flight-path arcs, stat panels (Monthly Delivered 1,021 / Yearly 4,603) and an AI-agent sidebar.

Qwen3.8-Max test: 3D earth dashboard rebuilt from a single prompt and reference image

Honest caveats: textures are fine, lighting is flat, glass panels have no depth, and some elements drift from the reference. It’s a “runs on the first try, structurally complete” kind of model — but on pixel-level fidelity it’s not at Fable 5’s level yet. That matches what I saw in the 3D tests: usable, one breath short of polished.

Test 2: Image Editing

For image work I ran a “cloud doodle” comparison: the same blue-sky photo, four models asked to turn the clouds into animals. Qwen3.8-Max drew a cat — pink inner ears, dot eyes, whiskers, a tiny mouth, plus a yellow pencil left mid-stroke. Kimi K3 drew a dachshund, GPT-5.6 Sol sketched a shaky elephant, and Claude Opus 5 drew a lion.

My subjective ranking: Qwen3.8-Max and Claude Opus 5 finished highest — and Qwen’s cat is notably clean, with no stray lines. For instruction-following in image editing, that’s a very solid debut for a freshly official flagship.

Qwen3.8-Max image editing test: cloud doodle comparison — Qwen3.8-Max draws a cat, Kimi K3 a dog, GPT-5.6 Sol an elephant, Claude Opus 5 a lion

Test 3: 3D & Game Generation

This is the part that genuinely surprised me. Everyone keeps saying “Qwen’s 3D is weak” — so I threw classic games and interactive apps at it. It produced five working products in one sitting.

Angry Birds Clone

A single “make an Angry Birds-style game” prompt produced a complete playable level: LEVEL 5 – PIGGY PALACE, a running score of 24,015, slingshot, trajectory guide, green pigs, wooden forts and three birds left to fire. It’s not a mockup — it actually plays, interacts and scores.

Plants vs Zombies

Plants vs Zombies is even more impressive: 400 sun points, a full plant-card bar (Sunflower, Peashooter, Snow Pea, Wall-nut, Cherry Bomb…), a 5-lane lawn, five lawnmowers, and a zombie already advancing — with the wave meter reading 10/10. The loop of selecting cards, planting and wave progression all runs in the browser.

Gold Miner

For Gold Miner it built a “Trading Post” shop: Dynamite (clears obstacles), Potion (+20% permanent speed), Clover and a Watch — each with descriptions, star ratings and prices. You hold $1,439 and the UI reminds you “next level target: $1,400 — money spent doesn’t count, spend wisely.” It even writes the economy strategy text for you.

Qwen3.8-Max Gold Miner test: Trading Post shop with four items and prices

macOS Simulator

It also built a web-based macOS simulator: menu bar (Notes, File, Edit, Go, Window, Help), a full Dock (Launchpad, Notes, Safari, Mail, Photos, Calendar, Trash), desktop wallpaper and a Finder window. The Notes app even says “this is a macOS Notes app implemented in a simulator — windows can be dragged, minimized and closed; Dock icons magnify on hover; right-click the desktop to change wallpaper.” It documents its own features.

Qwen3.8-Max macOS simulator test: menu bar, Finder, Notes and Dock

3D Projector Room

The headline test. Everyone says Qwen3.8 Max can’t do 3D — so I made it build a 3D projector music room: a four-tier vinyl shelf (every record with a gold label), a big central screen with a retro radio, and plants on both sides in a dark, immersive space. It even supports locally uploading PDF, MP4 or PPT files to beam onto the screen. The animation, interaction and detail work are genuinely good — not “can’t do 3D,” more like “can definitely do 3D.”

Qwen3.8-Max 3D test: projector room playing an Apple Business slide deck

Qwen3.8-Max 3D projector room: clicking a record shows its title and year

Test 4: Long-Horizon Agents

Beyond the visible generation demos, the real story is long-horizon agent work. Alibaba published several cases I found genuinely hard to believe:

  • 16 days of autonomous coding: starting from an empty folder, it ran ~16 days with no human hand-holding — 265 commits, 127 PRs, 151 issues — and shipped the self-evolving agent framework “oh-my-cli”;
  • Paper reproduction: ~125 hours and ~7,600 lines of code to reproduce a research paper, then proposed 18 improvements and gained another +2.7 on AIME 2024;
  • Chip design: ~500 interaction rounds cut a hardware accelerator from 8,298 logic gates to 678, shrinking area by 81%;
  • Quant research: from 6 factor descriptions, it orchestrated ~330 sub-agents and ran ~6,000 backtests.

I can’t independently verify those numbers, but they line up with my own experience of getting playable games from single prompts. The direction is consistent: this model finishes work instead of just answering.

Worth reading alongside: the newly open-sourced Kimi K3, and the free options Grok 4.5 and Gemini 3.6 Flash — each plays a different game.

How to Get Started

1. Subscribe to Token Plan

Open the Bailian console (or the Qwen AI platform), pick Token Plan, choose Lite / Standard / Pro, and grab your API key. During the promo, the night window offers an effective 0.2x rate — a good time to stock up.

2. Call the API

Use DashScope’s OpenAI-compatible endpoint with the model id qwen3.8-max (confirm in the console):

curl https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions \
  -H "Authorization: Bearer $DASHSCOPE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-max",
    "messages": [{"role": "user", "content": "Build a 3D projector with Three.js that can play an uploaded video"}],
    "max_tokens": 2048
  }'

3. Want a free taste?

Qoder and QoderWork already support Qwen3.8-Max, and the Qwen PC app offers free chat. For the full picture — image editing, 3D, long-horizon agents — grab a Token Plan.

FAQ

What is Qwen 3.8 Max?

Qwen3.8-Max is Alibaba’s flagship large model officially released on August 3, 2026. It has 2.4 trillion total parameters (95B active), a 1M-token context window, and accepts image, text and video input. Weights are scheduled to open next week on Hugging Face and ModelScope.

How much does Qwen 3.8 Max cost?

Domestic API pricing is 12 CNY per million input tokens and 36 CNY per million output tokens (1.5 CNY cached input). Internationally it is about $2 per million input and $6 per million output — roughly 40% and 24% of Opus 5’s international pricing. Token Plan subscriptions start at 39 CNY (~$6) per month.

What is the Qwen Token Plan?

Token Plan is Alibaba Cloud Bailian’s subscription tier. One API key covers 150+ models including Qwen3.8-Max and the HappyHorse-1.1 video model, with unified Credits across text, image and video tasks. Personal plans are 39/139/499 CNY per month, and night-time promo rates can drop to an effective 0.2x.

When will Qwen 3.8 Max be open source?

Alibaba confirmed the Qwen3.8-Max weights will be released the week after the August 3 launch, alongside a Qwen3.8-27B variant, on Hugging Face and ModelScope. It will be the first Max-level Qwen model to be open-sourced.

Can Qwen 3.8 Max generate games and 3D scenes?

Yes. In our hands-on tests it generated playable Angry Birds, Plants vs Zombies and Gold Miner-style games, a macOS simulator, and a 3D projector room with interactive controls, all from single prompts.

How does Qwen 3.8 Max compare with Claude Opus 5?

On the Frontend Code Arena it scores 1,668 (ranked #4, one point behind Claude Opus 5 High) and PaperBench 93.0, while international API pricing is 40%/24% of Opus 5. It still trails on SWE-bench Pro and long-video understanding benchmarks.

Verdict

CategoryScoreComment
Performance / Benchmarks4.6 / 5PaperBench 93.0, Frontend Code #4 — closing in on closed flagships
Price / Value4.9 / 540% / 24% of Opus 5 pricing; Token Plan from ~$6/mo
Ease of Use4.0 / 5Low API barrier, but no standalone consumer app
Multimodal / Generation4.3 / 5Image, games and 3D all ship; 3D polish still short
Ecosystem / Open Source4.5 / 5Open weights next week + 27B; Qwen-MM-Plugins for agents

One-line verdict: Qwen3.8-Max is the most “gets things done” domestic flagship this year — strong on frontend leaderboards, impressive at image editing, fun with game generation, and capable of shipping real 3D scenes, all at a price that feels like a typo. When the weights drop next week, it will likely become the self-host crowd’s new default. Just don’t expect pixel-perfect 3D or a perfect score on every benchmark yet.

If you’re on a budget but want the strongest Qwen: start with the 39 CNY Lite tier, run your heavy tasks after 22:00, and it honestly feels close to free.

Sources

All facts and figures in this review were cross-checked against official public materials:

🛡 How We Test: This review is based on hands-on testing. We independently purchase subscriptions and do not accept payment for reviews. Updated August 4, 2026
👤
About the Author — Frankie

Reviews are based on hands-on testing with real prompts and tasks. We pay for our own subscriptions. Learn about our methodology.