梦兽编程
AI_SUITE

Qwen3.8-Max Released: 2.4T-Parameter Coding Model, Open Weights Next Week

Qwen3.8-Max is out: 2.4T total params (95B active), the first Max-class model to open its weights (HF + ModelScope next week). Autonomous coding case: 265 commits in 16 days. PaperBench 93.0 leads all models; SWE-bench Pro still trails Claude Fable 5. Includes Claude Code / Codex setup.

Qwen3.8-Max Released: 2.4T-Parameter Coding Model, Open Weights Next Week

On August 3, Qwen3.8-Max was officially released: 2.4 trillion total parameters (95B active) — the strongest model in the Qwen family, and the first Max-class model ever to open its weights, arriving next week on Hugging Face and ModelScope.

Official autonomous-coding results: the model ran by itself for 16 days, leaving 265 commits, 127 PRs, and 151 issues on GitHub; it reproduced a research paper and then beat it, gaining another 2.7 points on AIME24; in a Tianchi competition it beat 458 human teams in 24 hours.

Remember three numbers: 2.4T params / 265 autonomous commits / open weights next week.

quote-card


0. Why developers should stop and read this

For the past year, AI coding has been a closed-source arms race: GPT-5.6-Sol, Claude Opus 4.8, and Fable 5 traded blows, while the open-source camp had no Max-class model in play.

Qwen3.8-Max breaks a house convention, not just a benchmark: “flagship Max-class model = closed source.” The weights go open next week, which means the community can deploy, evaluate, and fine-tune it themselves — every number on the scoreboard will face public scrutiny for the first time.

For working developers the practical impact is immediate: the model is already live on the Qianwen platform via API, compatible with both OpenAI and Anthropic protocols — your existing Claude Code or Codex setup can switch to Qwen3.8-Max with two environment variables. One more option means one more negotiating lever.

1. The release, in one table

ItemDetail
Total params2.4T (95B active)
ArchitectureBased on Qwen 3.5, scaled up
Open weightsFirst Max-class open release, HF + ModelScope next week
APILive now on the Qianwen platform
ProtocolsOpenAI-compatible (chat/responses) + Anthropic-compatible
Reasoning levelsreasoning_effort: xhigh / medium / low
FocusCoding, office work, research, long-horizon tasks, multimodal agents

In one sentence: the model is usable today, and runnable by you next week. At the 2.4T scale, this is the first time the open-source camp gets both at once.

2. Coding evidence: three cases with “no human in the loop”

Instead of showing single benchmark numbers, the team let the model complete three open-ended tasks — each requiring it to write its own code, run it, and iterate, with zero human help.

Case 1: Ten-day autonomous coding, a living project

The model had to create the oh-my-cli project from scratch and build a self-evolving harness: user feedback, community practice, and its own test results all feed into the engineering loop; requirements become issues, agents pick them up, implement, test, and merge.

  • Task state machine: ready → leased → active, with automatic rollback and re-verification on anomalies
  • Self-testing loop: Build, Unit Test, E2E, Desktop Lifecycle
  • After ~16 days of fully autonomous operation (as of 2026-07-30): 265 commits, 127 PRs, 151 issues

The full trace is public on GitHub: qwen-code-dev-bot/oh-my-cli. You can read its daily commits and watch it find and fix its own problems.

Case 2: Reproduce a paper, then surpass it

The model received one paper — Unified Data Selection for LLM Reasoning — and a batch of GPUs. No starter code, no existing pipeline. It had to design and write the data scripts, training code, and evaluation programs from zero.

  • ~125 continuous hours, ~7,600 lines of code, 1,100+ operations, 33 GPU training rounds
  • First 37 hours: built the full pipeline and reproduced all six main conclusions (the paper’s data-selection method beats random selection by +7.7% on AIME24)
  • Then a self-evolution loop: hypothesize → code → GPU → analyze → retry. Four rounds, 18 improvement ideas, ending with another +2.7 points on AIME24

Notice this detail: it chose to continue after “reproducing” succeeded. Nobody prompted it — the model itself defined “done” as “surpassed.”

Case 3: 24 hours, beating 458 human teams

The WWW2025 multimodal dialogue intent recognition challenge on Aliyun Tianchi, with 526 human teams.

  • 24-hour limit; the model read the rules and built its own solution
  • Text side: fine-tuned and ensembled BERT / MacBERT / RoBERTa; screenshots: fine-tuned Qwen2.5-VL-7B with Chinese-CLIP as fallback
  • Fused into a weighted voting system, weights calibrated via cross-validation
  • 45 submissions, accuracy climbing from 0.60 to 0.853 — beating 458 teams (87% of the field)

Together the three cases say one thing: Qwen3.8-Max isn’t a “model that writes code” — it’s a “model that finishes the project.”

The AI completing three open-ended coding tasks on its own

3. Benchmarks: where it wins, where it loses

The team published a full benchmark table (their own harness, some runs via Claude Code). The highlights:

BenchmarkQwen3.8-MaxStrongest rivalVerdict
PaperBench93.0GPT-5.6 Sol 90.5Win (best overall)
AndroidBench75.1Fable 5 84.5Loss
Terminal Bench 2.186.6GPT-5.6 Sol 88.8Slight loss
SWE-bench Pro67.7Fable 5 80.0Loss (clear gap)
DeepSWE 1.156.6GPT-5.6 Sol 73.0Loss (clear gap)
OSWorld-Verified86.1Fable 5 85.0Win (best overall)
CoWorkBench74.8Fable 5 75.9Roughly tied
JobBench53.4Fable 5 57.4Slight loss
IFBench82.8GPT-5.6 Sol 72.7Win (best overall)
HiPhO90.0GPT-5.6 Sol 86.8Win
MLVU90.8GPT-5.6 Sol 87.6Win
VideoMME90.4GPT-5.6 Sol 89.5Win

Honest takeaways:

  1. Long-horizon research, multimodal understanding, and office-agent tasks: Qwen3.8-Max genuinely sits in the top tier, first on several benchmarks.
  2. Classic software-engineering benchmarks still lag: SWE-bench Pro 67.7 vs Fable 5’s 80.0, DeepSWE 56.6 vs 73.0 — “fix real-repo bugs” is not yet its strong suit.
  3. The harness is Qwen’s own (Terminal Bench runs via Claude Code), so there’s a home-field factor.

In the closed era, benchmark tables are ads; after open weights, benchmarks are exams. Next week the community deploys and re-tests — that’s when we’ll really know.

4. Office and long-horizon: the model clocks in

The team also showed real-workflow demos (official demonstrations — treat them as capability ceilings, not production promises):

  • Corporate compliance lawyer: hundreds of documents, flagged 1,284 relevant clauses in one pass, done within an hour — traditionally a paralegal team’s week
  • Quant research: 6 factor families expanded into 50 research directions, ~330 sub-agents, ~6,000 backtests; selected factors with excess Sharpe 0.64–1.48
  • Chip design: GCD/RSA crypto accelerator, 500 interaction rounds and 71 evaluations, gate count cut from 8,298 to 678; physical layout area from 106×106µm² to 46×46µm² (-81%), timing from -4.46ns violation to timing closure at 500MHz
  • E-commerce simulation (365 days): started with ¥100,000, ended with ¥416,252 cash (4.16x return), +38% over runner-up GLM 5.2, +152% over Qwen3.7-Max

The pattern: the model’s long-horizon autonomy is no longer a demo — it’s reproducible engineering. Still, these are ceiling showcases; production repeatability awaits community verification after the open release.

5. Use it today: a two-minute setup

The model is live on the Qianwen platform (platform.qianwenai.com), compatible with OpenAI and Anthropic protocols.

Claude Code

export ANTHROPIC_MODEL="qwen3.8-max"
export ANTHROPIC_SMALL_FAST_MODEL="qwen3.8-max"
export ANTHROPIC_BASE_URL=https://dashscope.aliyuncs.com/apps/anthropic
export ANTHROPIC_AUTH_TOKEN=your-qianwen-api-key
claude

Codex (Responses protocol)

Register the model in ~/.codex/model-catalog.local.json and add a provider in ~/.codex/config.toml (full JSON in the official docs), then:

export OPENAI_API_KEY=your-qianwen-api-key
codex

Give your existing tools a stronger model

Reasoning levels

The API supports reasoning_effort: xhigh (default, deep analysis), medium (balanced), low (fast, token-cheap). Start with medium if you want to control cost.

One reminder: billing is per the Qianwen platform’s price list — check it before you get carried away by the word “free.”

6. Summary: three actions

  1. Curious? Sign up for the Qianwen platform today and run qwen3.8-max on your own task. Don’t trust anyone’s review — including this one.
  2. Cost-sensitive? Wait for next week’s open weights; self-host and re-run SWE-bench-class tasks before switching your main model.
  3. Watching the trend? Track three numbers: community evaluation count in the week after open release, GitHub stars, and how fast other vendors follow with open Max-class models.

The open-source camp finally holds a card of the same rank as closed models in the AI coding game. Whether it plays well — the open release next week will tell.


🎯 Found this useful?

  • Follow 「梦兽编程」 for hands-on AI coding tool tests and efficiency methods — only things you can actually adopt
  • Want a deep-dive evaluation of a specific model? Comment below and I’ll run it.


Sources:

Frequently Asked Questions

What is Qwen3.8-Max, and how is it different from previous Qwen flagship models?

Qwen3.8-Max is the strongest model in the Qwen family to date: 2.4 trillion total parameters and 95B activated parameters. It breaks the convention that "flagship Max-level model = closed source" and becomes the first Max-level model in history to open-source its weights—released next week on Hugging Face and ModelScope, and its benchmark scores will be tested by the community for the first time.

How did Qwen3.8-Max actually perform in programming tests?

The official team had the model independently complete three tasks with no human help: create the oh-my-cli project and build its own self-evolving harness, run autonomously for about 16 days and leave 265 commits, 127 PRs, and 151 issues; reproduce a paper and then self-evolve, gaining another 2.7 points on AIME24; and in a Tianchi competition, defeat 458 human teams within 24 hours.

Where does Qwen3.8-Max win and lose in benchmark scores?

It ranks first in PaperBench 93.0, OSWorld-Verified 86.1, and IFBench 82.8, and is in the top tier for long-horizon research, multimodal, and office agents; but it lags in software engineering: SWE-bench Pro 67.7 vs Fable 5 80.0, and DeepSWE 56.6 vs 73.0.

How can I connect Qwen3.8-Max to Claude Code or Codex?

The model is available on the Qwen platform. For Claude Code: set ANTHROPIC_MODEL=qwen3.8-max, point BASE_URL to the DashScope Anthropic endpoint, and fill AUTH_TOKEN with your Qwen API Key. For Codex: configure provider and register the model in config.toml.

What should ordinary programmers do now?

Three suggestions: if you want to try it, register on the Qwen platform and run your own tasks with qwen3.8-max—don't trust anyone's benchmark, including this one; if you want to save money, wait for the open-sourced weights next week, self-host inference, run SWE-bench-like tasks, and compare before deciding on your primary model; if you want to judge the trend, watch the number of community evaluations, GitHub stars, and vendor adoption speed within one week after open-sourcing. Check the price list before integrating.