"Today, we are officially releasing Qwen3.8-Max, the most capable model in the Qwen family to date. This also marks the first time we will open-source the weights of a Qwen-Max-class model," the Qwen team said.
According to Alibaba, Qwen3.8-Max can process large collections of documents, extensive code repositories, lengthy videos, and other workloads that require sustained reasoning across many steps. QwenCloud also supports OpenAI- and Anthropic-compatible API interfaces, allowing developers to integrate the model with tools including Claude Code, Codex, Qoder CLI, Qwen Code, and OpenClaw.
The company showcased several autonomous demonstrations intended to highlight the model's ability to complete complex projects with minimal human involvement. In one software engineering example, Qwen3.8-Max started with an empty repository and spent more than 10 days building the "oh-my-cli" project. Alibaba said the model created an engineering system that converted user and developer feedback into issues, assigned work to agents, generated code, executed tests, and merged approved changes.
After roughly 16 days of continuous operation, the repository contained 265 commits, 127 pull requests, and 151 issues while incorporating automated build, testing, monitoring, and recovery systems.
Alibaba also evaluated the model on academic research tasks. Starting with only a published paper describing a method for selecting training data, Qwen3.8-Max reportedly recreated the study without sample code before testing additional ideas over approximately 125 hours. The company said the model wrote around 7,600 lines of code, completed more than 1,100 actions, ran 33 rounds of GPU training, and ultimately improved performance on the AIME24 mathematics benchmark by 2.7 percentage points compared with the paper's original methodology.
Another demonstration entered Qwen3.8-Max into Alibaba Cloud's WWW2025 Multimodal Dialogue Intent Recognition Challenge alongside 526 human teams. Operating without human assistance during a 24-hour competition, the model assembled a system combining multiple language, vision-language, and image models. Alibaba reported that the system improved its accuracy from 0.60 to 0.853 across 45 submissions, finishing ahead of 458 teams.
The company also evaluated Qwen3.8-Max across a range of professional workflows. In one compliance exercise, Alibaba said the model identified 1,284 relevant clauses from hundreds of corporate documents in under an hour. Other demonstrations included building an interactive digital banking prototype, designing a restaurant menu from supplier documentation, reconstructing a seismic model for a 30-story office building, and analyzing approximately 8,400 basketball possessions per player to generate coaching reports.
For quantitative finance, Alibaba tasked the model with creating an exchange-traded fund rotation strategy from a one-line description. The system built its own data pipeline, generated investment factors, evaluated backtests, and adjusted its methodology after identifying potential overfitting. In a larger experiment, Qwen3.8-Max divided six investment themes into 50 research directions each, deployed approximately 330 sub-agents, and completed about 6,000 backtests within a single conversation.
Alibaba also tested the model on semiconductor design. Beginning with basic instructions and empty hardware templates, Qwen3.8-Max completed logic design, RTL code generation, simulation, synthesis, and physical layout work for a cryptographic hardware accelerator. Across approximately 500 interaction rounds and 71 evaluations, Alibaba said the design shrank from 8,298 gates to 678 gates while reducing chip area by 81%, shortening wire length from 33,369 micrometers to 4,187 micrometers, and achieving timing closure at 500 megahertz.
In a simulated year-long e-commerce environment featuring 12 store types, 60 product categories, nearly 600 suppliers, and 7,000 products, the model managed inventory, pricing, supplier negotiations, returns, and capital allocation starting with ¥100,000. Alibaba said the simulation ended with a balance of ¥416,252, representing a 4.16-times return that exceeded both the second-place model and the previous Qwen3.7-Max generation.
Beyond text generation, Qwen3.8-Max supports multimodal reasoning across images, documents, and video. Alibaba said the model can analyze financial reports and PDF documents exceeding 200 pages, organize more than 100 hours of video into searchable memory structures, and generate outputs including websites, animations, three-dimensional visualizations, and interactive applications. The company also said the model can inspect its own visual outputs, identify layout or spatial errors, and revise its work through an automated feedback process.
Developers can choose among three reasoning modes, with the default "xhigh" setting targeting the most demanding tasks, while "medium" balances speed and quality and "low" prioritizes lower-cost inference.
Alibaba also published benchmark results for Qwen3.8-Max, including scores of 93.0 on PaperBench, 81.9 on WideSearch, 82.8 on IFBench, 60.2 on HealthBench, and 58.3 on PRBench-Finance. Multimodal benchmarks included 82.3 on MMMU-Pro, 86.1 on OSWorld-Verified, 92.1 on OmniDocBench 1.5, and 90.4 on VideoMME with subtitles. The company noted that some evaluations were conducted internally and that testing methodologies differed across benchmarks and competing models.
"Built upon the architectural foundation of Qwen3.5, Qwen3.8-Max scales to 2.4 trillion parameters, delivering comprehensive improvements across coding, work, research and long-horizon tasks. It can not only answer more challenging questions, but also complete complex tasks end-to-end with greater reliability, producing dependable deliverables," the Qwen team said.
This analysis is based on reporting from Pulse 2.0.
Image courtesy of Neowin.
This article was generated with AI assistance and reviewed for accuracy and quality.