AI Pulse by Inblix

Alibaba’s Qwen3.8-Max autonomously ran a business for a year and beat 87% of human coders

The Decoder · Aug 3, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Alibaba’s Qwen3.8-Max autonomously ran a business for a year and beat 87% of human coders

Alibaba just dropped a model that doesn’t just answer prompts — it clocks in for a multi-day shift. Qwen3.8-Max, a 2.4-trillion-parameter behemoth with 95 billion active parameters per query, is designed for long-horizon autonomous work. The Qwen team is releasing the weights next week, marking the first time a Max-class model gets this open treatment.

Let’s skip the vague ‘agentic’ label and look at what it actually did. In one case study, the model spent 16 days single-handedly building a command-line tool called oh-my-cli, managing 265 commits and 127 pull requests without human intervention. Another test had it reproduce and then beat a research paper’s results on the AIME24 math benchmark by 2.7 points, iterating through 18 of its own ideas across 33 GPU training jobs. It’s a concrete demonstration of AI not just executing a plan, but formulating and testing its own hypotheses.

The most striking result is arguably the E-Commerce-Bench simulation. Operating on anonymized Taobao and Tmall data, Qwen3.8-Max ran multiple online stores for a simulated fiscal year, starting with 100,000 yuan. It negotiated with suppliers, adjusted pricing, managed returns, and navigated crises like supply chain disruptions — all while sniffing out 152 scammers hidden in the supplier pool. It quadrupled its money, finishing with 416,252 yuan, which is 38% more than the runner-up GLM 5.2 and over 2.5 times what its predecessor managed. The model made an aggressive early investment that paid off with a six-figure profit during the holiday season.

What’s genuinely new here isn’t just scale, but sustained autonomy. In a chip design task, the model started with a bloated 8,298-gate circuit and, over roughly 500 iterations, carved it down to just 678 gates — an 81% reduction in physical chip area after layout. The Qwen team noted it kept making deep structural changes even late in the process, rather than plateauing with surface-level tweaks. Pair that with a multimodal benchmark win against 458 out of 526 human teams in a WWW2025 competition, and you start to see a shift from AI as a tool you prompt to AI as a colleague you hand a multi-week project brief. The open-weight release next week will tell us if this is a real platform shift or just a very expensive demo.

💡 Key Takeaways

  1. Qwen3.8-Max quadrupled its starting capital to 416,252 yuan in a year-long e-commerce simulation, beating the next-best model by 38% by making aggressive early investments.
  2. The model autonomously managed a 16-day software project with 265 commits and zero human touch, demonstrating a shift from single-prompt answers to sustained, self-directed work.
  3. In a chip design test, the model reduced a cryptographic circuit from 8,298 gates to just 678 over 500 iterations, showing it can make deep structural improvements rather than just surface-level optimizations.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles