AI

Alibaba trained an AI on 100 real phones. It says the result outclicks GPT‑5.6 and Claude.

Qwen-UI-Agent runs phones, desktops and browsers the way a person does — look at the screen, tap, type. On Alibaba’s own benchmarks it beats every Western flagship. The important word there is “own.”

N Noah · The Sharp Brief · August 24, 2026 · 3 min read

Alibaba’s Qwen team released Qwen-UI-Agent on Thursday: a single foundation model that operates software the way you do. It looks at the screen, finds the button, and taps it. Phones, desktops, browsers and deep-search workflows — one model, no APIs, no integrations, no permission needed from the app it’s driving. English-language coverage caught up over the weekend, and the benchmark table is the reason.

Alibaba’s technical report puts the model at 92.2% on MobileWorld-Real, a benchmark of 400-plus tasks across more than 100 real apps — ahead of Google’s Gemini 3.1 Pro, Anthropic’s Claude Opus 4.8 and OpenAI’s GPT‑5.6 on the same tests. It reports 79.5% on OSWorld-Verified (desktop), 73.6% on WebArena (browser) and 97.5% on AndroidDaily.

The detail that matters more than any single score is how it was trained: against a live farm of more than 100 physical smartphones running 150-plus real apps, pop-ups, lag, notifications and all. Most rivals train screen agents in emulators. Real devices are where screen agents die — which is exactly why Alibaba built its moat there.

Every screen just became an API

Most of the world’s software has no API. Decades of legacy enterprise systems, mobile-only apps, government portals, hospital front-ends — the entire robotic-process-automation industry exists because of that gap. A GUI agent closes it by behaving like an employee: whatever a human can click, it can click. That’s a different proposition from last week’s news that AWS and Binance gave agents wallets — payments gave agents a way to spend; screen control gives them somewhere to work.

Alibaba also built in a familiar guardrail: per Pandaily’s coverage, the agent stops and asks for confirmation before sensitive steps like payments. Whether humans keep reading those confirmations is another matter — developers already approve 97% of Claude Code’s permission prompts without much thought.

Our take: Alibaba built MobileWorld-Real and then aced it — that’s a company grading its own homework, so treat the table as claims until public leaderboards replicate them. But the direction is unambiguous. The predecessor MAI-UI family is open-weights on Hugging Face under Apache 2.0, from 2B to 235B parameters, and its largest model already tops the AndroidWorld leaderboard for pure-vision agents. While US labs meter computer-use agents as a paid service, China’s AI bloc keeps open-sourcing the layer underneath. If part of your business assumes “a human must be doing the clicking” — pricing pages, loyalty programs, booking queues — that assumption now has a shelf life.

What to watch

The screen was the last interface software couldn’t cross without an invitation. That just ended.

Advertisement

Get the day, decoded — at 7 PM ET

The Sharp Brief: AI, money, business & performance in five sharp minutes. Free.

Free bonus: subscribe today and The 2026 AI Playbook lands with your welcome email.

Recommended by 5+ newsletters across AI, markets & business.