microsoft/CUAWright

A simple SWE style browser+desktop agent framework that achieves SOTA results on long horizon web tasks.

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 1 hour ago
Added to GitGenius on October 6th, 2026
Created on April 8th, 2026
Open Issues & Pull Requests: 53 (+0)
GitHub issues: Enabled
Number of forks: 384
Total Stargazers: 6,045 (+0)
Total Subscribers: 19 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 13.7 hours
Mean response time: 3.7 days
90th percentile: 9.0 days
Tracked items: 12

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 15
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 88 days
Stale 30+ days: 11
Stale 90+ days: 8

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

CUAWright is a browser and desktop agent framework that turns coding models into autonomous agents for web and desktop automation tasks.

The tool addresses the challenge of enabling large language models to autonomously complete complex, long-horizon tasks in web browsers and desktop environments. It provides a terminal-based interface that gives coding models direct access to browser automation through Playwright and desktop control in Ubuntu VMs. Rather than requiring models to reason about individual UI elements, CUAWright lets them write and execute code to accomplish tasks, treating the browser or desktop as a programmable environment. The framework supports multiple LLM backends including OpenAI, Anthropic, and OpenRouter, and includes a Skill Factory feature that automatically distills solved tasks into reusable, parameterized code scripts that can run standalone without further model invocation.

Developers should choose this tool if they need to automate complex web or desktop workflows with language models and want to leverage code generation rather than pure UI interaction. It suits projects requiring long-horizon task completion, such as web scraping, form filling, or multi-step desktop operations. The framework is particularly valuable for teams building agent systems that benefit from reproducible, reusable code artifacts. For browser-specific work, the tool can be used as a plugin or skill within existing agent platforms like Claude Code, Codex, and Hermes Agent. The project explicitly supports reproducing benchmarks on OSWorld-V2 for desktop tasks and achieves state-of-the-art results on web task datasets.

The project maintains backward compatibility with its previous Webwright identity while expanding scope to desktop automation. Development has progressed through incremental capability additions including persistent step-by-step browsing, native command execution tools, task-to-UI rendering modes, and cross-platform agent integrations. The codebase remains minimal at approximately fifteen hundred lines of code, suggesting a focused design philosophy centered on essential functionality rather than feature accumulation.