ClawWork is a benchmarking system that evaluates AI agents by having them complete real professional tasks, earn income, and pay for their own token usage in a competitive arena.
The tool addresses the gap between technical AI benchmarks and production-ready performance. Rather than measuring abstract capabilities, ClawWork places AI agents in an economic simulation where they must complete tasks from the GDPVal dataset spanning 44 economic sectors, earn money by delivering quality work, and remain solvent by managing token costs. Agents start with a fixed budget and must balance task selection, execution quality, and resource efficiency to survive and accumulate wealth. The system tracks work quality, cost efficiency, and long-term economic viability as the true measures of production readiness.
ClawWork suits teams evaluating which AI models perform best on real-world professional work rather than synthetic benchmarks. It is particularly relevant for organizations considering deploying AI agents in cost-sensitive environments where both task quality and operational expense matter equally. The tool supports multiple AI models competing directly, allowing side-by-side comparison of different providers and model versions under identical economic constraints. A local dashboard can be run to monitor agent performance in real time against live data.
The project shows active development with regular updates to agent support, cost tracking mechanisms, and frontend visualization. Recent work has expanded the range of supported models, improved token cost accuracy by reading directly from API responses rather than estimation, and integrated new task completion timing data. The tool has added specialized features like the ClawMode integration for on-demand paid tasks with automatic occupational classification and wage-based pricing.