Worlds
Six synthetic web apps agents are benchmarked against. Start a session to explore one by hand — with or without chaos.
Yuuki Shop
Search, filter and compare products; cart and checkout without ever taking a real payment.
7 benchmark tasks
Yuuki Travel
Search flights and hotels; assemble and review a trip without ever confirming a booking.
6 benchmark tasks
Yuuki Mail
A fake inbox — search, read, label, archive and draft. Nothing is ever actually sent.
6 benchmark tasks
Yuuki CRM
Customers, notes and status changes over a data table and detail records.
6 benchmark tasks
Yuuki Docs
Documents and folders with a search/tag/move/edit workflow.
6 benchmark tasks
Yuuki Cloud
A fake deployment platform — projects, env vars, deployments and a build log to diagnose.
6 benchmark tasks
This is a hand-driven session, not a run
Exploring a world creates a real, isolated session the same way a benchmark run does — but nothing here is scored or recorded as a run. Actual agent runs are produced by the CLI (pnpm lab run), which drives a real Playwright browser against this app; see the Overview page for that command.