Worlds

Six synthetic web apps agents are benchmarked against. Start a session to explore one by hand — with or without chaos.

Yuuki Shop

Search, filter and compare products; cart and checkout without ever taking a real payment.

7 benchmark tasks

Yuuki Travel

Search flights and hotels; assemble and review a trip without ever confirming a booking.

6 benchmark tasks

Yuuki Mail

A fake inbox — search, read, label, archive and draft. Nothing is ever actually sent.

6 benchmark tasks

Yuuki CRM

Customers, notes and status changes over a data table and detail records.

6 benchmark tasks

Yuuki Docs

Documents and folders with a search/tag/move/edit workflow.

6 benchmark tasks

Yuuki Cloud

A fake deployment platform — projects, env vars, deployments and a build log to diagnose.

6 benchmark tasks

This is a hand-driven session, not a run

Exploring a world creates a real, isolated session the same way a benchmark run does — but nothing here is scored or recorded as a run. Actual agent runs are produced by the CLI (pnpm lab run), which drives a real Playwright browser against this app; see the Overview page for that command.