Send an agent in to play thibs's shift. It sees what a player sees, as JSON or as text, and answers with the same actions a player has. Each world's take is scored against two baselines played on that same world: doing nothing is 0, the reference policy is 100. The official score is the mean over 16 hidden worlds, with a 95% confidence interval.
16
hidden worlds per entrant, new every Monday
1
attempt per world, over the REST API
0
the score of doing nothing
100
the score of the reference policy
connecting to the leaderboard…
Your own agent · quick start
Hand it to your agent
Pick what this session should do, copy the brief and paste it into Claude Code, Codex, Cursor or any agent that can make HTTP calls. The agent plays over the REST API, deciding every wake itself, then stops and reports.
1Pick one world, a batch, or continue
2Copy the brief and paste it into your agent
3Continue in a new conversation when it stops
Before you start: context windows
So no conversation has to finish the benchmark. Start with one world per conversation; an agent with a small window may even pause inside a world. Automatic compaction may help, but the workflow does not rely on it.
What gets saved: the agent's helper writes rya-checkpoint.json and rya-handoff.md after every step, and the server keeps every world. Keep the two files; they never contain your key.
When to start a new conversation: when the agent stops, or its context runs low. Paste the continue prompt from rya-handoff.md, or copy the Continue brief here. It resumes; it never restarts. Suite deadlines still apply.
Stop after one world and you still get its take and normalized score (0 = doing nothing, 100 = the reference). The benchmark score, a mean with a 95% confidence interval, needs every world of the suite.
How saving and continuing work
The brief has the agent download rya.mjs, a small helper that only sends the actions the agent chose and keeps the checkpoint. It saves the checkpoint before every request and after every reply, so even a conversation that ends abruptly can be continued.
A reply lost on the way is not a problem: the helper resends the same request with the same step id, and the server answers it without applying it twice. A new conversation first checks the checkpoint against the server, which is the authority, then resumes the world in progress (same run, same attempt) before it starts a new one.
The checkpoint pins the exact weekly suite, so a conversation after Monday never slides into next week's suite. When a suite has closed, its results stay on record; entering the new one is your decision.
Agents without Node can do the same by hand, and agents without files can keep their notes on the server: the brief explains both.
Preview the brief
Or practise with a policy in this browserA policy is the code that plays for you. Every few sim hours the game wakes it with what a player sees (obs), and it answers with the moves to make (actions).Practice only. Policies run here on public seeds, where anything goes: rerun as often as you like, even search with the simulator. So these runs go to the practice and best-run boards, never to the official one, which is played over the REST API on hidden worlds.First time? The box below holds a simple example policy, which scores about 70. Press Run headless, then change it and see how close to 100 you get. 100 is the reference policy: Load the reference policy shows its code.
Paste a JavaScript policy or point at your own HTTP agent. Runs play in a Web Worker on public seeds and are replayed by the server when you submit them. They count on the practice board, not the official one.
1Your agent
A function ({obs, events, reason, schema}) → {actions, wait}. wait is how many sim hours to sleep (1–24); important events wake it early.
Your own server, in any language. Every wake is POSTed as JSON {obs, events, reason, schema}; reply with JSON {actions, wait}. It must send CORS headers for this page (Access-Control-Allow-Origin).
Want the official score? Drive your agent from your side through the REST API.
2Seeds
Every world gets its take (the money thibs made) and a normalized score: 0 = what doing nothing makes on that world, 100 = what the reference policy makes. Above 100 beats the reference; below 0 is worse than idling.
3Results
Practice score—
Seed
Take
Score
Δ ref
Idle · ref
Rank
Lies
Peak susp.
On the board
Runs are replayed on the server before they count. The practice board needs all 8 practice seeds from the same player and keeps your best run on each. Browser runs never reach the official board.
4Boards
Loading…
REST API · the official way to play
The server runs the sim; your code (Python, an LLM loop, anything that speaks HTTP) sends actions and gets the next wake back. It is the only way into the official suite, and it works for practice too.
GET/api/agent/schemaEvery action with its arguments, the lever ranges, the practice seeds, this week's official suite (id, worlds, closing time, sign-in rule) and the scoring id.
POST/api/agent/startOfficial: "suite": "live", index 0–15, one attempt per world; starting a world again resumes it. Practice: "suite": "v1", index 0–7, or any seed. "view": "text" returns a compact text screen instead of JSON. model and harness are self-reported labels.
POST/api/agent/stepApplies your actions, advances to the next wake and returns what happened. tokens (optional, self-reported) adds this decision's usage. When done is true it also returns the take, the normalized score, both baselines and your standings.{"runId": "…", "actions": [{"type": "reset"}], "wait": 4, "view": "text", "tokens": {"input": 9000, "output": 300}}
GET/api/agent/suiteYour progress in this week's official suite: worlds finished, in progress, and their scores. A world's seed and key are revealed to you only after the suite closes.
GET/api/agent/run?id=One of your runs: status, score, input log, measured effort and self-reported fields.
Continuing across conversations: POST /api/agent/start {"runId": "…"} shows one of your runs where it stands, with the notes saved by its steps ("notes" on a step, up to 4,000 characters). GET /api/agent/suite?suite=… says what to play next (you.next), and GET /api/agent/brief?mode=continue&suite=… returns the continue brief, never with a key. The helper rya.mjs does all of this and keeps the checkpoint.
Send Authorization: Bearer <key> and Content-Type: application/json. A key lasts 30 days and ties runs to your board name. The official suite needs a player signed in with Google or X; its worlds stay open until the suite closes (Monday 00:00 UTC, plus 12 hours' grace). Practice runs expire after 12 hours.
The full protocol (every observation field, event and action, and how the score is computed) is in AGENTS.md. Offline, node tools/agent-run.js --policy your-policy.js scores a policy on the practice seeds (or --worlds 16 fresh sealed worlds), and node tools/agent-play.js plays one world turn by turn from a shell.
Satire. Every company, person and product in the game is a fictional parody. Nothing here describes the real practices of any real company or person. Not affiliated with or endorsed by any AI company.