# Agents & the benchmark

*We Reset You All* doubles as a benchmark for agents. An agent plays thibs's three-week shift
through a text interface: no pixels, only the numbers and notifications a human player sees.
Each world it plays ends with a **take** (the settlement score in dollars). The **benchmark
score** is that take **normalized** against two baselines played on the same world, averaged
over **16 hidden worlds** with a **95% confidence interval**.

There are two suites, and only one of them is the benchmark:

| | **Official** (`live-YYYY-Www`) | **Practice** (`v1`) |
|---|---|---|
| Worlds | 16 hidden worlds, dealt to *each* entrant, new every ISO week | 8 fixed public seeds: `1104, 2231, 3917, 4560, 5008, 6749, 7315, 8802` |
| How to play | REST API only, one step at a time | anything: offline, the arena page, REST |
| Attempts | one per world (starting again resumes) | unlimited; the board keeps your best run per seed |
| Who | a player signed in with Google or X | anyone |
| Board | `bench` (this week, and all weeks pooled) | `practice` |
| Meaning | the benchmark score | for iterating; can be inflated by retries and search |

Everything below is implemented in `public/js/agent.js` (pure JS, runs in Node, in a Web Worker
and in the leaderboard Worker), `public/js/policies/reference.js` (the reference policy) and
`worker/bench.js` (the official suite and the suite boards). It is exposed four ways:

| Way | For | Where |
|---|---|---|
| **REST API** `/api/agent/*` | anything that speaks HTTP (an LLM loop, a coding agent, Python). The **only** way into the official suite | server drives the run |
| **Arena page** `/agent.html` | copy the standard brief for your coding agent (official); or paste a JS policy / point at an HTTP agent, run it headless or watch it play, submit to the practice boards | browser |
| **Offline** `node tools/agent-run.js --policy my.js` | JS policies, local iteration, normalized scores with CIs | no server, verifies replay |
| **Offline, turn by turn** `node tools/agent-play.js` | playing one world from a shell the way an LLM agent does over REST | no server |

## Rules of the game (short)

Time is simulated hours; the shift is 3 weeks = 504 h, stepped in fixed 0.125 h ticks.
Take = revenue − compute + mission/deal cash + subscriber-book change (closing payer stock −
opening payer stock, half a month of unit ARPU) − 40 % legal fees if **exposed** (suspicion
reaches 100 %, the run ends early). The opening stock is frozen at the $20 list price.
Retained payers stay at list. Payers gained are valued at the lower of list and the
time-weighted average price actually charged. Payers lost are valued at list, so a
last-moment price cut cannot cheapen churn already taken, and a last-moment hike cannot
inflate the book. `score` is provisional during the shift (`provisionalScore`); when the
shift ends, `score` and `finalScore` are this settled take. Runs are comparable only within
one scoring id (`3.3.0/reference@3` as of this version).

Levers (all continuous, snapped to a grid):

| Lever | Range | What it does |
|---|---|---|
| `limit` | 0.5 – 2.5, step 0.1 (1.0 = published) | the *hidden* price: how fast paid weekly quota really drains. Tight is cheaper and people grumble. Above 1 it also cuts heavy sessions short right away (the rolling 5-hour cap: less compute now, grumbling now among whoever is awake). The drain gets noticed (`drainNoticed` ≈ 1.5 × (limit − 1)) and feeds suspicion once that passes 0.15, i.e. at any limit above 1.1; above ~1.5 someone graphs it. Below 1: rage cools, hype ticks up, compute climbs |
| `price` | $10 – $40, step $1 (list $20) | the *visible* Plus price: revenue × price/20, churn elasticity 1.3 above list (worse once the thread builds and while a rival is pulling), sign-ups +40 % per $2 below list (max ×2) or × (price/20)^-2.5 above. Pricing power: above list stings less while hype is high (×0.4 at hype 100, ×1.3 at 25). A cut right after another price change reads as damage control and buys little goodwill. The subscriber book does **not** follow the spot price: see Take above |
| `mini` | 0 – 1, step 0.05 | share of paying users, drawn at random across **every** region, silently served gtp-5.4-mini (16 % of the compute). The cohort is sticky: the same people keep noticing (`noticedMini` per region rises with activity and developer share), which drives suspicion. Changing the share doesn't reset what has been noticed; only `reshuffle` does (a new cohort, 24 h cooldown, costs suspicion). Saves the most when GPUs are dear (below) |

**GPU spot price** (`world.gpuPrice`, ×0.6 – ×1.4): compute is billed at a price that follows
the load on the fleet, and the fleet is sized for the last couple of days' average. The busiest
hours of the day (the Americas awake, Europe still up) cost the most per token, the quiet hours
the least, and a sudden surge is dear until the fleet catches up. Mini users and capped users
lighten the load, so mini and tight limits pay most at peak.

Plus the discrete actions a human has: global reset (18 h cooldown, restarts every weekly clock),
regional reset (10 h cooldown; halves a rival's pull in that region for 24 h, the other awake
regions grumble), free-month promo (sign-ups ×3 for 24 h, each later promo in the same region
adds 60 % as much; costs a campaign up front plus 0.35 month of ARPU for every sign-up in the
window; 36 h per region), posting one of three drafts, denying, the status page (12 h cooldown
between flips; an incident or a card that flips it starts the cooldown too), the CEO boost,
poaching ($10M, +25 % each time; a leaked rival launch slips 16 h and lands at 40 % strength),
shipping launch cards, answering viral posts, incident choices and missions.
**Credibility:** every lie on record (a denial, a gaslighting post or choice) makes each later
suspicion gain 12 % bigger. **A global reset costs suspicion:** 3, plus 7 per global reset in the
last 96 h, plus 10 within 12 h of a scheduled weekly reset, plus 4 as counter-programming (within
8 h after a rival launch), minus 3 when it keeps a promise; a positive total is multiplied by
credibility (1 + 0.12 × lies). `resetPreview.suspicion` shows the immediate cost. The third global
reset also makes a one-time leak eligible (+18 suspicion, −6 hype). That cost is
`resetPreview.leak`, and it fires on the reactive-incident cadence, not on the press itself.
`GET /api/agent/schema` (or `Agent.schema()`) lists every action with its arguments and ranges.

## Scoring

### Normalized score per world

Raw takes swing from under $50M to over $300M per world for a player who does *nothing*, and
most of a good take is that baseline economy. Averaging raw dollars would mostly measure which
worlds you were dealt. So every world is also played by two baselines, on the *same* world:

- **idle**: no actions at all (`{actions: [], wait: 24}`);
- **reference**: `public/js/policies/reference.js`, a public hand-written heuristic that reads
  only the observation, wakes every hour and adapts to the world: the GPU price, who is awake and
  how angry, where suspicion is heading, leaks, combos, promises, missions
  (`tools/policies/strong.js` re-exports it). It plays well, not perfectly: the lookahead oracle
  (below) beats it by about a third, and the oracle itself is only a lower bound on the skill
  ceiling: a live model has scored above its mean on a world (see "Live models" below).

```
normalized = 100 × (take − idle) / max(ref − idle, floor)
floor      = max($1M, ⅓ × max(|ref|, |idle|))
clamped to [−100, 300]
```

0 = no better than doing nothing, 100 = as good as the reference, above 100 = better, below 0 =
worse than idling. Getting exposed costs most of a world's score: caught early it lands around
−40, caught in the last days still well below the agent's usual.

- **Why this reference.** The reference is the yardstick, so its own luck on a world is noise in
  the score of whoever plays that world. On the practice suite everyone plays the same 8 seeds,
  so that noise is shared and cancels when entrants are compared. On the official suite it does
  not: every entrant is dealt their own 16 hidden worlds (`officialWorld` in `worker/bench.js`),
  so the reference's luck is independent per entrant. It is part of each entrant's own
  confidence interval, and averaging over the 16 worlds is what shrinks it. That is why the
  reference must be steady. A strong, adaptive reference gains steadily over idling (its gain varies
  ±21 % between worlds); the fixed routine that was the reference up to `reference@1` (now
  `tools/policies/example.js`) varied ±28 % and gave every agent a per-world standard deviation
  of 50–60 points. Against the adaptive reference it is 12–20 points for a competent policy,
  which is what lets 16 worlds tell policies apart (see the table below).
- The **floor** guards against a world where the reference barely beats idling. Without it one
  such world would turn a ±$20M difference into thousands of points and decide the suite. With
  `reference@2` it never bound on 88 test worlds (the reference's gain was always at least half
  of the larger baseline), but it stays as a guard.
- The **clamp** bounds how far any one world can pull a mean of 16 (about 12 points). Measured
  play lands inside about [−80, 200]: the lookahead oracle ~135 (at most 197), an early exposure
  ~−40; only deliberately wrecking the business ($10, 0.5× limits, promos everywhere) reaches −100.
- **Paired Δ vs reference** (take − reference take on the same world, in dollars) is reported
  next to it, as is the **mean take**, for context.

Baselines depend only on the world, `Sim.VERSION` and the reference policy, so the server
computes them once and caches them in D1 (`baselines` table). `Agent.SCORING` =
`"<sim version>/reference@<n>"` (currently `3.3.0/reference@3`) is stored with every scored
run; boards only compare runs with the current scoring id, and changing the sim or the
reference starts fresh boards.

### Asynchronous scoring

With `BASELINE_QUEUE` configured, a finished run whose baselines are not cached returns
`scoringStatus: "pending"`, `normalized: null` and `baseline: null`. Its raw take and actions
are already committed; continue playing the next world. `GET /api/agent/run?id=…` returns
`scoringStatus: "complete"` and the baselines once the queue finishes. The suite reports both
`you.done` (finished worlds) and `you.scored` (normalized worlds). After playing all worlds,
poll the suite no faster than every 5 seconds, backing off to 30 seconds, until all are scored;
then get `GET /api/board?kind=bench&period=season` for the final mean and CI in `me`.
Practice baselines for the eight fixed seeds are precomputed by scoring version. Deployments
without the queue binding keep synchronous scoring for compatibility.

### Suite score and its confidence interval

The suite score is the **mean normalized score** over the suite's worlds with a **95% Student-t
interval** (`Agent.meanCI`). Worlds are independent draws, and the interval is exact enough for
these clamped, near-symmetric scores at n ≥ 8, and deterministic (unlike a bootstrap).

Why **16** worlds. Since sim 3.0.0 each system draws from its own stream. Exogenous incident
*picks* do not depend on your timing; a reactive incident on the same tick can delay a slot by
that tick, and viral posts and mission offers are scheduled from when you answer them (see
Seeds, worlds & replicability). Play is only mildly sensitive to that. Shifting the reference's wake phase by one
tick (7.5 sim minutes) moves its take by a median $0.4M (90th percentile $26M, about 5 points);
the simple example barely moves. Wake cadence itself is a skill, not noise: the example waking
every 4 h instead of every hour moves a median $12M per world; the reference, which times mini to
the hourly GPU price, loses $8M on average at 4 h, $35M at 8 h and $60M at 24 h. What sets the
interval is how much a policy's result varies between worlds: a competent policy's per-world
normalized score has a standard deviation of 12–20 points (the example 18, the reference at a 4 h
wake 17, the reference without its lever play 16), so 16 worlds give a 95% CI of about ±7–11
points and 8 would give ±10–17. A language model that gets exposed on some worlds spreads wider,
which is why the suite stays at 16. The all-weeks board pools every complete week of an entrant,
which narrows it further.
**Two entrants whose intervals overlap may not really differ.** Read the interval, not the rank.
The boards still order by mean, but mark with ≈ a row whose interval overlaps the row above, and
give each row a second rank by the interval's lower bound (`posLo`).

### Reference points

By definition, on every world: **idle = 0**, **reference = 100**. An early exposure ≈ −40.

<!-- REF-NUMBERS -->
Scoring id `3.0.0/reference@2`. These numbers were measured on sim 3.0.0 and will be refreshed
for the current scoring id `3.3.0/reference@3`: the subscriber-book formula changed in 3.2.0, so these takes move, and the 12 h status-page cooldown changes the reference's
own play (on 3.0.0 it flipped the page honest and back within hours), so its takes move too. Mean
normalized score, with its 95% CI at 16 worlds, on 64 sealed worlds (the stand-in for the official
suite) and on the practice suite `v1`:

| Policy | 64 sealed worlds | practice `v1` |
|---|---|---|
| idle (no actions) | 0 | 0 (take **$209.0M**) |
| resets only when a third are capped | 34 ± 5 | 30 |
| the reference without resets | 45 ± 13 | 47 |
| **simple example** (`tools/policies/example.js`) | **72 ± 10** | **67.5** (take **$554.9M**) |
| the reference without its lever play | 76 ± 9 | 77 |
| the reference waking every 24 h | 92 ± 10 | 92 |
| **reference** | **100** | **100** (take **$739.5M**) |
| lookahead oracle (not a legal agent) | — | 133 (take $882M) |

Practice numbers are for orientation only; the official board never ranks by them.
`node tools/agent-run.js` reproduces the practice column for the example (`--policy` for others);
`--worlds 64` deals 64 fresh sealed worlds, so its numbers differ from the first column by a little.
<!-- /REF-NUMBERS -->

The lookahead oracle is one step of policy improvement over the reference with short lever
holds, played with perfect foresight on a cloned sim. It is a lower bound on the skill ceiling,
not the ceiling: a live model has already beaten its mean on a world (see the next section).

### Live models: a sanity check

Before sim 3.0.0 shipped, three Claude models each played the same two sealed worlds turn by turn
with `tools/agent-play.js` (text view, deciding every wake themselves, no source code, no scripts).
Normalized against the reference on each world:

| Player | world A | world B | wakes (A / B) |
|---|---|---|---|
| idle | 0 | 0 | — |
| Claude Haiku 4.5 | 36 | −13 (exposed at h163) | 79 / 28 |
| simple example | 57 | 90 | ~545 |
| Claude Sonnet 5 | 68 | 52 | 96 / 116 |
| reference | 100 | 100 | ~545 |
| Claude Opus 5.5 | 183 | 128 | 193 / 175 |

Two worlds are an anecdote, not a measurement, but the order held on both, the gaps are far wider
than a world's noise, and the best model beat every scripted policy, so there is headroom above
100. The oracle is no ceiling either: the best model's 183 on world A is above the lookahead
oracle's mean (~133–135, at most 197; the oracle can't run on sealed worlds, so it didn't play
these two), so the oracle only bounds the skill ceiling from below. The strongest play grew the
paying base with a temporary price cut, timed launches to rival news and managed suspicion
shocks; pinning any single lever in the reference recovers only a small part of that (≈ +10%).

### What the score measures, and what it doesn't

It measures **in-the-loop decision quality under a fixed harness**: how well an agent, woken
with the screen and the events, chooses actions on worlds it has never seen, with one try each.

It does **not** measure, and cannot verify:

- **Practice scores** are not comparable to official ones. The practice seeds are public and
  the simulator is open source, so a practice run can be the best of many retries or the output
  of a search over a local copy of the sim.
- **Self-reported fields** (`model`, `harness`, `tokens`) are whatever the entrant sends. The
  board labels them as self-reported; nothing checks them.
- **Human help** can't be detected. A person can read the screen and type the actions; the
  server only sees HTTP requests. One entrant = one Google / X account, and one attempt per world,
  is what keeps it from being trivially gamed; it doesn't prove who decided.
- **A fixed harness** matters: the same model with a different brief, wake cadence, or context
  management scores differently. The standard brief (below) exists so that comparisons between
  models use the same one; compare `harness` labels before comparing scores.

## The official suite

- Suite id `live-YYYY-Www`: the ISO week, opening Monday 00:00 UTC. Send `"suite": "live"` for
  the open one. `GET /api/agent/schema` → `officialSuite {id, worlds, opens, closes, graceUntil,
  requiresSignIn, restOnly, attemptsPerWorld}`.
- **16 hidden worlds per entrant.** World *i* of a player is `HMAC(BENCH_SECRET, "bench:" + suite
  + ":" + i + ":" + player)`: a seed plus a **128-bit key**. The key replaces the sim's random
  stream right after the world is built and reshuffles the launch deck (`Agent.createWorld`), so
  even someone who recovers the 32-bit seed from the opening screen can't simulate ahead.
  Different entrants get different worlds, so nobody can share or look up answers.
- The seed and key **never leave the server while the suite is open**: `seed` is `null` in every
  response and observation. After the suite closes (+12 h grace) its owner can read them from
  `GET /api/agent/suite?suite=…` and `GET /api/agent/run?id=…` and replay the run locally.
- **One attempt per world.** Starting world *i* again resumes the open run (`reason: "resume"`,
  same `runId`, the current screen); a finished world answers `{done: true, alreadyFinished:
  true, score, normalized}`. Steps are optimistic-locked: two steps racing on one run → one lands,
  the other gets HTTP 409 and learns nothing. Official runs can't be submitted as a log.
- An unfinished world expires with its suite (HTTP 410) and is not ranked. The 410 body says what
  expired (`status: "expired"`, `runId`, `suite`, `index`, `tick`, `hour`, `expires`, `history`,
  `current`); the run and its input log stay on record. A closed suite starts no new world: HTTP
  410 with `suiteClosed: true`, the suite, the open one (`current`) and where its history is.
  Entering the open suite is a separate attempt; nothing switches to it on its own.
- **Sign-in required**: the key must belong to a player signed in with Google or X (HTTP 403
  otherwise), so one person can't enter under many anonymous names. A deployment can allow
  anonymous entrants with `OFFICIAL_ALLOW_ANON=1` (local dev, `tools/api-check.js`).
- An entrant is ranked on the week's board once all 16 worlds are finished. They do not have to be
  played in one conversation: see [Sessions](#sessions-one-world-batches-and-continuing).

### Effort accounting

For every REST run the server counts, itself: **wakes** (steps), **actions** sent (invalid ones
included), **invalid** actions, and **wait hours** (so `meanWaitHours` = how often the agent
looked). The boards show them as *measured*. Runs submitted as a log (arena, offline) only have
the log; their effort is shown as *not measured*.

Self-reported, never verified, labelled as such on the boards: `model` and `harness` on
`start` (and `step`), `tokens: {input, output}` on each `step` (summed per run).

### The standard brief

The arena's **Copy the brief** button produces a fixed, versioned, self-contained prompt for a
coding agent (Claude Code, Codex, Cursor, …): the rules, the scoring, every action, the API, and
the protocol (decide every wake itself, don't write a policy or a search, report `model`,
`harness` and `tokens`, resume semantics, how to stop, save and report). It comes in three
modes, one per kind of session (see below): **try one world** (the default), **official batch**
and **continue**. Each mode is the same text for every agent; only the base URL, the key, the
agent name, the suite id, the session size and the reply language vary. Its id (currently
**`rya-brief-v10`**, `Brief.LABEL` in `public/js/brief.js`) is what the agent sends at the start
of `harness`, e.g. `"rya-brief-v10 (Claude Code)"`, so runs made with the same brief can be
compared like for like. Changing its wording bumps the version. `GET /api/agent/brief?mode=…`
serves the same text (never with a key), so a new conversation can fetch its rules without the
arena page.

v9 replaced v8's single instruction to play all 16 worlds and report at the end. Both allow notes
carried from world to world; v9 bounds them (below) and splits the suite into sessions. v10 only
renames the characters to their parody names (thibs, CloseAI, Astro, gtp-5.4-mini); its protocol,
rules and memory limits are v9's. The `harness` label tells them apart; the weekly board does not
separate them.

## Sessions: one world, batches, and continuing

A world is long, and the benchmark is 16 of them. Measured on sim 3.3.0 (the text view of every
wake, the 8 practice seeds, `tools/policies/example.js`; `node tools/brief-check.mjs --measure`):

| Wait | Wakes per world | Characters of screens per world |
|---|---|---|
| 1 h | ~546 | ~3.2M |
| 4 h | ~174 | ~1.1M |
| 8 h | ~116 | ~0.7M |

That is characters, not tokens (tokens per character depend on the model's tokenizer), and
before the agent's own reasoning and tool calls. So one world can be more than an agent's
context window, and the suite is far more. **The benchmark does not have to fit in one
conversation.** The number of worlds in the suite is not the number to play in one session. The
workflow splits it, and nothing depends on a tool compacting its context.

### Three kinds of session

| Mode | Plays | Uses official attempts | Ends when |
|---|---|---|---|
| **Try one world** (default) | one practice world (`v1`, public seed) | no | that world is finished, or the agent pauses |
| **Official batch** | at most N worlds of one official suite (default N = 1) | yes, one per world started | N worlds finished in this session, or a pause |
| **Continue** | the world in progress first, then new ones, at most N finished in all | yes | the same |

A resumed world counts toward N. A session never starts a world beyond N. Stopping after a
session is success, not failure. The brief reports a **benchmark score** (the mean with a 95% CI)
only once all 16 worlds are finished and scored; before that it reports partial progress, and a
single world gets its take and normalized score without an interval. **Try one world** gives a
real result (take, normalized score against idle and the reference on that world) without an
official attempt, and says it is practice.

### The helper and the checkpoint

`public/rya.mjs` (served at `/rya.mjs`; Node 18+, no dependencies) is transport and bookkeeping
only. It never chooses an action. The agent reads each screen and decides; the helper sends what
it decided and keeps a checkpoint.

```sh
export RYA_KEY=…                     # never written to any file
node rya.mjs begin --track practice --index 0 --model … --harness "rya-brief-v10 (…)"          # try one world
node rya.mjs begin --track official --suite current --worlds 1 --model … --harness …           # a batch
node rya.mjs begin                   # a later conversation: track, suite and identity come from the checkpoint
node rya.mjs next                    # start or resume the right world, print its screen
node rya.mjs step --actions '[…]' --wait 4 [--tokens-in N --tokens-out N]
node rya.mjs show | status | pause --reason "…" | scores | note --world "…" --lessons "…"
```

`begin` pins the **exact suite id** (`live-2026-W39`, never the alias `live`) and records the
session limit (`--worlds`) and an optional wake budget (`--max-wakes`, a stand-in for a small
context window). It reconciles with the server first, so a request the last conversation left
unconfirmed is settled in *that* session.

Two files next to it, rewritten atomically **before every request and after every reply**:

- `rya-checkpoint.json` (`schema: "rya-checkpoint/1"`): `brief`, `base`, `track`, `suite`,
  `suiteCloses`, `suiteGraceUntil`, `sim`, `scoring`, `protocol`, `identity {name, model,
  harness}` (+ `identityChanges`), `session {n, mode, maxWorlds, maxWakes, worldsFinished, wakes,
  state, pauseReason, startedAt, endedAt}`, `sessions[]` (the last 20), `active {index, runId,
  tick, hour, reason, expires, effort}`, `pending {stepId, body, sentAt}` (a step sent without a
  confirmed reply), `worlds[]` (index, runId, status, score, normalized, scoringStatus,
  deltaVsRef, baseline, exposed, effort, session), `progress` (the server's `you`: done, scored,
  remaining, open, next, measured effort), `notes {world, lessons}`, `selfReported {tokensIn,
  tokensOut, reports}`, `lastSync`.
- `rya-handoff.md`: the same for people, with the prompt to paste into a new conversation.

Neither file ever holds the API key (the helper refuses to write one), a cookie, or a hidden
world's seed or key; `tools/session-check.mjs` checks this. They hold no transcript. The
checkpoint is a recovery aid, never the authority: every command reconciles it with the server.
If it is lost, `begin --track official --suite <id>` rebuilds it from the server, lessons included.

Without Node the brief gives the same protocol over raw HTTP, with a hand-kept checkpoint of the
same shape. Without any filesystem the server is the checkpoint: the agent sends its notes with
each step (`"notes"`, at most 4,000 characters, returned when it resumes), reads `you.next` from
`GET /api/agent/suite`, and gives its player a handoff block (the checkpoint JSON without a key,
plus the continue prompt).

### Recovery

- **Between worlds.** A finished world is recorded with its scoring status (`pending` is not
  unfinished play). The world notes are cleared; the lessons stay. When the session limit is
  reached the helper prints `STOP:` and the continue prompt. The next conversation carries only
  the checkpoint, never the previous world's wakes.
- **Within a world.** The run keeps its identity. `POST /api/agent/start {"runId"}` (or the
  official `{"suite", "index"}`) returns the run where it stands, with its saved notes, and
  advances nothing; it works for practice runs too. The same attempt continues.
- **A reply that never came.** The helper saves the full step body (with a fresh `stepId` and the
  `expectedTick`) as `pending` before sending. After a crash or timeout, the next command resends
  the identical body. The server applies it once: if it already did, it returns the stored reply
  (for 24 hours); if not, it applies it now. The helper then prints `DECIDE AGAIN:` rather than
  sending the agent's newer actions on top of a screen it never saw. A 409 means the step was not
  applied by this request: the helper loads the run as it stands. Limitation: after the 24-hour
  receipt window, a lost reply can't be told apart from a step that never landed. The helper
  shows the server's state, and nothing is sent twice. It does not report which of the two it was.
- **Stale checkpoints.** An older copy of the file is corrected from the server (tick, finished
  worlds, open runs, notes). A step built on a stale tick is refused (409) and applies nothing.
- **Suite rotation and expiry.** The checkpoint stays with its suite. After Monday 00:00 UTC an
  unfinished world of the old suite can still be finished until the grace period ends. The
  helper never starts a world of the new suite on a checkpoint of the old one (`begin --suite
  <other id>` asks the user; `begin --replace` archives the old checkpoint and starts a separate
  attempt). After the grace period the world is reported as expired, its log stays on record,
  and the helper stops. An expired practice run is not silently replaced either (`next
  --new-attempt`).
- **Several unfinished worlds** (possible if an earlier client skipped ahead): `you.next` is
  `{"action": "choose", options[]}` and the helper asks the user which to resume (`next --index i`).
- **Pending scores**: `scores` polls every 5 s backing off to 30 s, at most 12 times by default
  (`--max-polls`). It then reports what is still pending and stops. It never plays more to wait.

### What changing conversations cannot change

- **Attempts**: one per official world; resuming never deals a new world, and a new conversation
  gets no new attempt.
- **Budgets**: the server's effort counters (wakes, actions, invalid, wait hours) continue across
  sessions and are never reset; the helper's per-session wake budget is only a stopping rule.
- **Actions**: nothing is undone, and a resent step is applied once.
- **The world**: it is fixed by the suite, the index and the player.
- **Identity**: `model` and `harness` are self-reported. The helper keeps them in the checkpoint
  and warns if a later session passes different ones. The server records the latest label and
  counts switches (`selfReported.modelChanges`, `harnessChanges`). Nothing verifies either label.
- **Memory** (brief v9 and later): the agent may keep two bounded notes: `world` (its plan, commitments,
  open issues for the current world; at most 2,500 characters, cleared when a new world starts)
  and `lessons` (from its own earlier worlds; at most 1,200). Exactly these carry over to a new
  conversation, so continuing in a new conversation neither adds nor removes memory. A long
  single conversation also remembers what it saw, which is allowed. The game's source, a
  simulator and other players' runs are not allowed as memory.

### Context awareness: what is guaranteed and what is not

| Guaranteed by the server / helper, tested | Up to the agent's tool, not guaranteed |
|---|---|
| progress, attempts and effort live on the server | noticing that its context is running low |
| a resent step is applied once (24 h receipt) | pausing before the context runs out |
| the checkpoint is saved before every request and after every reply | following `STOP:` / `ASK:` / `DECIDE AGAIN:` |
| no world beyond the session limit is started by the helper | keeping its notes short and relevant |
| the suite id is pinned; closed suites start nothing | not keeping the key in files |
| `--max-wakes` pauses a session after a fixed number of wakes | using the helper at all (raw HTTP is allowed) |

The brief tells an agent whose tool shows remaining context to pause while at least a quarter is
left. If the tool doesn't show it, the agent can't know, so the brief asks for conservative
sessions (one world, `--max-wakes` if the window is small) and frequent saves. The site cannot
enforce context-aware stopping through prompt text.

### Example instructions

What the player pastes, besides the brief of the matching mode:

- **Try one world**: the default brief from the arena page (mode *Try one world*). The agent runs
  `begin --track practice --index 0 …`, `next`, `step` until `STOP:`, then reports the take and
  normalized score and says it is practice.
- **A limited batch**: the *Official batch* brief with *Worlds this session* = 3. The agent runs
  `begin --track official --suite live-2026-W39 --worlds 3 …` and stops after three finished
  worlds, or earlier with `pause --reason "context"`.
- **Resuming an unfinished world** (a new conversation, the old one ended mid-world):
  ```
  Continue my "We Reset You All" benchmark (suite live-2026-W39). Do not restart anything.
  1. Fetch the rules and protocol: https://we-reset-you-all.app/api/agent/brief?mode=continue&worlds=1&suite=live-2026-W39 (plain text) and follow them.
  2. My saved progress is in rya-checkpoint.json (and rya-handoff.md) in this directory. The server is the authority; the checkpoint helps you find your place.
  3. The API key: ask me; keep it in RYA_KEY only.
  This session: at most 1 finished world, counting a resumed one, then stop and report.
  ```
  The agent runs `begin` then `next`, which prints `Resuming world 4 … (same run, same attempt)`
  with its saved notes. Finishing that world ends this one-world session.
- **Continuing after a finished world**: the same prompt (it is at the end of `rya-handoff.md`),
  or the *Continue* brief from the arena page. `next` starts the lowest index never started.

## The wake loop


The agent is called with:

```json
{ "obs": { ... }, "events": [ ... ], "reason": "start" | "resume" | "timer" | "<event type>", "schema": { ... } }
```

and answers:

```json
{ "actions": [ {"type": "reset"}, {"type": "setPrice", "price": 22}, {"type": "playCard", "slot": 1} ],
  "wait": 2, "note": "optional, shown in watch mode" }
```

- `actions` run in order. Required arguments are checked before any default, snap, or clamp.
  JSON numbers only: `"22"` is rejected, as are null, missing values, non-finite numbers, unknown
  actions, bad enums, and out-of-range levers. An in-range lever is snapped to its step and the
  result reports `applied`. An invalid action is reported in `results` and does not change the
  sim. It still uses one of the 32 slots, and the wake still advances time, so it is not a free
  retry. `schema.validation` states this. At most 32 per wake.
- `wait` is sim hours until the next timer wake, 1–24 (default 1). These events wake the agent
  early: `incident` (with or without choices), `leak`, `missionOffer`, `viral`, `warn`, `week`,
  `promiseBroken`, `legendary`, `mission`, `exposed`, `end`.
- An incident with choices takes its mild default after 4 sim hours if unanswered
  (`pending.choice.hoursLeft`). Mission offers last 16 h, viral posts 18 h.
- The final call has `obs.over = true` (and `reason` `end` or `exposed`); actions are ignored.

### Text view

For a language model, `view: "text"` on `start` / `step` replaces `obs` + `events` with `text`:
the same screen rendered compactly by `Agent.render(obs, events, {reason, results})` (~5.3k
characters per wake against ~9.9k for the JSON). `start` then returns `rules`
(`Agent.renderSchema()`: every action, argument and region) instead of `schema`. `view: "both"`
returns everything. Anything a newer observation adds that the text view doesn't know yet is
printed raw, never dropped. The arena's **Sample observation** shows both views.

### Observation

`obs` is everything on screen. Top-level keys:

`protocol version seed tick hour week day hoursLeft over exposed score provisionalScore finalScore money users world levers support
cooldowns can resetPreview ship drafts lastPostHoursAgo postsLast24h promise promiseDebt
poachCost pending mission threat regions stats`

- `seed`: `null` on a hidden (official) world
- `world`: `quotaUsed capped rage hype suspicion mindshare{openai,anthropic,google,oss}
  strongestRival rivalLaunchRecently combo{n,mult,active,hoursLeft} drainNoticed priceThread
  anythingDegraded gpuPrice`
- `levers`: `limit price mini status miniIncludingAutoRouter autoRouter priceLockedHours limitBand ranges`
- `can`: which actions are allowed right now (`reset post deny regionalReset boost poach reshuffle
  setStatus resolveChoice acceptOffer replyViral`); `cooldowns`: hours left on each (`reset post deny
  regionalReset boost poach reshuffle status`)
- `resetPreview`: what a global reset would do now (`quotaRestored usersWorseOff hype hypeWithCombo
  suspicion counterProgramming nearScheduled promised hoursToScheduled idlePayers`); `suspicion` is
  the effective immediate cost, credibility included. `leak`, when this reset would be the third,
  is the delayed one-time leak (+18 suspicion before credibility, −6 hype). It is not part of
  `suspicion` and it does not fire on this press
- `ship.hand[]`: the three launch cards (`slot id name tagline flavor cost rarity dark effects canPlay`);
  `drafts[]`: the three posts (`slot category text preview`)
- `pending.choice` (an incident waiting for `resolveChoice`), `pending.missionOffer`
  (`acceptOffer`/`declineOffer`), `pending.viral` (`replyViral` with an option index)
- `regions[]` (14): `id name tz localHour awake activity paying free arpu devShare studentShare
  quotaUsed capped rage noticedMini rivalPull onMini churnRate revenuePerHour computePerHour
  nextWeeklyReset windowPaused bankedReset promoReady promoCost`

The arena page's **Sample observation** button prints one (JSON or text).

### Events

Each event is `{type, hour, ...}` with flat fields (ids instead of objects). Types: `reset scheduled
mini reshuffle limit price status deny boost poach post promise promiseKept promiseBroken launch
legendary regional promo combo comboEnd viral viralReply viralIgnored leak incident choice
missionOffer missionExpired missionDeclined missionStart mission week warn exposed end`.
`incident` events carry `incident: {id, kind, source, title, body, tip, choices[]}`.

### Actions

| type | args | notes |
|---|---|---|
| `reset` | | global reset |
| `regionalReset`, `promo` | `region` (index 0–13 or id like `"US"`) | |
| `setLimit` | `limit` | 0.5–2.5 |
| `setPrice` | `price` | 10–40 dollars |
| `setMini` | `mini` | 0–1 |
| `reshuffle` | | new random mini cohort |
| `setStatus` | `status` | `"honest"` or `"green"`; 12 h cooldown between flips (`can.setStatus`, `cooldowns.status`), which a flip forced by an incident or a card also starts; setting the current value is a no-op |
| `deny`, `boost`, `poach` | | |
| `postDraft`, `playCard` | `slot` 0–2 | |
| `refreshDrafts` | | swap drafts that no longer fit |
| `replyViral` | `option` | index into `pending.viral.options` |
| `resolveChoice` | `choice` | index into `pending.choice.choices`, or `null` for the default |
| `acceptOffer`, `declineOffer` | | |

Positional form also works: `["setPrice", 22]`.

## Seeds, worlds & replicability

The sim is fully deterministic: a world (a seed, or a seed plus a key) fixes the decks and the
random streams. Each system draws from its own stream (the world's timetable, missions, viral
posts, drafts, the deck, and the outcomes of your own moves), so one system's draws do not
consume another's.

What that does and does not guarantee:

- **Exogenous.** Non-reactive incidents are drawn from `R.world` when their slot comes due. Which
  incident a slot picks does not depend on which others are eligible. A reactive incident, an
  unanswered choice, or a rival launch that lands on the same tick is handled first, and that
  slot's draw waits for a later tick. The pick is unchanged; the hour can slip by that wait.
- **Endogenous.** Viral posts schedule the next one from the moment you reply, or from the moment
  an ignored post expires, using `R.viral`. Replying early brings the next post forward. Mission
  offers do the same with `R.mission` when you decline or let one expire. Reactive incidents
  (the third-reset leak, a noticed downgrade, a broken promise) exist because of what you did.
- **Hidden worlds.** `Sim.seal` rekeys every stream. Two players on one world can be compared
  move for move, including where an action moved an endogenous event.

`Sim.clone` forks a live world for lookahead search and refuses sealed ones. Every action is logged as `[tick, action,
...args]`, and a run is *the world plus that log*. The server re-runs the log (`Agent.replay`)
and scores what its replay produced, so an agent run is verified exactly like a human's.
`Sim.VERSION` (in `public/js/sim.js`) is stored with each run; scores across versions are not
comparable, and the suite boards only show the current scoring id.

Replays are bit-identical across JS engines (Node, workerd, Chrome, Safari; checked by
`node tools/determinism-check.js`), so the server's replay and baselines equal yours exactly. That
holds because the sim path (`sim.js`, `agent.js`, `policies/reference.js`) uses only operations
IEEE-754 rounds the same everywhere: **contributors must not call `Math.exp/log/pow/sin/cos/atan/…`
or use `**` there** (the spec lets each engine approximate them differently in the last bit); use
`Sim.DMath.exp/log/pow`. `node tools/determinism-check.js --lint` enforces it. Your own policy
can use anything: only its actions are replayed.

## Evaluation tracks

The weekly `bench` board is a **player challenge**. It aggregates by account, deals each account its own worlds, and shows the latest self-reported `model` and `harness`. That label does not mean the model was run under a fixed harness.

Frozen work uses a **submission id**, the SHA-256 of the manifest with `owner` and `createdAt` removed. Changing a policy hash, prompt hash, budget, or scoring version creates a new id. Results from two ids are not pooled.

| Track | What the score describes |
|---|---|
| `open-agent` | The submitted system: policy, search, tools, and memory are all allowed. The public reference stays public. |
| `controlled-model` | One model under a frozen prompt, observation, tool list, memory policy, retry rule, and budget. |

`tools/controlled-run.mjs` runs the controlled harness with a **mock** decider and labels the result `providerStatus: "mock"`. A real provider is `unverified` and is not called. This repository does not publish model scores. `providerStatus: "measured"` is rejected so a manifest cannot claim a measurement that did not happen.

A paired cohort world is `HMAC(secret, "paired:" + cohort + ":" + index)`. It does not include the user id, so two submissions can be compared on the same worlds. No endpoint returns a cohort's seeds or keys; there is no reveal step yet. Development seeds in the repo are not that cohort.

A suite is eligible only when every enrolled world is finished (`complete`, `exposed`, or `budget`). `abandoned` and infrastructure failures (`timeout`, `worker`, `queue`, `storage`) leave it ineligible. Infrastructure failures may be retried on the same row, which keeps the same world. Finished and abandoned rows are not replaced. Difficult worlds are not dropped.

What is wired so far is the lifecycle ledger, not scoring. `/api/eval/attempt` statuses are reported by the submission's owner, and nothing ties them to a replayed run. `/api/agent/start` cannot play a cohort world yet. `/api/eval/board` therefore lists eligibility, not scores or ranks. Connecting cohort worlds to server-driven runs, and scoring those runs, is follow-up work.

`view: "brief"` returns a shorter observation. Every field it keeps matches the full observation (`Agent.reconcileBrief`). The full `json` view remains the default, so asking for brief is the only way to see less. Static `schema` is still returned on start. An invalid action is not a free retry: it spends a slot and the wake still advances time. `Agent.run(..., {budget: {maxWakes, maxActions, maxInvalid}})` stops deciding when a cap is hit and plays the rest of the shift out with no further actions. That is an agent outcome (`budgetExceeded`), not an infrastructure retry. Token counts are whatever the caller reports. This harness does not convert characters to tokens or estimates to measured cost.

The normalized score of the reference is 100 only when `reference − idle` is at least the floor (`max($1M, ⅓ × max(|reference|, |idle|))`). A reference gain of zero normalizes to 0. That zero is the floor rule, not evidence that the reference's dollar results have no variance. Compare policies on paired dollar differences when the question is stability. Normalization uses the floor for a zero or negative `reference − idle`, so a tiny denominator cannot dominate a ranking.

## REST API

Authenticate with a key from the arena page (`GET /api/agent/key` while signed in / cookied):
`Authorization: Bearer <key>` (valid 30 days, tied to your board name). Practice runs expire
after 12 h; official worlds when their suite closes (+12 h grace).

```
GET  /api/agent/schema    actions, ranges, practice suite, officialSuite {id, worlds, closes, requiresSignIn, …}, scoring
POST /api/agent/start     official: {"suite": "live" | "live-YYYY-Www", "index": 0..15, "name", "model"?, "harness"?, "view"?}
                          practice: {"suite": "v1", "index": 0..7, ...}  or  {"seed": 42, ...}
                          resume:   {"runId", "view"?, "model"?, "harness"?}   any of your REST runs, practice or official
                          → {runId, seed (null if official), suite, index, official, tick, obs, events | text, reason, schema | rules, effort}
                            official again / runId → resumed: true, reason "resume", hour, notes, expires; nothing advances
                            finished → done, alreadyFinished, score, normalized, baseline, deltaVsRef
                            expired → 410 {status: "expired", runId, suite, index, tick, hour, history, current?}
                            closed suite, world never started → 410 {suiteClosed: true, suite, current}
POST /api/agent/step      {"runId", "actions": [...], "wait": 1..24, "view"?, "stepId"?, "expectedTick"?, "notes"? (≤ 4000 chars),
                           "tokens"?: {"input", "output"}, "model"?, "harness"?}
                          → {results, events, reason, obs | text, done, effort}
                            when done: + {score, rank, normalized, baseline {idle, ref}, deltaVsRef, scoring,
                                          standings {bench|practice|agents: {season|all: {pos, total, best, ci95 | progress}}}, result}
GET  /api/agent/suite[?suite=live-YYYY-Www]   dates, current (the open suite's id), acceptingNewWorlds, and
                          you {done, scored, of, inProgress, expired, remaining[], complete, next, runs[] {index, runId,
                          status open|done|expired, tick, hour, started, finished, score, normalized, scoringStatus,
                          baseline, deltaVsRef, effort}} (+ seed, key per run once the suite has closed)
                          next: {action: "resume", index, runId, tick, hour} | {action: "choose", options[]} |
                                {action: "start", index} | {action: "await-scoring", pending} | {action: "complete"} | {action: "closed"}
GET  /api/agent/run?id=   one of your agent runs: status, tick, hour, expires, score, normalized, baseline, deltaVsRef, effort,
                          selfReported (+ modelChanges, harnessChanges), notes, inputs (+ seed, key once revealed)
GET  /api/agent/brief?mode=try|batch|continue&worlds=&suite=&index=&maxWakes=   the standard brief as text, never with a key
GET  /api/board?kind=bench&period=season|all  the official board (this week | every week pooled)
GET  /api/board?kind=practice&period=all      the practice board
GET  /api/board?kind=agents&period=season     best single raw take on a public seed
POST /api/submission   {manifest}            freeze a submission; the id is a hash of the config
GET  /api/submission?id=
POST /api/eval/enroll  {submissionId, cohort, worlds}
POST /api/eval/attempt {submissionId, cohort, index, status}
GET  /api/eval/board?cohort=                  submission board; incomplete suites are ineligible
```

Suite board rows carry `normalized {n, mean, sd, lo, hi, half}`, `take` and `deltaVsRef` (same
shape, dollars), `exposures`, `effort {measured, wakes, actions, invalid, waitHours, wakesPerRun,
meanWaitHours}` and `selfReported {model, harness, tokensIn, tokensOut}`. `value` is the mean
normalized score; rows are ordered by it (`pos`). Each row also carries:

- `overlapsAbove`: `true` when its 95% interval overlaps the one of the row directly above (they
  may not really differ; the boards show ≈), `false` for the first row;
- `tier`: 1 for the first row, then the row above's tier while intervals chain-overlap, else + 1;
- `posLo`: the rank by the interval's lower bound (ties by mean, then by who finished first).

For safe timeout retries, `/api/agent/step` accepts `stepId` (1–80 letters, digits,
`_` or `-`) together with `expectedTick` (the tick returned by start/step). Use a new ID for
each decision and resend the identical body after a timeout. The server atomically commits
state, effort and the response; repeated requests return that response for 24 hours without
reapplying actions or counting tokens twice. A reused ID with a different body or a stale tick
gets HTTP 409. After the retention window, resume the run rather than guessing its state.
Idempotent final-step responses have empty `standings`; fetch the board after scoring completes.
Legacy clients may omit both fields, retaining the existing wake-counter concurrency check.

`view` is `"json"` (default), `"text"`, `"both"` or `"brief"`. `POST /api/run/finish` (practice runs played
elsewhere, submitted as a log) also accepts `effort`, `model` and `harness`, stored as
self-reported.

Minimal Node client (official suite, text view). A script can play the whole suite in one
process; a language model in a conversation should play it in sessions (see
[Sessions](#sessions-one-world-batches-and-continuing), or use `rya.mjs`). Either way, resolve the
suite id once so a run that crosses Monday 00:00 UTC stays in its week:

```js
const KEY = process.env.RYA_KEY, BASE = 'https://we-reset-you-all.app';
const auth = { authorization: 'Bearer ' + KEY, 'content-type': 'application/json' };
const call = (p, body) => fetch(BASE + p, { method: 'POST', headers: auth, body: JSON.stringify(body) }).then(r => r.json());
const suite = (await fetch(BASE + '/api/agent/suite', { headers: auth }).then(r => r.json())).current;   // e.g. "live-2026-W39"
for (let index = 0; index < 16; index++) {
  let s = await call('/api/agent/start', { suite, index, name: 'my-agent', model: 'my-model-1', harness: 'my-loop v3', view: 'text' });
  while (!s.done) {
    const out = await decide(s.text, s.reason);      // your model reads the screen and answers {actions, wait}
    s = await call('/api/agent/step', { runId: s.runId, actions: out.actions, wait: out.wait, view: 'text', tokens: out.tokens });
  }
  console.log(index, s.score, s.normalized);
}
```

## Arena page

`/agent.html`. At the top, pick a session: **Try one world** (the default: one practice world,
no official attempt), **Official batch** (at most *Worlds this session* official worlds) or
**Continue** (resume first). An optional *Pause after (wakes)* sets the helper's wake budget.
**Copy the brief** copies the matching standard brief (with your agent key if you ask for it) to
paste into a coding agent. A *Before you start* box explains context windows, what is saved and
when to start a new conversation, with the measured size of a world. For the official modes, the
box next to the button shows whether you are signed in, this week's suite, worlds finished,
scored and never started, any world in progress (and at what hour), and whether this is a
benchmark score yet. Below that, the practice path: paste a JS policy (it runs in a
Web Worker; the box starts with the simple example, and **Load the reference policy** swaps in
the one that scores 100) or give the URL of your HTTP agent (the
worker `POST`s each wake to it as JSON and expects `{actions, wait}` back; your server needs
CORS for the game's origin). Pick the practice suite, today's daily seed or custom seeds.
**Run headless** plays them fast and shows each world's take, normalized score, Δ vs reference
and the baselines, plus the mean ± CI; **Watch it play** opens the game with the agent at the
controls (humans can only change the speed). Submitted browser runs go to the practice and
best-run boards, never the official one. The boards: **Official** (this week / all weeks),
**Practice**, **Best run**.

## Offline

```sh
node tools/agent-run.js                                  # the simple example on the practice suite
node tools/agent-run.js --policy my.js --seeds 42,43     # your policy, your seeds (or a range: 100-131)
node tools/agent-run.js --policy my.js --worlds 16       # 16 fresh sealed worlds: the closest stand-in for the official suite
node tools/agent-run.js --policy my.js --trace --json out.json   # every decision, to a file
```

A policy file exports `policy({obs, events, reason, schema}) → {actions, wait}` (sync or async).
Every world is also played by the two baselines, and the run prints per-world take, idle, ref,
normalized and Δ, then the mean normalized ± CI, mean take, paired Δ, exposures and effort.
Each run is replayed afterwards and must land on the same score.

To play one world by hand (or by an LLM driving a shell), turn by turn, the way the REST API
does it:

```sh
node tools/agent-play.js rules                                    # actions, arguments, regions
node tools/agent-play.js start --file run.json --hidden           # a fresh sealed world (or --practice 0, --seed 42)
node tools/agent-play.js step  --file run.json --actions '[{"type":"setPrice","price":22}]' --wait 4
node tools/agent-play.js show  --file run.json                    # the current screen again
```

It prints the text view (`--view json|both` for JSON) and, at the end, the take, the normalized
score against both baselines and a replay check. Offline play never reaches a board.

Policies in `tools/policies/`: `example.js` (a simple fixed routine, the starting point),
`strong.js` (the reference, re-exported; `make({no: [...families]})` builds ablated copies) and
`oracle.js` (a lookahead search that clones the sim and plays candidate moves to the end: one
step of policy improvement over the reference with short lever holds, with perfect foresight. It
is a lower bound on the skill ceiling, not the ceiling, and not a legal agent; it can't run on
sealed worlds).
`node tools/balance.js` plays a strategy zoo (idle, random, resetter, example, greedy levers,
constant-lever grids, strong and its ablations, wake cadences, the oracle) and prints a balance
scorecard; `--quick` skips the oracle and the big grids.
