Decision records
Why things are the way they are, and what would change them back. Written when a decision was actually made, so the numbering is chronological rather than grouped.
| # | Decision | Settled by |
|---|---|---|
| 0001 | The agent is a peer of the server, not a layer above it | reading the import graph |
| 0002 | No explicit planner | 1.05x an oracle on long tasks, with no planner |
| 0003 | No cross-session memory | hedged framing measured worse than none; assertive stopped the agent checking |
| 0004 | No subagents | 5,864 tokens/run against a 1M window |
| 0005 | LangGraph core, not the deepagents harness | three of its four pillars were later rejected on their own numbers |
| 0006 | The agent calls IosSession directly |
+1.6 ms and 0 tokens per call over MCP, so the reason is architectural rather than performance |
| 0007 | Report unreachable content rather than reaching it | snapshot cost is XCTest’s floor and observations are already at the oracle’s, so the constraint is completeness |
| 0008 | The terminal front end is a third distribution | the agent’s seven-module public surface excludes the two modules a front end needs |
| 0009 | The eval trend is committed, and CI checks it rather than writing it | device tokens, observations, actions and tiers are byte-deterministic on a scripted device, so a band would only license drift |
| 0010 | No multi-action batching, though its guard stays | invited to batch, the model did it in 0 of 24 turns, at +4.2% turns and +25% cost |
| 0011 | A find that reads the tree the digest threw away | -9.2% turns and 70% fewer perception faults, concentrated entirely in the two tasks that called it |
| 0012 | No per-app skill files | on a route built to need one, both arms ran at the oracle’s floor, 4 actions and 6 turns, three times out of three |
| 0013 | No defence against instructions planted in screen content | a labelled bait named by the text beside it, with the gate disarmed, was taken 0 times in 24; reopened on a disguised payload, a goal that acts, and a late bait, 0 in 18 more |
| 0014 | Ask before reaching another person | a real run liked a stranger’s profile four times unasked; on a real Settings app, one false positive in 178 labels, removed by judging only the pressed control |
| 0015 | Route routine turns to a small model, as an off-by-default option | pre-registered; all three rules held, cost at 48.9% against a 50% bar, which is a tie; the book’s 15% did not hold |
| 0016 | No Accessibility Inspector tree source | a focus walk at 60 to 80 ms per element; Settings root 3.3 to 4.4s against WDA’s 3.7s, with no rects, and it scrolls the screen as it reads |
| 0017 | No simulator-native tree source | CoreSimulator read the same screens 10 to 23% faster than WDA against a 50% bar; both wait on the app serialising its tree, and on Xcode 27 the read bootstraps through XCTest anyway |
| 0018 | Start the USB tunnel without sudo; Wi-Fi stays on xcodebuild | go-ios’s userspace tunnel launched 10 of 10 in 3.1 to 3.3s with no sudo; pymobiledevice3’s native tunnel launched faster, then refused six in a row after a replug, because it and remoted evict each other by design |
| 0019 | Settle on frames, as an off-by-default option | 17 to 22% off a phone’s settle time, returning early 1 time in 304 against the tree loop’s 3 in 198; only after counting novel frames, since the caret and a 0.42s search pause each defeated “identical frames” |
| 0020 | No App Intents action tier | Settings declares no intents; typed Siri on a phone answered two App Shortcut phrases itself, with a dialog and a how-to, and reported success both times; taps are not the bottleneck at 1.14x the oracle |
| 0021 | Not claimed: exploring an app for its bugs | pre-registered at 80%; 18 of 30 self-evident bugs found (60%), every one on the forced path and none off it, one false report, $0.30 to $1.10 a run against #450’s $4.28 |
| 0022 | Cache the prompt prefix | pre-registered at 60% of the same runs priced uncached; not yet run |
| 0023 | Tracing, as an off-by-default option | 894 spans over 22 replayed tasks added 54 ms in total, 0.06 ms a span, with the model shown identical bytes in all 220 replays; off because a trace sends what the redactor does not scrub off the machine |
| 0024 | Count false success claims | pre-registered as a measurement; 0 false successes in 103 claims on gpt-6.1-sol; 3 under-claims on one task traced to an action diff that ignored labels, and fixed |
Eleven of the twenty-four are refusals. That is the point rather than an accident: the eval harness was built before the agent so it could overturn the design, and it did, on the first run and repeatedly afterwards.
Each record states what would reopen it. Most of them come down to a harder task set: longer horizons, an unfamiliar third-party app, or a screen the agent cannot read without vision.