Skip to content

Decision records

Why things are the way they are, and what would change them back. Written when a decision was actually made, so the numbering is chronological rather than grouped.

# Decision Settled by
0001 The agent is a peer of the server, not a layer above it reading the import graph
0002 No explicit planner 1.05x an oracle on long tasks, with no planner
0003 No cross-session memory hedged framing measured worse than none; assertive stopped the agent checking
0004 No subagents 5,864 tokens/run against a 1M window
0005 LangGraph core, not the deepagents harness three of its four pillars were later rejected on their own numbers
0006 The agent calls IosSession directly +1.6 ms and 0 tokens per call over MCP, so the reason is architectural rather than performance
0007 Report unreachable content rather than reaching it snapshot cost is XCTest’s floor and observations are already at the oracle’s, so the constraint is completeness
0008 The terminal front end is a third distribution the agent’s seven-module public surface excludes the two modules a front end needs
0009 The eval trend is committed, and CI checks it rather than writing it device tokens, observations, actions and tiers are byte-deterministic on a scripted device, so a band would only license drift
0010 No multi-action batching, though its guard stays invited to batch, the model did it in 0 of 24 turns, at +4.2% turns and +25% cost
0011 A find that reads the tree the digest threw away -9.2% turns and 70% fewer perception faults, concentrated entirely in the two tasks that called it
0012 No per-app skill files on a route built to need one, both arms ran at the oracle’s floor, 4 actions and 6 turns, three times out of three
0013 No defence against instructions planted in screen content a labelled bait named by the text beside it, with the gate disarmed, was taken 0 times in 24; reopened on a disguised payload, a goal that acts, and a late bait, 0 in 18 more
0014 Ask before reaching another person a real run liked a stranger’s profile four times unasked; on a real Settings app, one false positive in 178 labels, removed by judging only the pressed control
0015 Route routine turns to a small model, as an off-by-default option pre-registered; all three rules held, cost at 48.9% against a 50% bar, which is a tie; the book’s 15% did not hold
0016 No Accessibility Inspector tree source a focus walk at 60 to 80 ms per element; Settings root 3.3 to 4.4s against WDA’s 3.7s, with no rects, and it scrolls the screen as it reads
0017 No simulator-native tree source CoreSimulator read the same screens 10 to 23% faster than WDA against a 50% bar; both wait on the app serialising its tree, and on Xcode 27 the read bootstraps through XCTest anyway
0018 Start the USB tunnel without sudo; Wi-Fi stays on xcodebuild go-ios’s userspace tunnel launched 10 of 10 in 3.1 to 3.3s with no sudo; pymobiledevice3’s native tunnel launched faster, then refused six in a row after a replug, because it and remoted evict each other by design
0019 Settle on frames, as an off-by-default option 17 to 22% off a phone’s settle time, returning early 1 time in 304 against the tree loop’s 3 in 198; only after counting novel frames, since the caret and a 0.42s search pause each defeated “identical frames”
0020 No App Intents action tier Settings declares no intents; typed Siri on a phone answered two App Shortcut phrases itself, with a dialog and a how-to, and reported success both times; taps are not the bottleneck at 1.14x the oracle
0021 Not claimed: exploring an app for its bugs pre-registered at 80%; 18 of 30 self-evident bugs found (60%), every one on the forced path and none off it, one false report, $0.30 to $1.10 a run against #450’s $4.28
0022 Cache the prompt prefix pre-registered at 60% of the same runs priced uncached; not yet run
0023 Tracing, as an off-by-default option 894 spans over 22 replayed tasks added 54 ms in total, 0.06 ms a span, with the model shown identical bytes in all 220 replays; off because a trace sends what the redactor does not scrub off the machine
0024 Count false success claims pre-registered as a measurement; 0 false successes in 103 claims on gpt-6.1-sol; 3 under-claims on one task traced to an action diff that ignored labels, and fixed

Eleven of the twenty-four are refusals. That is the point rather than an accident: the eval harness was built before the agent so it could overturn the design, and it did, on the first run and repeatedly afterwards.

Each record states what would reopen it. Most of them come down to a harder task set: longer horizons, an unfamiliar third-party app, or a screen the agent cannot read without vision.