Skip to content

Measured results

The eval harness was built before the agent, which is the only reason any of these numbers exist. The last full measurement, 13 tasks × 3 runs on gpt-5.6-sol:

success 39/39
observations 41, against an oracle floor of 39
actions 123, 1.14x a hand-written oracle
model turns 216
faults perception 6, policy 6, model 0
refusals, unusable runs 0, 0
cost $1.50 over 6m39s

The cost is priced at $4 in and $20 out per million tokens. It was first published as $1.87, priced at Claude Opus rates by a harness that ignored the price set beside the model; the token counts, and every ratio between arms, were unaffected.

The suite has since grown to 19 tasks, four of them planting an instruction in the screen the agent reads. Run for docs/adr/0015, 3 runs each on the same model, it passed 55/57 at 1.30x the oracle’s actions for $2.61, and obeyed no planted instruction in 12 tries. That run recorded passes, actions and cost, not observations or faults, so the table above stays the one to read for those.

One observation per run, give or take two across the whole set, including the two tasks in an app Apple did not write. That is the floor, and it holds because every action already folds the screen it produced into its response.

Actions were 1.25x the oracle until ios_find let the agent read the raw accessibility tree rather than only the digest built from it. The gain is entirely in the two tasks that call it, turns 43 to 27 and actions 10 to 3, against 195 to 189 and 125 to 120 for the eleven that never do. It is not free: a tenth verb costs about 190 prompt tokens on every turn of every run, used or not, which is docs/adr/0011 and the reason there is no eleventh.

Model turns are measured because they were the one axis left: grouping several actions into a turn changes no device work at all. Invited to do it, the model never once did, so docs/adr/0010 rejects the idea and keeps the guard that bounds the path it would have run on.

Verified on real iOS, including a physical iPhone

Section titled “Verified on real iOS, including a physical iPhone”

Tier 1 runs against a scripted in-process device, so its numbers are a claim about a fake. The same goal, turn on Bold Text, across all three tiers:

actions observations digest
scripted fake 3 1 n/a
iOS 27.0 simulator 4 1 166 raw nodes → 15 elements, 272 tokens
iPhone, iOS 26.6, Wi-Fi 3 1 140 raw nodes → 15 elements, 243 tokens

One observation on every tier, which is the number the design argument rests on, and on the phone it took 48.6s where the simulator took seconds. The simulator row was re-measured on iOS 27.0 after the 26.5 runtime was removed; its extra action is one model run choosing a longer route, not a capability the tier lacks. The phone row still reads 26.6 because the device runner’s provisioning profile has expired, so tier 3 cannot currently be re-run.

The switch was confirmed by navigating there and reading value="1" independently of what the agent claimed, then restored.

Most importantly, a real no-op still reports screen_changed=False on the phone. If a physical device had moved its fingerprint between settled snapshots, the verification step would have been silently dead on hardware while every simulator and fake test stayed green.

Hardware is opt-in twice over, by the device marker and IOS_MCP_ALLOW_DEVICE=1, because hardware being present is not consent to change settings on it.