Skip to content
screen: Preferences "Display & Text Size" e1 button "Accessibility" e2 switch "Bold Text" =0 e3 button "Larger Text" agent tap(e2) result screen_changed=True e2 switch "Bold Text" =1 251 raw nodes to 12 elements 50 to 474 tokens per step

Give an AI agent an iPhone. It checks every step it takes.

Use an iOS Simulator or a real iPhone over a cable or Wi-Fi, from the terminal, Claude Code, any MCP client, or Python. Every action returns the screen it produced, so the agent knows when a tap did nothing.
Terminal window
uv sync && uv run ios-agent quickstart
uv run ios-agent "turn on bold text"
ios-agent answering a question by driving Apple Maps

One goal, start to finish, at 2.5x: 4 actions, 1 observation, 2,767 device tokens, 49.8s of real time. The terminal is the agent’s own transcript; the phone is an iOS Simulator being driven by it.

Hand an agent a goal

ios-agent "turn on bold text" and watch it work: reasoning, the screen it reads, actions and cost, live in the terminal.

Give your MCP client hands

One entry in .mcp.json gives Claude Code, Cursor or any MCP client 31 tools for a simulator or a phone.

Use your own iPhone

Over a cable or Wi-Fi, verified on an iPhone 17 Pro Max, with anything risky asking first.

Drive it by hand

ios-agent manual needs no API key, and shows exactly what an agent would see on an app nobody has pointed this at before.

39/39agent runs succeeded
1.14xa hand-written oracle’s actions
1observation per run, the floor
$1.50for all 39 runs

13 tasks, three runs each, against an oracle whose action counts are asserted so they cannot drift. Read the results.

Four decisions, and everything follows from them

Section titled “Four decisions, and everything follows from them”

Raw WebDriverAgent page source for a 200-row list runs to roughly 37,000 tokens. Re-reading that after every tap exhausts a context window in a handful of steps.

Perception is budget-aware

Screens arrive as a compact digest, not accessibility XML: 251 raw nodes to 12 elements on a real third-party screen, 50 to 474 tokens per step.

Actions return the screen they produced

Half the round-trips, and a delta when the screen is similar, so a long flow stays cheap.

Resolution runs on the host

Six tiers re-find a stale ref by identity, so a retry costs zero model tokens where a round-trip costs a whole turn.

The gate asks before acting

Send, Pay, Delete, and anything that reaches another person need approval before they happen, while the answer still means something.