Features
It knows when an action did not work
Section titled “It knows when an action did not work”The most common complaint about agent tools for phones is an action that reports success while nothing happened. Here:
- Every action returns the screen it produced, with
screen_changed, so a tap that did nothing is visible at once. A real no-op reportsscreen_changed=Falseon a physical iPhone too, not only on a simulator. - Typed text is read back from the field. Text that did not land returns
ok: falsewith what the field shows, rather than being reported as typed. - The agent names elements, never coordinates. It passes a ref like
e2; if the screen moved, the host re-finds the same element by identity through six tiers, and refuses when a ref now points at something else. - Switches are set, not toggled. Asking for
onwhen a switch is already on does nothing, instead of turning it off. - Retries are safe. Idempotency keys mean an agent framework that replays a step does not tap Send twice.
It works on a real iPhone, over a cable or Wi-Fi
Section titled “It works on a real iPhone, over a cable or Wi-Fi”- Verified on an iPhone 17 Pro Max on iOS 26.6, over USB and over Wi-Fi, at the same action count as the simulator.
ios-agent doctorchecks the toolchain, the runtimes, the tunnel and the runner’s signing expiry, and gives a remedy for each failure.- The device runner stops with the process that started it, a SIGKILL
included, so a crashed run does not hold the phone.
ios-mcp resetclears anything else. - The device picker never pre-selects a physical phone, and
/deviceswitches devices mid-session.
It is cheap on tokens
Section titled “It is cheap on tokens”- Screens arrive as a compact digest, not accessibility XML: 251 raw nodes to 12 elements on a real third-party screen, 50 to 474 tokens per step across thirteen golden flows.
- When the next screen is similar, an action returns only what changed.
- A 300-row Contacts list comes back in one observation of 438 tokens.
ios_findsearches the full tree for anything the digest left out, so compaction never hides a control for good.- The bundled agent looks at the screen once per run, because every action already hands back the screen it produced.
It asks before anything risky
Section titled “It asks before anything risky”- Send, Pay, Buy, Delete, Confirm, Sign Out, and anything that reaches another person (Like, Follow, Share, Message) need approval before they happen. With no one to ask, they are refused.
- Passwords come from the Mac’s keychain and go straight to the device. They
never enter a prompt or the audit trail, and once typed, every screen the
session returns shows
[secret]in their place. - Card numbers and email addresses are redacted before any client sees the screen, and every action is recorded in an exportable audit trail.
It keeps working when Xcode updates
Section titled “It keeps working when Xcode updates”Everything goes through XCTest, the framework Apple ships for UI testing. Faster routes through private frameworks were measured and turned down, and Xcode 27 broke several tools that took them. See Why only Apple’s public APIs.
Three ways in, and any model
Section titled “Three ways in, and any model”- A terminal app that streams the model’s reasoning beside the screen it is reading, with actions, tokens and cost on screen as they climb.
- An MCP server: 31 tools, 5 resources and an
ios_operatorprompt, over stdio or HTTP. See the tool reference. - A Python library,
IosSession, that the server and the agent both sit on, so your own agent can skip the protocol. - Anthropic, OpenAI, Azure OpenAI, Gemini, Vertex AI, Bedrock, Groq, Mistral or a local Ollama model, switched by two environment variables.
It is measured, and you can rerun the numbers
Section titled “It is measured, and you can rerun the numbers”39 of 39 agent runs succeeded at 1.14x the actions a hand-written oracle needs, for $1.50 in total. Golden flows track tokens, time and how each element was found, and CI fails if the free series moves without a recorded reason. See Measured on real hardware.
What it does not do
Section titled “What it does not do”- No Android, no other platforms, no builds or profiling.
- Flutter canvases and WebViews have no tree to read. The digest says so and points at a screenshot, rather than returning a screen that looks empty. See ADR 0007.
- The approval gate is not a defence against instructions planted in a screen. See ADR 0013.
- A physical iPhone needs a signed runner, and a free Apple ID’s profile lasts seven days.