Skip to content

Safety

The controls, and their limits. What they defend against, and which configuration a given setup is in, is in the threat model.

Driving a person’s phone with their real accounts is a different risk from test automation. The tap that dismisses a dialog in CI can send a message, make a payment, or delete a photo library here. The policy layer is on by default and sits in front of every action.

Actions are classified before they run, so approval is asked while the operation is still preventable rather than reported afterwards.

Two questions are asked, each with its own switch:

  • Does it destroy data or cost money? policy.destructive_labels, such as send, pay, delete, sign out. Checked against the control’s label, its id, and any text being typed.
  • Does it reach another person? policy.person_labels, such as like, follow, comment, reply, share. Checked against the control being pressed only. Typing is not judged, because nothing typed reaches anyone until it is sent, and a paragraph of static text is nobody’s button. The prompt says the action “would reach another person”, so the person approving knows the consequence rather than only that a rule matched.

The second exists because the first missed a real one: an agent asked to read a dating profile liked it four times, and nothing asked, because a like destroys nothing and costs nothing. See ADR 0014.

Matching is on whole words. “Sender” and “Undelete” do not trip the “send” and “delete” rules, because a gate that prompts on everything trains an operator to approve reflexively, which is worse than no gate.

Two modes:

  • The MCP client supports elicitation: the human is asked directly, with the action and target named.
  • It does not, or the question cannot be delivered: the call raises action_requires_approval carrying a signature, and nothing happens on the device. The caller confirms with the user, then repeats the call with approve=<signature>. This is the path an external human-in-the-loop layer uses.

Nothing gated runs until someone has said yes. An unanswerable question is not consent, and it is not a refusal either: the caller is told nobody was asked, rather than that the user declined.

Approval is scoped to one specific action. Approving Send does not approve Delete, and a refusal is never cached as consent.

ios_type_secret takes a reference, not a value:

ios_type_secret(secret_ref="icloud-password", ref="e4")

The value is read from the host keychain and sent straight to the device. It appears in no prompt, no tool result, and no audit entry. It is also deliberately not run through the destructive-text rules, since a password containing the word “delete” is not an instruction.

A password field shows dots, so the screen never holds the value. Any other field shows it, and an app may repeat it (“No results for …”). So from the moment a secret is typed, the session replaces it with [secret] in every screen, search result, read and audit entry it returns, for the rest of the session. Two limits: a value shorter than four characters is not scrubbed, because removing every occurrence of it would wreck the screen, and a screenshot is a picture, which nothing here edits. Type secrets into password fields.

Store one with:

Terminal window
security add-generic-password -s ios-mcp -a icloud-password -w

or set IOS_MCP_SECRET_ICLOUD_PASSWORD in the environment.

Never put a real credential in ios_type, where it would enter the transcript.

Apps holding payment or credential data are blocked by default (policy.app_blocklist). An allowlist may be set instead, in which case everything else is refused.

An accessibility tree contains whatever is on screen, which on a real phone means message bodies, card numbers, and email addresses. Card-like numbers and email addresses are stripped from what a session hands out: digests, action results including any alert they carry, finds, text reads, clipboard reads and the audit trail. Card numbers are matched whether printed contiguously or grouped by spaces or hyphens. Patterns are configurable via policy.redact_patterns.

This is applied inside the session, so it holds for every consumer: the MCP server, the bundled agent and the terminal front end alike. It used to live at the server’s boundary, which is why the agent that ships with this project was not redacted at all. Device logs, which only the server reads, are redacted there.

One exit is left raw on purpose: IosSession.alert() returns the alert as WebDriverAgent reported it, because alert buttons are matched by their labels and nothing this project ships reads it directly. An alert that reaches the agent does so inside an action result, which is redacted.

Redaction is text only. A screenshot is returned as captured, because redacting an image needs a model of where things are on it and this project does not have one. An earlier version of this page documented a redact_screenshots setting; it was never read by anything, and the configuration ignores unknown keys, so setting it changed nothing and said nothing.

  • Repeated consecutive failures halt the session rather than letting an agent flail at a screen it does not understand.
  • A detected loop halts it too: if actions keep landing on the same few screens, the agent is stuck. Observations do not count, since re-reading a screen is careful behaviour, not thrashing.
  • ios_halt stops a session immediately; ios_resume clears it once a human has decided it is safe.

Every session records an ordered list of what it did: tool, arguments, resolution tier, screen fingerprint, outcome. ios_export_trace returns it. This serves three purposes at once: explaining what happened, replaying a run as a regression test, and providing worked examples for a future agent.

[policy]
enabled = true
confirm_destructive = true
confirm_reaching_a_person = true
app_allowlist = ["com.apple.Preferences"] # empty means "anything not blocked"
max_consecutive_failures = 5

Or IOS_MCP_POLICY__CONFIRM_DESTRUCTIVE=false in the environment, or in a .env beside the repository. See .env.example, which documents every setting and leaves each safety default switched on, so copying it unedited changes nothing.

Turning the gate off is a deliberate choice and a reasonable one for a simulator running a test suite. It is not reasonable for a device carrying someone’s real accounts.

The gate is a heuristic over labels and roles. It will not recognise a destructive action whose control is unlabelled or misleadingly named, and it cannot know that tapping a particular row costs money. Approval mode and an app allowlist are the real controls; the label rules are a convenience on top.