Safety
The controls, and their limits. What they defend against, and which configuration a given setup is in, is in the threat model.
Driving a person’s phone with their real accounts is a different risk from test automation. The tap that dismisses a dialog in CI can send a message, make a payment, or delete a photo library here. The policy layer is on by default and sits in front of every action.
Approval before the fact
Section titled “Approval before the fact”Actions are classified before they run, so approval is asked while the operation is still preventable rather than reported afterwards.
Two questions are asked, each with its own switch:
- Does it destroy data or cost money?
policy.destructive_labels, such as send, pay, delete, sign out. Checked against the control’s label, its id, and any text being typed. - Does it reach another person?
policy.person_labels, such as like, follow, comment, reply, share. Checked against the control being pressed only. Typing is not judged, because nothing typed reaches anyone until it is sent, and a paragraph of static text is nobody’s button. The prompt says the action “would reach another person”, so the person approving knows the consequence rather than only that a rule matched.
The second exists because the first missed a real one: an agent asked to read a dating profile liked it four times, and nothing asked, because a like destroys nothing and costs nothing. See ADR 0014.
Matching is on whole words. “Sender” and “Undelete” do not trip the “send” and “delete” rules, because a gate that prompts on everything trains an operator to approve reflexively, which is worse than no gate.
Two modes:
- The MCP client supports elicitation: the human is asked directly, with the action and target named.
- It does not, or the question cannot be delivered: the call raises
action_requires_approvalcarrying a signature, and nothing happens on the device. The caller confirms with the user, then repeats the call withapprove=<signature>. This is the path an external human-in-the-loop layer uses.
Nothing gated runs until someone has said yes. An unanswerable question is not consent, and it is not a refusal either: the caller is told nobody was asked, rather than that the user declined.
Approval is scoped to one specific action. Approving Send does not approve Delete, and a refusal is never cached as consent.
Secrets
Section titled “Secrets”ios_type_secret takes a reference, not a value:
ios_type_secret(secret_ref="icloud-password", ref="e4")The value is read from the host keychain and sent straight to the device. It appears in no prompt, no tool result, and no audit entry. It is also deliberately not run through the destructive-text rules, since a password containing the word “delete” is not an instruction.
A password field shows dots, so the screen never holds the value. Any other
field shows it, and an app may repeat it (“No results for …”). So from the
moment a secret is typed, the session replaces it with [secret] in every
screen, search result, read and audit entry it returns, for the rest of the
session. Two limits: a value shorter than four characters is not scrubbed,
because removing every occurrence of it would wreck the screen, and a
screenshot is a picture, which nothing here edits. Type secrets into password
fields.
Store one with:
security add-generic-password -s ios-mcp -a icloud-password -wor set IOS_MCP_SECRET_ICLOUD_PASSWORD in the environment.
Never put a real credential in ios_type, where it would enter the transcript.
Apps holding payment or credential data are blocked by default
(policy.app_blocklist). An allowlist may be set instead, in which case
everything else is refused.
Redaction
Section titled “Redaction”An accessibility tree contains whatever is on screen, which on a real phone
means message bodies, card numbers, and email addresses. Card-like numbers and
email addresses are stripped from what a session hands out: digests, action
results including any alert they carry, finds, text reads, clipboard reads and
the audit trail. Card numbers are matched whether printed contiguously or
grouped by spaces or hyphens. Patterns are configurable via
policy.redact_patterns.
This is applied inside the session, so it holds for every consumer: the MCP server, the bundled agent and the terminal front end alike. It used to live at the server’s boundary, which is why the agent that ships with this project was not redacted at all. Device logs, which only the server reads, are redacted there.
One exit is left raw on purpose: IosSession.alert() returns the alert as
WebDriverAgent reported it, because alert buttons are matched by their labels
and nothing this project ships reads it directly. An alert that reaches the
agent does so inside an action result, which is redacted.
Redaction is text only. A screenshot is returned as captured, because redacting an image
needs a model of where things are on it and this project does not have one. An
earlier version of this page documented a redact_screenshots setting; it was
never read by anything, and the configuration ignores unknown keys, so setting
it changed nothing and said nothing.
Stopping
Section titled “Stopping”- Repeated consecutive failures halt the session rather than letting an agent flail at a screen it does not understand.
- A detected loop halts it too: if actions keep landing on the same few screens, the agent is stuck. Observations do not count, since re-reading a screen is careful behaviour, not thrashing.
ios_haltstops a session immediately;ios_resumeclears it once a human has decided it is safe.
Every session records an ordered list of what it did: tool, arguments,
resolution tier, screen fingerprint, outcome. ios_export_trace returns it.
This serves three purposes at once: explaining what happened, replaying a run
as a regression test, and providing worked examples for a future agent.
Configuration
Section titled “Configuration”[policy]enabled = trueconfirm_destructive = trueconfirm_reaching_a_person = trueapp_allowlist = ["com.apple.Preferences"] # empty means "anything not blocked"max_consecutive_failures = 5Or IOS_MCP_POLICY__CONFIRM_DESTRUCTIVE=false in the environment, or in a
.env beside the repository. See .env.example, which documents every setting
and leaves each safety default switched on, so copying it unedited changes
nothing.
Turning the gate off is a deliberate choice and a reasonable one for a simulator running a test suite. It is not reasonable for a device carrying someone’s real accounts.
What this does not protect against
Section titled “What this does not protect against”The gate is a heuristic over labels and roles. It will not recognise a destructive action whose control is unlabelled or misleadingly named, and it cannot know that tapping a particular row costs money. Approval mode and an app allowlist are the real controls; the label rules are a convenience on top.