Agent control of a real iPhone
Provides the screen as text and lets an AI agent tap, type, and scroll through WebDriverAgent, including in apps without APIs.
Loading page…
Coding / product dossier
Lets AI agents operate a real iPhone via HTTP or MCP and replay recorded flows without a model.
Product brief
iphone-use gives an AI agent a real iPhone: the screen as text, tap, type and scroll over WebDriverAgent, with an honest answer for every action (applied, not sent, or unknown). Payment apps that block screenshots come back as a wireframe. A task done once becomes a flow that replays with no model. HTTP API, a 21-tool MCP server for Claude Code, and remote control from a browser or iOS app. Open source, MIT.
Why we selected it
Concrete, MIT-licensed infrastructure for operating real iPhones, including apps without APIs. Action-status reporting, screenshot-blocking wireframes, and model-free flow replay provide technical substance beyond a chat
Product preview saved with our daily selection
Capability scan
Only capabilities supported by the product information we collected are listed here.
Provides the screen as text and lets an AI agent tap, type, and scroll through WebDriverAgent, including in apps without APIs.
Reports each action as applied, not sent, or unknown.
Returns a wireframe for payment apps that block screenshots.
Turns a task completed once into a flow that can replay without a model.
Offers an HTTP API, a 21-tool MCP server for Claude Code, and remote control from a browser or iOS app.
Best-fit use cases
FAQ
Yes. It gives an AI agent access to a real iPhone, with screen text and tap, type, and scroll actions through WebDriverAgent.
Each action receives one of three statuses: applied, not sent, or unknown.
The description states that a task completed once can become a flow that replays without a model.
Yes. It is open source under the MIT license.