iphone-use
Product Hunt 郭立
iphone-use lets AI agents control a real iPhone, not just an app’s API. It can replay tasks and even handles locked-down payment screens as wireframes.
Based on reporting by Product Hunt, 郭立 — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
iphone-use is a new open-source project that gives an AI agent a real iPhone to work with. Instead of talking to an app through an API, it reads the screen as text and can tap, type, and scroll through WebDriverAgent.
The project leans hard into honesty about what the agent actually did. Every action gets a clear result: applied, not sent, or unknown. That sounds small, but it matters when you’re trying to debug automation that can fail in messy, half-visible ways.
Some apps make life awkward on purpose. Payment apps that block screenshots are handled as wireframes, so the agent still has something to reason over. And once a task has been done once, it can be turned into a flow that replays without needing the model again.
There’s also a broader toolset around it. iphone-use exposes an HTTP API, ships with a 21-tool MCP server for Claude Code, and can be controlled remotely from a browser or an iOS app. It’s released under MIT, which should make it easier for people to poke at, extend, and break in public.
My take — AI-written commentary, not fact-checked reporting
This is the right kind of open source: practical, a little scrappy, and aimed at the annoying parts of real automation instead of the demo-friendly bits. Closed systems love to wave at “agents”; open tools like this have to admit when they’ve merely tapped, failed, or stared at a wireframe. That honesty is the feature.
Read more about this at: Product Hunt