Drive a real iPhone from code — screenshot it, tap it, swipe it, type into it — so an AI agent can use the phone the way a person would.
rPlay mirrors an iPhone to a computer and forwards mouse, keyboard, and touch input back to it. This repository is the programmable half: the client SDKs, an MCP server that plugs into Claude Code and OpenAI Codex, and working agent samples for Claude, OpenAI, and Gemini.
from rplay_client import RPlayClient
c = RPlayClient()
dev = c.first_device()
open("screen.jpg", "wb").write(c.screenshot()) # what's on the phone
c.tap(dev["screen_size"]["w"] // 2, 400) # touch itEverything runs against a physical iPhone, not a simulator. No jailbreak, no app installed on the phone, no developer profile.
They are genuinely different underneath, and the difference matters when choosing:
| macOS | Linux | |
|---|---|---|
| How | TCP socket server inside rPlay for Mac | xdotool + mss acting on the mirror window |
| Client | macos/rplay_client.py |
linux/linux_iphone_sdk.py |
| Reference | docs/api-macos.md | docs/api-linux.md |
| Screenshot / tap | ✅ | ✅ |
| Swipe, type text, Home | ✅ | ✅ |
| Needs a display attached | no | yes |
| Moves your real mouse cursor | no | yes |
Both cover the same operations. macOS is the better-behaved of the pair —
no display required, and it never touches your cursor — so prefer it
where you have the choice. Write against the shared
screenshot / tap / swipe / type_text set and one agent serves
both.
-
Install rPlay for Mac and start a mirror session — Wi-Fi (AirPlay) or USB cable, either works.
-
Confirm the socket is live. The server binds when you click Start, not at app launch:
printf '{"method":"ping"}\n' | nc 127.0.0.1 9876
You want
{"ok":true,"result":{"pong":true}}. -
Run the smoke test:
cd macos && python3 01_smoke_test.py
It pings, lists devices, saves a screenshot, and taps the centre of the screen. No API key and no model needed — get this working before involving an LLM.
sudo apt install xdotool
pip install mss Pillow
cd linux && python3 linux_iphone_sdk.pywith airplaydemo / youcast_wd running and the iPhone mirroring. Full
setup in docs/api-linux.md.
The failure that costs people the most time is a tap that returns success and does nothing at all.
Absolute touch input only works while the iPhone has an Accessibility pointer feature turned on. iOS accepts the position-carrying HID reports either way, and silently discards them otherwise — no error, no cursor, no tap.
On the iPhone, turn on Settings ▸ Accessibility ▸ Zoom (preferred) or Settings ▸ Accessibility ▸ Touch ▸ AssistiveTouch. Zoom is the better choice for agent work: it does not leave a floating button on the screen where it can end up in your screenshots and confuse a vision model. After enabling Zoom, three-finger double-tap to turn magnification back off — the feature stays enabled, which is all that is required.
If taps still do nothing after that, reboot the iPhone. iOS's accessibility state wedges after repeated reconnects more often than you would expect, and no amount of debugging your own code will fix it.
Three ways to point a model at a phone, in increasing order of how much of the loop you own:
| What it is | Read | |
|---|---|---|
| MCP server | Six tools inside a Claude Code or OpenAI Codex session. No API key, no loop to write — just chat. | docs/mcp.md |
| Standalone agents | claude_ios_agent.py / openai_ios_agent.py / gemini_ios_agent.py — direct API, your own loop, lower latency, runs unattended. |
docs/agents.md |
| Minimal samples | macos/02_claude_agent.py, 03_gemini_agent.py, 04_openai_agent.py — ~150 lines each, the whole loop visible at once. |
source |
The three standalone agents are deliberately interchangeable — same six tools, same loop — so switching provider is a different script, not a different program.
The minimal samples are the ones to read first if you are writing your own. The loop is genuinely small: screenshot → model → tap → repeat.
cd macos
pip install anthropic # or: openai, google-genai
export ANTHROPIC_API_KEY=sk-ant-...
python3 02_claude_agent.py "open the camera app"
# same loop, different provider
export OPENAI_API_KEY=...
python3 04_openai_agent.py "open the camera app"There is no human in the loop. The agent keeps acting until the model
says it is done or the step budget runs out. Use --dry-run to see what
it would tap before letting it touch anything.
- Reaction games do not work. A vision model needs one to three seconds per turn. Anything needing reflexes is out.
- Turn-based and slow-paced work well — puzzles, form filling, navigating settings, walking a UI to reproduce a bug.
- Grid-precision is the weak point. Models land taps a few tens of
pixels off, which for a grid UI means the wrong cell entirely.
linux/blockblast_solver.pyshows the workaround: detect the grid geometry from the screenshot in ordinary code, and let the model reason in(row, col)rather than pixels.
linux/blockblast_agent.py is the fullest worked example — an agent that
plays a real game end to end, with the solver doing the geometry and the
model doing the strategy. Write-ups:
docs/blockblast.md and
docs/blockblast-solver.md.
macos/ rplay_client.py + four samples — the socket API
linux/ linux_iphone_sdk.py, the MCP server, agents, BlockBlast demo
docs/ API references, MCP and agent guides, the v1 design spec
- docs/api-macos.md — the macOS socket API. Four methods, the coordinate model, and the prerequisites that make taps actually land.
- docs/api-linux.md — the Python SDK, and the three constraints that come with driving a window from outside.
- docs/mcp.md — MCP server setup for Claude Code and OpenAI Codex.
- docs/agents.md — the three standalone agents, model choices, CLI flags, and how to race them against each other.
- docs/blockblast.md and docs/blockblast-solver.md — the worked end-to-end example, including how the grid geometry is solved in code so the model never has to guess pixels.
- docs/design-spec-v1.md — a WebSocket API with auth, grid helpers, and an event stream. Designed, not built. It is published for the reasoning, not to code against; the status table at the top says exactly what exists.
This is v0. The API is small and it will change.
Two things to know before you build on it:
- There is no authentication on the macOS socket. Any local process can screenshot your phone and tap anything on it. That is a deliberate tradeoff for a loopback-only developer tool, but do not leave a session running on a shared machine.
- Device IDs are positional (
device-0) and are not stable across sessions. Re-read them; never persist them.
Unimplemented methods return unsupported_method rather than failing
silently, so feature detection is a round trip.
The most useful gaps, roughly in order:
- An MCP server for the macOS API. Wrap
RPlayClientwith FastMCP the waylinux/mcp_ios_server.pywrapsIPhoneSDK. - The grid helpers from the design spec. They fix a real, measurable failure mode rather than adding surface area.
- The remaining
press_buttontargets — lock, volume, mute. The iAP consumer usages exist; they are just reachable only through function keys in the keyboard path today, not by name.
MIT — see LICENSE.