Skip to content
Eric-huang799Public

Repository files navigation

Duplex

build Glama awesome-mcp-servers freemcp.space

Listed in awesome-mcp-servers · freemcp.space · Glama

Duplex mascot

One browser shared by a human and an AI — the human sees the rendered page, the AI reads the DOM and page source. Same tabs, same live session, at the same time.

English | 中文

Duplex start page, dark theme

Duplex start page (dark theme). A light theme is also included.

Demo

Duplex demo — the AI searches Bilibili for a paper explainer, opens the video and plays it, while the side panel mirrors every step

The AI drives: it searches Bilibili, opens the video and starts playback — the side panel streams every step.

Full demo — 3 min, bilingual subtitles, with music:

Click to play the full demo

One continuous walkthrough: Act 1 — the built-in agent searches Bilibili and plays a paper-explainer video; Act 2 — Kimi browses to Claude and asks what CUDA is, then summarizes the answer; feature highlights — multi-model APIs, skills, external agents, emergency stop; and where to get it.

Controlling Excel for the web with Duplex

Controlling Microsoft Excel for the web — no dedicated spreadsheet-editing skill installed; only the stock API and a few simple automation skills 😅

A feasibility demo, not a recommended workflow. With no task-specific skill to lean on, Duplex worked with the stock API plus a few simple automation skills alone — improvising clipboard read/write channels, DOM probing and the like to push a whole grade sheet into Excel for the web. It got there, but the process wasted a lot of time and tokens. 😅

⚠️ Usage notes

  • Security boundaries (high stakes — please read). Duplex lets the AI drive your real browser session, including login state, cookies and local data. Under this architecture, a wrong AI move can touch real accounts and data (sending messages, submitting forms, modifying or deleting content), potentially with serious consequences. Don't leave the agent running unattended in environments where sensitive accounts are signed in. The emergency-stop hotkeys (default F2 / Ctrl+Shift+K, customizable in settings) are a last-resort human brake — they are no substitute for your own judgement about what the AI should be allowed to touch.

  • You and the AI share one page — and you come first. While you are scrolling, typing or interacting, the AI keeps observing but pauses its own modifications on that page instead of fighting you for the same input. A status bar lists the pages under human control, each with a one-click resume for when you're ready. The emergency stop remains the hard brake: one key cuts every AI action, the CLI processes it started, and all queued messages.

What's new in v0.2.9

  • Human-first collaboration — the AI yields to you. When you scroll, type, tap or focus something, it keeps reading and thinking but pauses its own modifications on that page; a status bar shows which pages are under human control and hands each one back with a click. Your scroll position and focus are preserved even while the AI acts.
  • Playwright under the hood — page actions now drive the browser's own Chromium through the debugging protocol: Playwright locators survive DOM changes, iframes and open shadow DOM work, waits and script execution are cancellable, and AI input is marked so the app can tell it apart from yours.
  • Task-tab binding & message integrity — each AI task is bound to its tab: switching tabs no longer redirects the AI, and closing the task tab reports an error instead of silently acting on another page. Panel messages and annotations carry their source tab / document / target session; injections are claimed once (60 s lease + renew + ACK), so nothing leaks to the wrong session and retries don't duplicate.
  • Session restore — on exit Duplex saves your tabs, their order and the active page; the next launch brings them back (blank pages included; local file: / data: pages are not persisted).
  • Sharper stops — stopping a built-in task cancels its model run, tools and pending confirmations; the emergency stop additionally terminates CLI children started by the app and clears queued injections.

Also in this build (from the 0.2.6 batch): fully customizable shortcuts, the complete bookmark manager, native context menus, the in-flow find bar, the F2 / Ctrl+Shift+K emergency stop, and honest annotation delivery.

How Duplex compares

Duplex browser-use Browser MCP Playwright MCP AI browsers (Atlas / Comet)
Form Desktop browser (Electron app) Python automation framework MCP server (browser extension) MCP server (Microsoft) Closed-source product
Who uses the browser Human and AI share the same tab and the same live session AI only (separate automation instance) AI drives your current Chrome AI only (Playwright instance) AI assistant alongside/operating
What the AI sees DOM outline snapshot + source + screenshots Vision + DOM Screenshots + a11y tree Accessibility tree Internal
Human collaboration Real-time side-by-side; the emergency-stop key interrupts the AI anytime Logs afterwards Human spectates Human spectates Limited intervention
When both act on one page The AI yields while you interact; per-page pause + one-click resume — — — —
Connectable AI Built-in models + opencode / Codex / Claude Code / Gemini / Qwen / any MCP client Bring your own LLM Any MCP client Any MCP client Official model only
Conversation visibility Live side-panel mirror (including external CLIs' chats and tool calls) Logs/terminal In the client In the client In-app
Data Fully local Local/cloud Local Local Cloud

In one line: browser-use / Playwright MCP let the AI run a flow for you; Duplex lets the AI use the browser together with you — same tab, same session, and when you both reach for the page, you go first.

Highlights

  • A real browser — tabs, address bar with search (Baidu / Bing / Google), back / forward / reload, loading state, themes, a wallpaper start page, customizable keyboard shortcuts and a full bookmark manager (folders, edit, sort, search).
  • Human-first collaboration — the AI works on its own bound tab, reads while you interact, pauses its page modifications when you take over, and hands pages back one by one. Same tab, same live session — without stepping on each other.
  • 24 MCP tools, Playwright-backed — snapshot compresses any page into a compact DOM outline with [eN] refs; clicks, typing, drags and reads run through Playwright locators (stable across DOM changes, iframe & open-shadow capable, cancellable waits). The rest covers tabs, navigation, upload, console logs, JS evaluation and page annotations.
  • Zero-setup bridge — a stdio MCP bridge (mcp-bridge) auto-launches the browser on the first tool call. Works with opencode, Claude Code, or any MCP client.
  • Live session mirror — when your AI works through opencode, its replies, reasoning and tool-call cards stream into the side panel in real time. Type in the panel to inject a message into the same session.
  • AI action visualization — a translucent cursor, element highlight and a status bar ("AI is clicking «…» — F2 to take over") are drawn in a Shadow-DOM overlay, so you always see what the AI is doing on the page.
  • Emergency stop — press the configured emergency-stop key (default F2 / Ctrl+Shift+K) or click the status bar to take control instantly: running tool calls are aborted, external processes started from the panel are terminated, pending confirmations are denied and queued messages are dropped. The pause survives a restart — send a message or click "Resume" to continue.
  • Page annotations — press Ctrl+Shift+A (or the ✎ button in the toolbar, or use annotation_mode) to draw a box / circle / arrow / point on any page and attach a question. The annotation is compiled into a structured text brief (DOM outline + visible text + selectors + geometry) and sent to the AI — with delivery receipts, and stale annotations are rejected.
  • Built-in agent (optional) — connect any OpenAI-compatible API (DeepSeek, Kimi, Qwen, GLM, Ollama, …) and let the browser drive itself. Provider management supports one-click import from opencode.
  • Conversation history & session restore — built-in agent sessions are saved locally and can be reopened any time; browser tabs and the active page come back after a restart.

Feature tour

1. Start page

Dark start page

The start page shows a clock, date and search box over a wallpaper that follows the active theme (light and dark wallpapers switch automatically with the system).

2. Built-in agent drives the browser

Built-in agent in action

Ask the built-in model in the side panel — in Chinese or English. It calls tools (navigate, snapshot, …), with each call shown as a card, and reports back in the panel while you watch the page change.

3. opencode session mirror

opencode session mirror

When your AI runs through opencode, its messages and tool calls are mirrored into the panel while it drives the same visible browser. The "opencode" / "built-in" tabs at the top of the panel switch between the two ways of working.

4. Session picker

Session picker

Pick which opencode session the panel is connected to. "Auto" follows the most recent conversation.

5. Model providers

Model providers

Manage OpenAI-compatible providers for the built-in agent: add, edit, delete, or import providers directly from your opencode configuration. Local models via Ollama (http://localhost:11434/v1) work out of the box.

6. Conversation history

Conversation history

Built-in agent conversations are stored locally and can be reopened at any time.

7. AI action visualization and emergency stop

AI action visualization

Every AI action is drawn on the page: a cursor ring, an element highlight and a status bar. Press the emergency-stop key (default F2) at any moment to take the browser back — the AI stops immediately and waits for your instruction.

8. Page annotations

Page annotations

Draw a box (or circle / arrow / point) around anything and ask a question about it. The annotation is converted into a structured text brief for the AI, including the DOM outline of the region, visible text, selectors and geometry — so even text-only AIs can "see" what you mean.

Quick start

Installer (Windows / macOS / Linux)

  1. Download the latest installer from Releases: Duplex Setup x.y.z.exe (Windows), .dmg (macOS, Apple Silicon / Intel) or .AppImage (Linux).
  2. Run it and launch Duplex.

Prefer no installer? A portable zip (unzip anywhere, run Duplex.exe) is attached to releases as well.

Build from source

npm install
npm run build          # Electron app -> out/
npm run build:bridge   # stdio MCP bridge -> dist-bridge/index.cjs

If Electron's binary download fails (e.g. behind a slow mirror), retry with: $env:ELECTRON_MIRROR="https://npmmirror.com/mirrors/electron/"; node node_modules/electron/install.js

Run it:

npm run dev    # development mode with HMR
# or
npm start      # preview the production build

Running a dev build next to the installed app? Set DUPLEX_DATA_DIR to another directory so the two don't share settings and the local endpoint.

Connect opencode (recommended)

  1. Install the mirror plugin. Copy integrations/opencode/plugins/cobrowse-mirror.ts to your global opencode plugin directory (~/.config/opencode/plugins/), or into .opencode/plugins/ of a project.
  2. Register the MCP server in your opencode config (opencode.json), project-level or global — see opencode.example.json:
{
  "mcp": {
    "duplex": {
      "type": "local",
      "command": ["node", "C:\\path\\to\\Duplex\\dist-bridge\\index.cjs"],
      "enabled": true
    }
  }
}
  1. In opencode, ask the AI to browse: "open example.com and describe the page". The browser launches automatically on the first tool call (no need to start it manually).
  2. In the Duplex side panel choose the opencode tab and select the session you want to follow (or keep "Auto").

Use the built-in model (no opencode required)

Panel → Built-in tab → Model settings → add an OpenAI-compatible provider (base URL, API key, model name), or click Import from opencode to reuse your existing provider configuration. For local models, point the base URL to Ollama, e.g. http://localhost:11434/v1.

MCP tools

Tool Description
list_tabs / new_tab / close_tab / switch_tab Tab management — human and AI share the same tabs
navigate / search / history Open a URL (auto-searches if not a URL), search (Baidu default, or Bing / Google), back / forward / reload
snapshot Page as text — compact DOM outline; interactive elements carry [eN] refs
get_html / query Raw HTML, or detailed info for elements matching a CSS selector
screenshot PNG screenshot of the visible page (for multimodal models)
click / dblclick / hover Click / double-click / hover through Playwright; target accepts an eN ref or a CSS selector, and the action waits for the element to be actionable
type / press Type text (optionally submitting) and press keys, incl. combos like Control+A
drag Real drag & drop from one point/element to another
select_option / upload Native <select> options, and file upload by absolute path
scroll Scroll the page or a specific element into view
wait Wait for time / selector / text — cancellable; yields with a reason while you are interacting
get_console Read page console logs (errors and warnings)
annotation_mode Enter/exit the annotation overlay (human draws a box/circle/arrow/point + question)
evaluate Evaluate JS in the page, returns JSON-serializable results

How it works

Architecture

  • The Electron main process serves a small local HTTP API on 127.0.0.1 protected by a per-launch bearer token (endpoint info is written to ~/.cobrowse/endpoint.json).
  • dist-bridge/index.cjs is a stdio MCP server that proxies tool calls to that API and auto-launches the app when it is not running.
  • Page automation is Playwright over the browser's own Chromium: the app reads DevToolsActivePort, matches targetId to each tab and drives locators, iframes, open shadow DOM, cancellable waits and script evaluation — the same page you are looking at, not a separate automation instance.
  • A collaboration layer binds every task to its tab, serializes same-page writes, watches for real human input (wheel / touch / keyboard / focus) and pauses AI modifications per page; the status bar reflects it and offers one-click resume per page.
  • The opencode plugin pushes session events (text, reasoning, tool calls) into the panel and long-polls for queued messages; deliveries are claimed once (60 s lease, renewed, ACK'd) and target the bound session rather than "whichever session is active".
  • Panel messages and page annotations are injected with identity: annotations capture the source tab, URL and a document marker, and stale ones are rejected — they cannot cross into the wrong session.
  • Codex and Claude Code are mirrored from their local session transcripts (history + live tail); replies sent from the panel headlessly resume the same session (codex exec resume / claude --resume) and the new turns stream back through the same tail. Launches are matched by structured session IDs, not "most recently modified file".
  • Browser state (tab order, HTTP(S) URLs, the active page) is saved on exit and restored on the next start.
  • External scripts can also push messages into the session: POST /api/chat { "text": "..." }.

Known limitations

  • While you are interacting with a page, AI input on it is synthesized DOM events (isTrusted=false); sites that depend on real key/mouse events, contenteditable, or custom widgets may need manual coordination — and side effects already applied cannot be undone.
  • Cancellation granularity: in-flight Playwright calls are bounded by 200–250 ms polling — an action that already completed (click, submit, write) cannot be rolled back.
  • evaluate cleans up timers / RAF / fetch inside its scope, but cannot revoke callbacks registered elsewhere (e.g. via document.defaultView); advanced uses are being constrained further.
  • Element refs are snapshot-scoped: after navigation or large DOM changes take a fresh snapshot (stale refs fail with a clear message instead of clicking something wrong).
  • Direct MCP clients share a default caller identity; an explicit task start / end / handoff protocol is planned for 0.3.0.
  • opencode message de-duplication relies on the plugin; if it stays disconnected longer than the lease, an accepted-but-unacknowledged message may be delivered twice — exact dedupe needs upstream message IDs.
  • Custom CLIs without a structured launch identity are accepted only when a unique new transcript appears; ambiguous cases are refused rather than guessed.
  • Built-in multimodal image context, site-specific custom widgets and cross-platform interaction have not been fully verified yet.
  • The built-in agent is a convenience option: local / smaller models are noticeably less reliable at long tool-use chains than a full opencode setup.
  • CLI tools that open pages through the OS (Claude Code, Codex, …) follow the system default browser. To route them to Duplex, pick Duplex once in your OS settings — ⋯ → 设为默认浏览器… opens that page (it is never forced). CLIs and commands launched from Duplex also get a BROWSER=duplex-open shim, which covers tools that honour $BROWSER.

Roadmap

  • v0.3.0 — collaboration, continued. Next up:
    • An explicit task-lifecycle / handoff protocol for MCP callers (start, end, restore) and finer-grained cancellable action steps.
    • Fallbacks for contenteditable and custom widgets; a richer CLI session lifecycle; multimodal verification.
    • Community backlog: command palette, tab search, reader mode, bookmark HTML import / export, and more — tracked in docs/待办与用户反馈.md.

Development

npm run dev        # dev mode (electron-vite, HMR for the renderer)
npm run typecheck  # TypeScript checks (node + web)
npm test           # unit tests (vitest, 300+ cases)
npm run smoke      # end-to-end smoke test (launches the bridge and a real browser)
npm run dist       # build the Windows installer (electron-builder)

Self-contained smokes (a hidden Electron plus a temporary data directory — they never touch your own sessions):

node tests/playwright-connection-smoke.mjs   # the Playwright layer against a hidden fixture browser
node tests/private-029-smoke.mjs             # 11 main-process collaboration scenarios

Debug helpers:

  • GET /api/debug/ui-snapshot — captures the current window to ~/.cobrowse/ui-snapshot.png.
  • POST /api/debug/panel-eval — runs JS in the panel renderer (only when the app is started with COBROWSE_DEBUG_UI=1).
  • POST /api/debug/ui-action — drives panel actions (used by tests, e.g. {"action":"panel-mode:opencode"}).
  • Logs: ~/.cobrowse/app.log (main + renderer log lines).

A note from the author

I'm not a professional developer — just a regular computer user who got curious about AI. Duplex is a hobby project I've been building in my spare time, learning as I go (with AI coding tools lending a hand).

It is still rough. There are bugs I haven't found, designs that will annoy you, and rough edges I simply didn't notice. If you're willing to give it a try, any feedback means a great deal to me: a bug, a crash, a confusing step, or simply "it didn't start for me" — please open an Issue. English or Chinese, either is fine.

Thank you for reading this far — and for giving it a shot.

License

MIT

Releases

Packages

Contributors

Languages