Skip to content

Repository files navigation

ChromeRPC

Writing gRPC adapters for https://chromedevtools.github.io/devtools-protocol/ (definition at eg https://source.chromium.org/chromium/chromium/src/+/main:third_party/blink/public/devtools_protocol/domains/Page.pdl) in a way that's compatible with the rest of our rpc tooling.

The dream:

./runrpc Stream.captureScreenshotRequest pages.binarypb | ./runrpc Page.captureScreenshot > screenshots.binarypb

Starting Out

Milestone1: able to send a grpc to headless multiclient (https://developer.chrome.com/blog/new-in-devtools-63/#multi-client) chrome to:

We're writing our binaries in Go, and setting up a common linker in https://github.com/accretional/rpcfun - invest as little as possible in main.go, we want to basically just define and implement grpc services with one service per "domain" per directory, one .go implementation of that service per directory.

Might be worth using https://github.com/bitfield/script to chain commands/convert to http calls.

HeadlessBrowser Automation

The HeadlessBrowserService is a high-level automation layer built on top of the CDP domain services. Instead of wiring together individual gRPC calls, you define automation as a sequence of steps in a text proto file, then execute the whole sequence with a single RPC.

Quick Start

  1. Start the server:
make run   # launches headless Chrome + gRPC on :50051

Or connect to an existing Chrome instance with remote debugging enabled:

# Get the WebSocket URL from a running Chrome
WS_URL=$(curl -s http://127.0.0.1:9222/json/version | python3 -c \
  "import sys,json; print(json.load(sys.stdin)['webSocketDebuggerUrl'])")

./bin/chromerpc --ws-url "$WS_URL" --port 50051
  1. Write an automation file (my_automation.textproto):
name: "screenshot_example"

steps: {
  label: "set_viewport"
  set_viewport: {
    width: 1280
    height: 800
    device_scale_factor: 2
  }
}

steps: {
  label: "navigate"
  navigate: {
    url: "https://example.com"
  }
}

steps: {
  label: "wait_for_render"
  wait: {
    milliseconds: 500
  }
}

steps: {
  label: "capture"
  screenshot: {
    output_path: "screenshot.png"
    format: "png"
  }
}
  1. Run it:
go run ./cmd/automate -input my_automation.textproto

Available Step Types

Step Description Key Fields
set_viewport Set browser viewport size width, height, device_scale_factor, mobile
navigate Navigate to a URL url
wait Pause for a fixed duration milliseconds
screenshot Capture the visible page as an image output_path, format (png/jpeg), quality, full_page
full_page_screenshot Capture the entire scrollable page output_path, format, quality
record Record an audio+visual capture of the tab to a video file output_path, audio_path, pre_delay_ms, max_duration_ms, stop_condition, start_script, output_fps
evaluate_script Run JavaScript in the page expression
click Click at coordinates or a CSS selector x, y, selector
type_text Insert text into a focused element or selector text, selector
type_key_by_key Type text character-by-character with realistic delays text, delay_ms, selector
press_key Press a special key (Enter, Tab, Escape, arrows, etc.) key
wait_for_selector Wait until a CSS selector appears in the DOM selector, timeout_ms
reload Reload the current page ignore_cache
scroll_to Scroll to coordinates x, y
open_tab Open a URL in a new browser tab url
switch_tab Switch CDP session to a different tab target_id
close_tab Close a browser tab target_id
download_file Download a file via browser-native download url, output_path

RPCs

The service exposes two RPCs:

service HeadlessBrowserService {
  // Run a full sequence of steps.
  rpc RunAutomation(AutomationSequence) returns (AutomationResult);
  // Run a single step (for orchestrators that need to branch on results).
  rpc ExecuteStep(AutomationStep) returns (StepResult);
}

RunAutomation executes a linear sequence and stops on first failure. ExecuteStep runs one step at a time, returning the result so the caller can make decisions (e.g., extract links from a page, then open each in a loop). This makes it possible to build complex orchestrators as standalone Go programs that call ExecuteStep in a loop.

Multi-Tab Support

The open_tab, switch_tab, and close_tab steps enable multi-tab workflows. When you open a new tab, the returned StepResult.script_result contains the target ID. Pass this to switch_tab to route subsequent commands to that tab, and close_tab to clean up.

open_tab(url) → target_id
switch_tab(target_id) → session_id (commands now go to this tab)
... do work in the tab ...
close_tab(target_id) → tab destroyed

The server manages CDP sessions internally via Target.attachToTarget with flatten=true.

File Downloads

The download_file step handles browser-native downloads. It opens the URL in a new tab, sets Browser.setDownloadBehavior to auto-save to the output directory, finds and clicks the download button (supporting pdf.js viewer's #download button, generic download buttons, and <a download> links), then waits for the file to appear on disk. This preserves the browser's cookies and session, avoiding issues with authenticated or CDN-protected resources.

Audio + Visual Recording

The record step captures an audio+visual recording of the current tab to a video file (WebM/MP4) on the server.

  • Video is captured with Page.startScreencast — real rendered frames, so it works on any page (DOM, SVG, canvas) in headless Chrome. Frames arrive at the page's actual repaint rate, which in headless is bounded by page weight and how often the page repaints (a media page driven by ontimeupdate repaints only a few times/second; a light page screencasts faster).
  • Audio is muxed in from a caller-supplied source file (audio_path) with ffmpeg. Headless Chrome exposes no way to capture the tab's own audio over CDP — there is no tab-audio-capture command, screencast is video-only, and getDisplayMedia/chrome.tabCapture need a GUI/extension/user gesture or a real audio device. Muxing the source audio is exact (not a lossy re-recording) and is the natural model whenever you already own the audio you fed the page. Omit audio_path for a video-only clip.

Lifecycle: enable Page → wait pre_delay_ms → run start_script (e.g. begin playback) → screencast → collect frames until a stop condition: max_duration_ms elapses, stop_condition (a polled JS expression) returns truthy, or the request context is cancelled (e.g. the interactive bidi stream disconnects — the partial recording is still encoded and written). The step returns a JSON summary in script_result: {"output","bytes","frames","video_seconds","has_audio","stop":"max_duration|stop_condition|disconnect"}.

Requirements: ffmpeg on the server's PATH, and — to let start_script begin <audio>/<video> playback without a user gesture — run the server with --autoplay (adds --autoplay-policy=no-user-gesture-required). Paths in output_path/audio_path are on the server's filesystem, so this is best on the local/bidi path rather than the stateless Cloud Run deployment. Container format is inferred from the extension (.mp4 → H.264/AAC, .webm → VP9/Opus).

{ "record": {
    "output_path": "out/clip.mp4",
    "audio_path": "source.wav",
    "pre_delay_ms": 400,
    "max_duration_ms": 12000,
    "start_script": "document.querySelector('audio,video').play()",
    "stop_condition": "document.querySelector('audio,video').ended",
    "output_fps": 30
} }

See recipes/record_av.textproto.

Modularity

Automations are plain text proto files (AutomationSequence messages). This means you can:

  • Reorder steps by moving steps: { ... } blocks around.
  • Compose sequences by concatenating multiple .textproto files or merging them with tooling.
  • Version control your automations alongside code — they're human-readable diffs.
  • Extend with new step types by adding a new action to the AutomationStep oneof in proto/cdp/headlessbrowser/headlessbrowser.proto and implementing the handler in internal/server/headlessbrowser/headlessbrowser.go.

Example Automations

See the automations/ directory for ready-to-use text proto files.

Connecting to an Existing Chrome

For sites with bot detection, you can connect to a real (non-headless) Chrome instance:

# Launch Chrome with remote debugging
"/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" \
  --remote-debugging-port=9222 &

# Connect chromerpc to it
WS_URL=$(curl -s http://127.0.0.1:9222/json/version | python3 -c \
  "import sys,json; print(json.load(sys.stdin)['webSocketDebuggerUrl'])")
./bin/chromerpc --ws-url "$WS_URL"

The server includes --disable-blink-features=AutomationControlled by default and supports --user-agent overrides.

Low-Level CDP Interface (all domains)

HeadlessBrowserService is the high-level, ergonomic surface. Underneath, every Chrome DevTools Protocol domain is also exposed as its own gRPC service — cdp.page.PageService, cdp.runtime.RuntimeService, cdp.network.NetworkService, cdp.dom.DOMService, cdp.input.InputService, cdp.target.TargetService, and ~50 more. Each method maps 1:1 to a CDP command (e.g. PageService.CaptureScreenshot, NetworkService.SetCookie, RuntimeService.Evaluate), and PageService alone has 46 methods.

gRPC reflection is enabled, so you can discover the full surface without any local .proto files. Capabilities are methods, not services — listing services only gives you the ~55 domains; the actual ~480 commands live in the methods inside each service. Discovery has three levels:

grpcurl $ADDR list                                          # 1. all ~55 services (domains)
grpcurl $ADDR list cdp.network.NetworkService               # 2. that domain's methods = capabilities
grpcurl $ADDR describe cdp.network.NetworkService.SetCookie # 3. a method's request/response
grpcurl $ADDR describe cdp.network.SetCookieRequest         #    a message's fields

To dump every method across every domain (the whole capability list):

for s in $(grpcurl $ADDR list | grep '^cdp\.'); do grpcurl $ADDR list "$s"; done

(Locally use grpcurl -plaintext localhost:50051 …; on Cloud Run add -H "$AUTH" and target $HOST:443 — see below.)

When to use which:

  • RunAutomation (high-level) — recommended default. Self-contained sequences, and each call is isolated in its own browser context, so it's safe under concurrency.
  • Low-level domain services — full, fine-grained CDP power for things the step types don't cover. Note these currently share the process-wide default session and are not per-call isolated, so they're best for single-session or local use (or low, careful concurrency). For most automation, prefer RunAutomation, and reach for the domains when you need a specific CDP command.

Running on Cloud Run

chromerpc ships as a self-contained container (Go server + bundled google-chrome-stable) and runs on Google Cloud Run as a hosted gRPC service with reflection enabled. Full details — image/version tagging, IAM, scaling, and the dev workflow — are in DEPLOY.md.

gcloud auth login
gcloud config set project <YOUR_PROJECT>
make deploy            # Cloud Build -> Artifact Registry -> Cloud Run (IAM-gated)

Key properties of the deployed service:

  • gRPC over TLS/HTTP2 on :443 (Cloud Run terminates TLS at the edge).
  • IAM-gated — every call needs a bearer identity token; unauthenticated callers get 403. (Deploy publicly with INVOKER_AUTH=allow only if you accept that a public browser-automation endpoint is effectively an open fetch/SSRF proxy.)
  • Per-call isolation — each RunAutomation/ExecuteStep runs in its own incognito browser context, so concurrent calls don't share cookies/storage.
  • Scale-to-zero, concurrency 8 — the first call after idle pays a cold start (Chrome launch, a few seconds); subsequent calls are fast.

Calling the deployed service

Set the endpoint and a token once (the principal must have roles/run.invoker):

HOST=$(gcloud run services describe chromerpc --region us-central1 \
        --format='value(status.url)'); HOST=${HOST#https://}
TOKEN=$(gcloud auth print-identity-token)
AUTH="authorization: Bearer ${TOKEN}"

Discover the API (reflection is on — no local .proto needed):

grpcurl -H "$AUTH" $HOST:443 list
grpcurl -H "$AUTH" $HOST:443 describe cdp.headlessbrowser.HeadlessBrowserService

Run an automation and save the screenshot (PNG bytes come back inline in stepResults[].screenshotData, base64):

grpcurl -H "$AUTH" -d '{
  "steps": [
    { "navigate": { "url": "https://example.com", "wait_until": "networkidle" } },
    { "screenshot": { "format": "png" } }
  ]
}' $HOST:443 cdp.headlessbrowser.HeadlessBrowserService/RunAutomation \
  | jq -r '.stepResults[]|select(.screenshotData).screenshotData' \
  | base64 -d > out.png && open out.png

Run a saved recipe (handles textproto→JSON→call→save/open for you):

HOST=$HOST ./scripts/recipe-run.sh recipes/search_and_screenshot.textproto

See recipes/ for reusable playbooks (load-then-screenshot, search, dismiss-consent, scroll-to-lazy-load) and recipes/README.md for the building blocks.

Grant another caller access:

gcloud run services add-iam-policy-binding chromerpc --region us-central1 \
  --member 'user:[email protected]' --role roles/run.invoker

Calling from code (Go)

Any gRPC client works — point it at HOST:443 with TLS and attach a bearer identity token per call:

import (
    "crypto/tls"
    "google.golang.org/grpc"
    "google.golang.org/grpc/credentials"
    "google.golang.org/grpc/credentials/oauth"
    "google.golang.org/api/idtoken"
    pb "github.com/accretional/chromerpc/proto/cdp/headlessbrowser"
)

audience := "https://" + host // the Cloud Run service URL
ts, _ := idtoken.NewTokenSource(ctx, audience)        // service-account or ADC creds
conn, _ := grpc.NewClient(host+":443",
    grpc.WithTransportCredentials(credentials.NewTLS(&tls.Config{})),
    grpc.WithPerRPCCredentials(oauth.TokenSource{TokenSource: ts}))
defer conn.Close()

client := pb.NewHeadlessBrowserServiceClient(conn)
res, err := client.RunAutomation(ctx, &pb.AutomationSequence{
    Steps: []*pb.AutomationStep{
        {Action: &pb.AutomationStep_Navigate{Navigate: &pb.Navigate{
            Url: "https://example.com", WaitUntil: "networkidle"}}},
        {Action: &pb.AutomationStep_Screenshot{Screenshot: &pb.Screenshot{Format: "png"}}},
    },
})
// res.StepResults[i].ScreenshotData holds the PNG bytes.

For other languages: open a TLS channel to :443 and send an authorization: Bearer <id-token> metadata header on each RPC, using the service URL as the token audience.

Treat every call as a self-contained, isolated session — don't rely on state (cookies, navigation, open tabs) persisting across calls; put the whole flow (navigate → act → screenshot) in one RunAutomation.

Chrome Testing

The chrome-testing/ folder contains a self-contained screenshot testing module. It handles the full lifecycle — building chromerpc, launching Chrome, serving HTML, capturing PNGs, and tearing down — in a single script.

./chrome-testing/snap.sh my-page.html screenshots/my-page.png

See chrome-testing/USAGE_INSTRUCTIONS.md for the full guide, including how any other project can copy this folder and use it for their own visual validation.

Example output

chromerpc demo

Resources / Notes

Nodejs implementaiton of the chrome remote interface: https://github.com/cyrus-and/chrome-remote-interface

VERY USEFUL: entire browser_protocol.json for the chrome remote interface https://github.com/ChromeDevTools/devtools-protocol/blob/master/json/browser_protocol.json

https://buf.build/docs/reference/descriptors/#what-are-descriptors this could be useful for converting individual domains or commands into .protos programmatically via https://github.com/protocolbuffers/protobuf/blob/main/src/google/protobuf/descriptor.proto and https://pkg.go.dev/google.golang.org/protobuf/reflect/protoreflect and https://github.com/jhump/protoreflect/tree/main/protoprint

About

gRPC adapters for the Chrome DevTools Protocol

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages