Writing gRPC adapters for https://chromedevtools.github.io/devtools-protocol/ (definition at eg https://source.chromium.org/chromium/chromium/src/+/main:third_party/blink/public/devtools_protocol/domains/Page.pdl) in a way that's compatible with the rest of our rpc tooling.
The dream:
./runrpc Stream.captureScreenshotRequest pages.binarypb | ./runrpc Page.captureScreenshot > screenshots.binarypbMilestone1: able to send a grpc to headless multiclient (https://developer.chrome.com/blog/new-in-devtools-63/#multi-client) chrome to:
-
captureSnapshot (https://source.chromium.org/chromium/chromium/src/+/main:third_party/blink/public/devtools_protocol/domains/Page.pdl;l=632-642)
-
captureScreenshot (https://source.chromium.org/chromium/chromium/src/+/main:third_party/blink/public/devtools_protocol/domains/Page.pdl;l=611-630)
-
maybe printToPdf (https://source.chromium.org/chromium/chromium/src/+/main:third_party/blink/public/devtools_protocol/domains/Page.pdl;l=940-998)
-
any other commands/infrastructure to get these working
We're writing our binaries in Go, and setting up a common linker in https://github.com/accretional/rpcfun - invest as little as possible in main.go, we want to basically just define and implement grpc services with one service per "domain" per directory, one .go implementation of that service per directory.
Might be worth using https://github.com/bitfield/script to chain commands/convert to http calls.
The HeadlessBrowserService is a high-level automation layer built on top of the CDP domain services. Instead of wiring together individual gRPC calls, you define automation as a sequence of steps in a text proto file, then execute the whole sequence with a single RPC.
- Start the server:
make run # launches headless Chrome + gRPC on :50051Or connect to an existing Chrome instance with remote debugging enabled:
# Get the WebSocket URL from a running Chrome
WS_URL=$(curl -s http://127.0.0.1:9222/json/version | python3 -c \
"import sys,json; print(json.load(sys.stdin)['webSocketDebuggerUrl'])")
./bin/chromerpc --ws-url "$WS_URL" --port 50051- Write an automation file (
my_automation.textproto):
name: "screenshot_example"
steps: {
label: "set_viewport"
set_viewport: {
width: 1280
height: 800
device_scale_factor: 2
}
}
steps: {
label: "navigate"
navigate: {
url: "https://example.com"
}
}
steps: {
label: "wait_for_render"
wait: {
milliseconds: 500
}
}
steps: {
label: "capture"
screenshot: {
output_path: "screenshot.png"
format: "png"
}
}- Run it:
go run ./cmd/automate -input my_automation.textproto| Step | Description | Key Fields |
|---|---|---|
set_viewport |
Set browser viewport size | width, height, device_scale_factor, mobile |
navigate |
Navigate to a URL | url |
wait |
Pause for a fixed duration | milliseconds |
screenshot |
Capture the visible page as an image | output_path, format (png/jpeg), quality, full_page |
full_page_screenshot |
Capture the entire scrollable page | output_path, format, quality |
record |
Record an audio+visual capture of the tab to a video file | output_path, audio_path, pre_delay_ms, max_duration_ms, stop_condition, start_script, output_fps |
evaluate_script |
Run JavaScript in the page | expression |
click |
Click at coordinates or a CSS selector | x, y, selector |
type_text |
Insert text into a focused element or selector | text, selector |
type_key_by_key |
Type text character-by-character with realistic delays | text, delay_ms, selector |
press_key |
Press a special key (Enter, Tab, Escape, arrows, etc.) | key |
wait_for_selector |
Wait until a CSS selector appears in the DOM | selector, timeout_ms |
reload |
Reload the current page | ignore_cache |
scroll_to |
Scroll to coordinates | x, y |
open_tab |
Open a URL in a new browser tab | url |
switch_tab |
Switch CDP session to a different tab | target_id |
close_tab |
Close a browser tab | target_id |
download_file |
Download a file via browser-native download | url, output_path |
The service exposes two RPCs:
service HeadlessBrowserService {
// Run a full sequence of steps.
rpc RunAutomation(AutomationSequence) returns (AutomationResult);
// Run a single step (for orchestrators that need to branch on results).
rpc ExecuteStep(AutomationStep) returns (StepResult);
}RunAutomation executes a linear sequence and stops on first failure. ExecuteStep runs one step at a time, returning the result so the caller can make decisions (e.g., extract links from a page, then open each in a loop). This makes it possible to build complex orchestrators as standalone Go programs that call ExecuteStep in a loop.
The open_tab, switch_tab, and close_tab steps enable multi-tab workflows. When you open a new tab, the returned StepResult.script_result contains the target ID. Pass this to switch_tab to route subsequent commands to that tab, and close_tab to clean up.
open_tab(url) → target_id
switch_tab(target_id) → session_id (commands now go to this tab)
... do work in the tab ...
close_tab(target_id) → tab destroyed
The server manages CDP sessions internally via Target.attachToTarget with flatten=true.
The download_file step handles browser-native downloads. It opens the URL in a new tab, sets Browser.setDownloadBehavior to auto-save to the output directory, finds and clicks the download button (supporting pdf.js viewer's #download button, generic download buttons, and <a download> links), then waits for the file to appear on disk. This preserves the browser's cookies and session, avoiding issues with authenticated or CDN-protected resources.
The record step captures an audio+visual recording of the current tab to a
video file (WebM/MP4) on the server.
- Video is captured with
Page.startScreencast— real rendered frames, so it works on any page (DOM, SVG, canvas) in headless Chrome. Frames arrive at the page's actual repaint rate, which in headless is bounded by page weight and how often the page repaints (a media page driven byontimeupdaterepaints only a few times/second; a light page screencasts faster). - Audio is muxed in from a caller-supplied source file (
audio_path) with ffmpeg. Headless Chrome exposes no way to capture the tab's own audio over CDP — there is no tab-audio-capture command, screencast is video-only, andgetDisplayMedia/chrome.tabCaptureneed a GUI/extension/user gesture or a real audio device. Muxing the source audio is exact (not a lossy re-recording) and is the natural model whenever you already own the audio you fed the page. Omitaudio_pathfor a video-only clip.
Lifecycle: enable Page → wait pre_delay_ms → run start_script (e.g. begin
playback) → screencast → collect frames until a stop condition:
max_duration_ms elapses, stop_condition (a polled JS expression) returns
truthy, or the request context is cancelled (e.g. the interactive bidi stream
disconnects — the partial recording is still encoded and written). The step
returns a JSON summary in script_result:
{"output","bytes","frames","video_seconds","has_audio","stop":"max_duration|stop_condition|disconnect"}.
Requirements: ffmpeg on the server's PATH, and — to let start_script
begin <audio>/<video> playback without a user gesture — run the server with
--autoplay (adds --autoplay-policy=no-user-gesture-required). Paths in
output_path/audio_path are on the server's filesystem, so this is best on
the local/bidi path rather than the stateless Cloud Run deployment. Container
format is inferred from the extension (.mp4 → H.264/AAC, .webm → VP9/Opus).
{ "record": {
"output_path": "out/clip.mp4",
"audio_path": "source.wav",
"pre_delay_ms": 400,
"max_duration_ms": 12000,
"start_script": "document.querySelector('audio,video').play()",
"stop_condition": "document.querySelector('audio,video').ended",
"output_fps": 30
} }See recipes/record_av.textproto.
Automations are plain text proto files (AutomationSequence messages). This means you can:
- Reorder steps by moving
steps: { ... }blocks around. - Compose sequences by concatenating multiple
.textprotofiles or merging them with tooling. - Version control your automations alongside code — they're human-readable diffs.
- Extend with new step types by adding a new action to the
AutomationSteponeof inproto/cdp/headlessbrowser/headlessbrowser.protoand implementing the handler ininternal/server/headlessbrowser/headlessbrowser.go.
See the automations/ directory for ready-to-use text proto files.
For sites with bot detection, you can connect to a real (non-headless) Chrome instance:
# Launch Chrome with remote debugging
"/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" \
--remote-debugging-port=9222 &
# Connect chromerpc to it
WS_URL=$(curl -s http://127.0.0.1:9222/json/version | python3 -c \
"import sys,json; print(json.load(sys.stdin)['webSocketDebuggerUrl'])")
./bin/chromerpc --ws-url "$WS_URL"The server includes --disable-blink-features=AutomationControlled by default and supports --user-agent overrides.
HeadlessBrowserService is the high-level, ergonomic surface. Underneath, every
Chrome DevTools Protocol domain is also exposed as its own gRPC service —
cdp.page.PageService, cdp.runtime.RuntimeService, cdp.network.NetworkService,
cdp.dom.DOMService, cdp.input.InputService, cdp.target.TargetService, and
~50 more. Each method maps 1:1 to a CDP command (e.g.
PageService.CaptureScreenshot, NetworkService.SetCookie,
RuntimeService.Evaluate), and PageService alone has 46 methods.
gRPC reflection is enabled, so you can discover the full surface without any
local .proto files. Capabilities are methods, not services — listing
services only gives you the ~55 domains; the actual ~480 commands live in the
methods inside each service. Discovery has three levels:
grpcurl $ADDR list # 1. all ~55 services (domains)
grpcurl $ADDR list cdp.network.NetworkService # 2. that domain's methods = capabilities
grpcurl $ADDR describe cdp.network.NetworkService.SetCookie # 3. a method's request/response
grpcurl $ADDR describe cdp.network.SetCookieRequest # a message's fieldsTo dump every method across every domain (the whole capability list):
for s in $(grpcurl $ADDR list | grep '^cdp\.'); do grpcurl $ADDR list "$s"; done(Locally use grpcurl -plaintext localhost:50051 …; on Cloud Run add
-H "$AUTH" and target $HOST:443 — see below.)
When to use which:
RunAutomation(high-level) — recommended default. Self-contained sequences, and each call is isolated in its own browser context, so it's safe under concurrency.- Low-level domain services — full, fine-grained CDP power for things the
step types don't cover. Note these currently share the process-wide default
session and are not per-call isolated, so they're best for single-session
or local use (or low, careful concurrency). For most automation, prefer
RunAutomation, and reach for the domains when you need a specific CDP command.
chromerpc ships as a self-contained container (Go server + bundled
google-chrome-stable) and runs on Google Cloud Run as a hosted gRPC service
with reflection enabled. Full details — image/version tagging, IAM, scaling, and
the dev workflow — are in DEPLOY.md.
gcloud auth login
gcloud config set project <YOUR_PROJECT>
make deploy # Cloud Build -> Artifact Registry -> Cloud Run (IAM-gated)Key properties of the deployed service:
- gRPC over TLS/HTTP2 on
:443(Cloud Run terminates TLS at the edge). - IAM-gated — every call needs a bearer identity token; unauthenticated
callers get
403. (Deploy publicly withINVOKER_AUTH=allowonly if you accept that a public browser-automation endpoint is effectively an open fetch/SSRF proxy.) - Per-call isolation — each
RunAutomation/ExecuteStepruns in its own incognito browser context, so concurrent calls don't share cookies/storage. - Scale-to-zero, concurrency 8 — the first call after idle pays a cold start (Chrome launch, a few seconds); subsequent calls are fast.
Set the endpoint and a token once (the principal must have roles/run.invoker):
HOST=$(gcloud run services describe chromerpc --region us-central1 \
--format='value(status.url)'); HOST=${HOST#https://}
TOKEN=$(gcloud auth print-identity-token)
AUTH="authorization: Bearer ${TOKEN}"Discover the API (reflection is on — no local .proto needed):
grpcurl -H "$AUTH" $HOST:443 list
grpcurl -H "$AUTH" $HOST:443 describe cdp.headlessbrowser.HeadlessBrowserServiceRun an automation and save the screenshot (PNG bytes come back inline in
stepResults[].screenshotData, base64):
grpcurl -H "$AUTH" -d '{
"steps": [
{ "navigate": { "url": "https://example.com", "wait_until": "networkidle" } },
{ "screenshot": { "format": "png" } }
]
}' $HOST:443 cdp.headlessbrowser.HeadlessBrowserService/RunAutomation \
| jq -r '.stepResults[]|select(.screenshotData).screenshotData' \
| base64 -d > out.png && open out.pngRun a saved recipe (handles textproto→JSON→call→save/open for you):
HOST=$HOST ./scripts/recipe-run.sh recipes/search_and_screenshot.textprotoSee recipes/ for reusable playbooks (load-then-screenshot, search,
dismiss-consent, scroll-to-lazy-load) and recipes/README.md
for the building blocks.
Grant another caller access:
gcloud run services add-iam-policy-binding chromerpc --region us-central1 \
--member 'user:[email protected]' --role roles/run.invokerAny gRPC client works — point it at HOST:443 with TLS and attach a bearer
identity token per call:
import (
"crypto/tls"
"google.golang.org/grpc"
"google.golang.org/grpc/credentials"
"google.golang.org/grpc/credentials/oauth"
"google.golang.org/api/idtoken"
pb "github.com/accretional/chromerpc/proto/cdp/headlessbrowser"
)
audience := "https://" + host // the Cloud Run service URL
ts, _ := idtoken.NewTokenSource(ctx, audience) // service-account or ADC creds
conn, _ := grpc.NewClient(host+":443",
grpc.WithTransportCredentials(credentials.NewTLS(&tls.Config{})),
grpc.WithPerRPCCredentials(oauth.TokenSource{TokenSource: ts}))
defer conn.Close()
client := pb.NewHeadlessBrowserServiceClient(conn)
res, err := client.RunAutomation(ctx, &pb.AutomationSequence{
Steps: []*pb.AutomationStep{
{Action: &pb.AutomationStep_Navigate{Navigate: &pb.Navigate{
Url: "https://example.com", WaitUntil: "networkidle"}}},
{Action: &pb.AutomationStep_Screenshot{Screenshot: &pb.Screenshot{Format: "png"}}},
},
})
// res.StepResults[i].ScreenshotData holds the PNG bytes.For other languages: open a TLS channel to :443 and send an
authorization: Bearer <id-token> metadata header on each RPC, using the
service URL as the token audience.
Treat every call as a self-contained, isolated session — don't rely on state (cookies, navigation, open tabs) persisting across calls; put the whole flow (navigate → act → screenshot) in one
RunAutomation.
The chrome-testing/ folder contains a self-contained screenshot testing module. It handles the full lifecycle — building chromerpc, launching Chrome, serving HTML, capturing PNGs, and tearing down — in a single script.
./chrome-testing/snap.sh my-page.html screenshots/my-page.pngSee chrome-testing/USAGE_INSTRUCTIONS.md for the full guide, including how any other project can copy this folder and use it for their own visual validation.
Nodejs implementaiton of the chrome remote interface: https://github.com/cyrus-and/chrome-remote-interface
VERY USEFUL: entire browser_protocol.json for the chrome remote interface https://github.com/ChromeDevTools/devtools-protocol/blob/master/json/browser_protocol.json
https://buf.build/docs/reference/descriptors/#what-are-descriptors this could be useful for converting individual domains or commands into .protos programmatically via https://github.com/protocolbuffers/protobuf/blob/main/src/google/protobuf/descriptor.proto and https://pkg.go.dev/google.golang.org/protobuf/reflect/protoreflect and https://github.com/jhump/protoreflect/tree/main/protoprint
