Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
NVIDIA_API_KEY=
NVIDIA_MODEL=nvidia/nemotron-3.5-lightning-30b-a3b
NVIDIA_BASE_URL=https://integrate.api.nvidia.com/v1

2 changes: 1 addition & 1 deletion LICENSE
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
MIT License

Copyright (c) 2026 Jeesh
Copyright (c) 2026 Jeethesh Reddy Gattupalli singalreddy

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
Expand Down
36 changes: 32 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ LLM-driven computer-use system that discovers UI workflows, records them as reus

## Status

The capability schema, local core-servicing app, deterministic replay, policy, and human handoff are in place. Discovery is not built yet.
Discovery, deterministic replay, policy, and human handoff are in place. A live model call is required only for discovery.

## Setup

Expand All @@ -13,17 +13,45 @@ Requires Node 22.12+.
```bash
npm install
npx playwright install chromium
cp .env.example .env
npm run check
```

Start the local target app (no API keys):
Add NVIDIA API key in `.env` as `NVIDIA_API_KEY` to run discovery. Create a key at [https://build.nvidia.com/models](https://build.nvidia.com/models).
Replay and `npm run check` do not need a key.

`npm run app` is only if you want to click the UI in a browser. `discover` and `replay` start their own copy of the app; you do not need this running for the demo commands.

```bash
npm run app
```

It listens on `http://127.0.0.1:4173/`. Member `10001` has a savings balance; any other id returns "Member not found".
It listens on `http://127.0.0.1:4173/`. Members `10001` (`$1,240.50`) and `10002` (`$50.00`) have savings balances; an unknown id returns "Member not found". A successful lookup shows a session warning; replay dismisses it and logs a `recovered` event.

## Demo

Not available yet. This section will have the commands to discover a goal and replay the resulting capability.
Discovery talks to NVIDIA NIM (`nvidia/nemotron-3.5-lightning-30b-a3b`) and writes a capability under `evidence/`. Replay does not call a model. Override the model with `NVIDIA_MODEL` in `.env` (must be a chat NIM with tool calling).

Recorded run (do not re-run `discover` into these paths unless you intend to replace it):

- `evidence/lookup-member-savings.json` — compiled capability (the contract)
- `evidence/discovery.json` — raw model log from that run
- `evidence/discovery.png` — screenshot at the end of discovery
- `evidence/replay-success.json` — replay for `10001` (`$1,240.50`, plus `recovered`)
- `evidence/replay-member-not-found.json` — replay for `99999`

```bash
npm run discover -- --goal "Look up the member savings balance" --param memberId=10001 --out evidence/lookup-member-savings.json
npm run replay -- --capability evidence/lookup-member-savings.json --param memberId=10001 --out evidence/replay-success.json
npm run replay -- --capability evidence/lookup-member-savings.json --param memberId=99999 --out evidence/replay-member-not-found.json
```

The same capability with `--param memberId=10002` returns `$50.00` (no second discovery). Human handoff is exercised in tests; the operator UI is mocked.

Without a key, `npm run check` still exercises discovery against the live app using a scripted model. To dry-run the CLI itself, write somewhere other than `evidence/`:

```bash
npm run discover -- --model scripted --param memberId=10001 --out /tmp/lookup-member-savings.json
```

`--model scripted` is a test double. Evidence meant to show a real discovery run must use the default NIM path.
88 changes: 88 additions & 0 deletions REPORT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,88 @@
# REPORT

## Architecture

Single process, TypeScript, CLI. Discovery and replay share one `Surface` (today `PlaywrightWebSurface`) and the same policy gate. The LLM is a recording engineer, not the production executor.

```
goal + entry URL
│
▼
discover() ──observe a11y inventory──► NVIDIA NIM (one tool call / turn)
│ fill | click | extract | done
▼
compile() ──parameterize, scrub, normalize──► capability JSON
│
▼
replay() ──no model──► success | business_outcome | escalated | failed
│
└── stuck / irreversible ──► HITL on the same RunSession
```

Zod validates the capability at the boundary. Playwright maps intents onto a live page. Policy runs **before** every act. `npm run check` (typecheck, lint, test) does not need an API key; CI uses a `ScriptedModel`. A genuine discovery run used `nvidia/nemotron-3.5-lightning-30b-a3b` via OpenAI-compatible `/chat/completions` (`fetch`, no extra SDK).

Rejected: queues, a capability catalog service, Python/Pydantic (one language for schema, replay, and the browser adapter). Rejected saving the model transcript as the artifact.

The target is a local server-rendered core-servicing mock (`apps/core-servicing`, port 4173): nested layout tables, ids `f1`/`cmd1`, no test IDs. Search → member detail. Known members: `10001` savings `$1,240.50`, `10002` savings `$50.00`. An unknown id (`99999` in evidence) returns “Member not found”. A successful lookup also shows a session-expiring dialog; the balance table stays hidden until Continue. Rejected a public cart demo (ToS, no controlled not-found). Rejected framesets in v1 (replay would need frame targeting before the first replay worked).

## Artifact schema

The file a calling agent invokes is `schemaVersion` 1.0.0 JSON (`src/schema/capability.ts`). `schemaVersion` is the format; `revision` is this flow. `strictObject` so a reviewed file cannot silently drop fields.

A capability is a **contract**: typed `parameters` and `outputs` with sensitivity; ordered `steps`; a success `checkpoint`. Values are refs (`param` / `literal` / `entry`), never a discovery-time member number. Fill is `{ kind: "param", name: "memberId" }`, not `"10001"`.

Steps are a discriminated union on `action`. Invalid combos (fill without a value, extract without an output) fail at parse time. Each step has `risk`: `read` | `reversible` | `irreversible`.

Locators are an ordered candidate list, biased to a surface with no clean DOM: `role_name` → `label` → `nearby_text` → `table_cell` → `structural` → `css`. CSS is last-resort. Test IDs are not a strategy. `table_cell` is row text plus column header, not a CSS nth-child.

`on` clauses on steps declare what replay should do when the page matches: `business_outcome`, `recover`, `escalate`, or `fail`. Replay does not invent that taxonomy at runtime.

The compiler (`src/discover/compile.ts`) is the seam between a noisy model run and that contract. A live Nemotron session extracted the same Savings cell under names `string` and `1240.50`, used `$1,240.50` as a checkpoint, and put a member id in the description. Compile scrubs param values, dedupes table cells, renames invalid outputs to `savingsBalance`, defaults `member_not_found` on Search, and replaces a money-amount checkpoint with heading `Member detail`. `evidence/discovery.json` is the raw log; `evidence/lookup-member-savings.json` is what production keeps.

## Determinism & error handling

Replay walks the artifact with no LLM. For each step it tries locators in rank order, then evaluates `on` against the observed page. Fill events log `param.memberId`, not the value.

Nested layout tables made a naive `table tr` extract return the whole page; the adapter targets **direct rows only** (`:scope > tbody > tr`). That is a replay bug, not “the UI drifted.”

Terminal statuses: `success` | `business_outcome` | `escalated` | `failed`.

- **`business_outcome`** is a legitimate caller result (`member_not_found`), not a crash. Evidence: `evidence/replay-member-not-found.json`.
- **`recovered`** is an **event**, not a status. After Search, the mock shows a session-expiring dialog and hides the balance table. Replay matches `on` → `recover` → dismisses Continue, logs `recovered`, then extracts. If dismiss works, the caller still gets success (or a business outcome). Evidence: `evidence/replay-success.json` includes that event.
- **`failed`** always includes `stepId`, `expected`, `observed`, and optionally a screenshot. Unknown UI (locator miss with no handoff, failed checkpoint) stops here. Replay does not call the model to improvise.
- **`escalated`** is control transfer (irreversible step, locator miss with a handoff that aborts, or an `on` → escalate), not “we are stuck internally.”

Happy-path evidence: `evidence/replay-success.json` (`savingsBalance: "$1,240.50"`). The same artifact with `memberId=10002` returns `$50.00` — parameterization, not a second recording. Locator miss without a handoff: broken Search name → `failed` on `click-search` with expected/observed. With a handoff, the same miss cedes the live page; the operator clicks Search and replay still extracts `$1,240.50`.

## Heterogeneity & multi-tenant

**Surface.** The artifact stores intents (`click`, `fill`, `extract`) and locator *candidates*, not Playwright selectors. `Surface` is how we perceive and act (goto, click, fill, extract, a11y `inventory`, observe). Today one implementation: Chromium. A desktop adapter would implement the same type and resolve `role_name` against the OS accessibility tree; `table_cell` would mean “row/column in the focused grid.” CSS candidates would no-op or be ignored. The replay engine would not change.

Framesets, extra document contexts, and screenshot+coordinates are the next surface problems, not schema problems. Coordinates were rejected as the default locator: they fail under DPI and layout shift; role+name matches how a human finds the control.

**Tenants.** The base artifact has `app.family` + `surface`, **no `tenantId`**. Hundreds of credit unions on the same vendor product should share one capability. Drift belongs in an overlay later: per-tenant locator inserts (an extra `label` candidate), copy variants for `on` match strings, entry URL. Detection: replay `failed` with the same `stepId` and a changed `observed` across tenants is a locator/copy drift signal, not a reason to re-record the whole flow. Re-record when the *intent* changed (a new confirmation step), not when a button’s accessible name gained a suffix.

## Escalation & handoff

Stuck means: policy refuses an irreversible act, an `on` handler says escalate, a locator miss when a handoff is present, or discovery `give_up` / max steps. Locator-miss with no handoff stays `failed`. With a handoff, the same session is ceded **once per step**; a second miss on that step fails instead of looping.

`RunSession` owns the live `Surface` and `owner: automation | human`. Handoff is `cede` → `intervene` → `resume` on that object. A new browser is not a handoff.

The intervention carries capability/goal, current step, page text, optional screenshot, and why it stopped. The operator acts on `session.surface`. Resume is `skip_step` (human finished the blocked step — correct after an irreversible click), `retry_step`, or `abort`. Human work is `events[].type === "human"`.

Operator UI is mocked: `ScriptedOperator` in tests (Playwright runs for irreversible Search and for a broken Search locator; the operator clicks the real Search on the **same** page and replay extracts `$1,240.50`); `PromptOperator` is headed browser plus stdin `skip|retry|abort`. A co-browsing console would subscribe to the same `Handoff` interface.

## Safety

Policy (`src/policy`) runs before the surface moves: allowlisted action types, host, path prefix. Default allowlist is `127.0.0.1` / `localhost`. Irreversible defaults to **escalate and do not click**; `onIrreversible: "block"` is the stricter option. Discovery is bound by the same gate.

Artifacts and logs must not persist secrets or raw identifiers. Param values with sensitivity `identifier` / `financial` / `secret` / `full_pii` are stripped from failure `observed` (`[memberId]`, not `10001`). Success **outputs** stay intact — that is the capability contract (the caller asked for the balance). Screenshots are not pixel-redacted (cut). Discovery prompts redact inventory values the same way so the model sees “filled” without a raw member number in our logs.

## Cuts

- **Operator UI** — mock; control transfer is real.
- **Desktop / framesets / multi-tenant overlays** — designed, not built.
- **Screenshot redaction, capability catalog API, codegen, N-run flakiness, bounded LLM fallback on replay** — stretch; skipped until this thread was evidenced.
- **CI model** — `ScriptedModel`. The NIM path is real; CI does not call it.

Next, if this were production: a tenant overlay for copy/locator drift.
17 changes: 16 additions & 1 deletion apps/core-servicing/server.ts
Original file line number Diff line number Diff line change
Expand Up @@ -119,6 +119,13 @@ function memberPage(rawId: string) {
)
.join("");
return shell(`
<div role="dialog" aria-modal="true" aria-labelledby="sess-title">
<table class="panel" cellpadding="8" cellspacing="0">
<tr><td id="sess-title"><b>Your session is about to expire.</b></td></tr>
<tr><td><input type="button" id="cmd2" value="Continue"></td></tr>
</table>
</div>
<div id="member-body" hidden>
<h1>Member detail</h1>
<table cellpadding="4" cellspacing="0">
<tr><td>Member</td><td>${escapeHtml(id)}</td></tr>
Expand All @@ -129,7 +136,15 @@ function memberPage(rawId: string) {
<tr><th>Account</th><th>Balance</th></tr>
${rows}
</table>
<p><a href="/">New search</a></p>`);
<p><a href="/">New search</a></p>
</div>
<script>
document.getElementById("cmd2").onclick = function () {
var dialog = document.querySelector('[role="dialog"]');
if (dialog) dialog.remove();
document.getElementById("member-body").hidden = false;
};
</script>`);
}

const isDirectRun =
Expand Down
Empty file added evidence/.gitkeep
Empty file.
78 changes: 78 additions & 0 deletions evidence/discovery.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,78 @@
{
"goal": "Look up the member savings balance",
"model": "nvidia/nemotron-3.5-lightning-30b-a3b",
"events": [
{
"type": "acted",
"detail": "navigate entry"
},
{
"type": "decided",
"detail": "fill Member ID"
},
{
"type": "acted",
"detail": "fill Member ID"
},
{
"type": "decided",
"detail": "click Search"
},
{
"type": "acted",
"detail": "click Search"
},
{
"type": "decided",
"detail": "extract string"
},
{
"type": "acted",
"detail": "extract string"
},
{
"type": "decided",
"detail": "extract 1240.50"
},
{
"type": "acted",
"detail": "extract 1240.50"
},
{
"type": "decided",
"detail": "extract 1240.50"
},
{
"type": "acted",
"detail": "extract 1240.50"
},
{
"type": "decided",
"detail": "extract 1240.50"
},
{
"type": "acted",
"detail": "extract 1240.50"
},
{
"type": "decided",
"detail": "extract 1240.50"
},
{
"type": "acted",
"detail": "extract 1240.50"
},
{
"type": "decided",
"detail": "extract 1240.50"
},
{
"type": "acted",
"detail": "extract 1240.50"
},
{
"type": "decided",
"detail": "done: Goal accomplished - savings balance extracted from member detail page"
}
]
}
Binary file added evidence/discovery.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
127 changes: 127 additions & 0 deletions evidence/lookup-member-savings.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,127 @@
{
"schemaVersion": "1.0.0",
"id": "member-savings-balance-lookup",
"name": "Member savings balance lookup",
"description": "Retrieve the savings balance for member [memberId]",
"revision": 1,
"app": {
"family": "core-servicing",
"surface": "web",
"entryPoint": "http://127.0.0.1:4173/"
},
"parameters": [
{
"name": "memberId",
"type": "string",
"sensitivity": "identifier",
"description": "Invocation parameter memberId"
}
],
"outputs": [
{
"name": "savingsBalance",
"type": "money",
"sensitivity": "financial",
"description": "Value of Savings / Balance"
}
],
"steps": [
{
"id": "open-app",
"risk": "read",
"action": "navigate",
"url": {
"kind": "entry"
}
},
{
"id": "fill-member-id",
"risk": "reversible",
"action": "fill",
"target": {
"candidates": [
{
"strategy": "role_name",
"role": "textbox",
"name": "Member ID"
},
{
"strategy": "label",
"label": "Member ID"
}
]
},
"value": {
"kind": "param",
"name": "memberId"
}
},
{
"id": "click-search",
"risk": "read",
"on": [
{
"match": {
"kind": "text",
"value": "Member not found"
},
"then": {
"type": "business_outcome",
"code": "member_not_found"
}
},
{
"match": {
"kind": "dialog",
"value": "Your session is about to expire."
},
"then": {
"type": "recover",
"action": "dismiss",
"target": {
"candidates": [
{
"strategy": "role_name",
"role": "button",
"name": "Continue"
}
]
}
}
}
],
"action": "click",
"target": {
"candidates": [
{
"strategy": "role_name",
"role": "button",
"name": "Search"
}
]
}
},
{
"id": "extract-savingsbalance",
"risk": "read",
"action": "extract",
"target": {
"candidates": [
{
"strategy": "table_cell",
"rowText": "Savings",
"columnHeader": "Balance"
}
]
},
"output": "savingsBalance"
}
],
"success": {
"checkpoint": {
"kind": "heading",
"value": "Member detail"
},
"outputs": ["savingsBalance"]
}
}
Loading
Loading