diff --git a/CHANGELOG.md b/CHANGELOG.md index 7b1eca6..a23d748 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,19 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); ver ### Added +- Bounded recovery on the browser front: `--rethink on|off` (default off), `--rethink-attempts` (3) and + `--rethink-timeout` (15 s). A stall (no page change, an A-B-A-B loop, a repeated URL, a WAIT that moved nothing) + re-probes the page read-only through the same runtime permission and asks the chat model for a plan under a + per-task `RecoveryLimits` budget; the next normal decision still picks the action, the plan never executes and + cannot widen the offered tools or add `unsafe_dev`. Only the detection windows reset after an attempt; attempts, + seconds, history and ticks stay. `report()` and `decision_ticks.json` keep the recovery events, counts and + termination, and a timed-out, failed or exhausted recovery is a clear `BLOCKED` even when a summary carries text. + `--model llm` with `--rethink on` is rejected before the agent or browser is built. +- The desktop agent's bounded recovery takes the same `--rethink-attempts` and `--rethink-timeout` names and stalls + after 3 actions without progress. Recovery failures include an operator next action; failed or cancelled chat + calls remain counted, and incomplete token usage is reported as unknown cost. +- Small repeatable browser and native Windows recovery on/off fixtures, with independent completion checks, + bounded failure cases and paired reports. These use scripted models to verify mechanisms, not model accuracy. - The MCP `decide` tool accepts `model="jev"|"laya"|"cua"`, defaulting to `jev`. Local backends use their optional extras and need no Jev API key. - `docs/benchmarks.md`: the Google Flights driver comparison rerun on 2026-09-23 from Poland, every arm three times on diff --git a/docs/agents.md b/docs/agents.md index c7ff39a..180f0e5 100644 --- a/docs/agents.md +++ b/docs/agents.md @@ -18,7 +18,10 @@ hook of a running agent. The injection guard rail fails closed: a decision error Every tool agent takes `--model jev|laya|cua|llm|random|rule`, `--rethink on|off`, `--episodes N`, `--seed S`, `--max-steps` and `--timeout`, and writes a Harbor-shaped job folder under `evals/results//`. A browser agent -takes `--model jev|laya|cua|llm` and `--goal`. A rail takes `--model jev|laya`, the two models that answer `noul`. +takes `--model jev|laya|cua|llm`, `--goal` and the same `--rethink` flag. The desktop agent adds the bounded-recovery +`--rethink-attempts` and `--rethink-timeout` and stalls after 3 actions; the browser front takes the same three names, +and `docs/browser-front.md` decision 19 describes its branch. A rail takes `--model jev|laya`, the two models that +answer `noul`. `uv run python -m evals.table evals/results` aggregates every job folder per eval and model into one table. Every `run` prints one JSON object on stdout and nothing else there; `s1a-mcp` serves the same agents over stdio @@ -26,12 +29,19 @@ with `list_agents`, `run_agent` and `decide`. Flags, exit codes and the job-fold [architecture.md](architecture.md). The extras each agent needs and the keys: `CONTRIBUTING.md`. The `--model` values and the models behind them: [architecture.md](architecture.md#models). +With bounded recovery enabled, a failed refresh or plan records its reason and a suggested next action for the +operator. A task the policy still answers `BLOCKED` after one or more replans is a failure too, even when a partial +answer was fetched: the answer is context only and the run carries the block reason and a next action, without being +reported as a failed or exhausted recovery. Failed and cancelled chat calls still count toward the run's call total. +If a call's token usage is unknown, the reported cost stays unknown rather than becoming zero. + ## Agent-specific flags `s1a run --help` lists every flag with its default. Beyond the shared ones: `flights` and `allrecipes` take `--goal`, -`--batch on|off`, `--prefetch on|off`, `--goal-values on|off`, `--profile-out` and `--logs-dir`; `desktop` takes -`--app`, `--goal`, `--expect`, `--execute`, `--plan` and `--clear`; `ticket_router` takes `--dataset` and -`--batch-size`; `injection_guard` takes `--labelled-set`. The four games take no flag of their own. +`--batch on|off`, `--prefetch on|off`, `--goal-values on|off`, `--rethink on|off`, `--rethink-attempts`, +`--rethink-timeout`, `--profile-out` and `--logs-dir`; `desktop` takes `--app`, `--goal`, `--expect`, `--execute`, +`--plan`, `--clear` and the same three `--rethink` flags; `ticket_router` takes `--dataset` and `--batch-size`; +`injection_guard` takes `--labelled-set`. The four games take no flag of their own. ## Allrecipes diff --git a/docs/browser-front.md b/docs/browser-front.md index 026ff92..0c3e2cb 100644 --- a/docs/browser-front.md +++ b/docs/browser-front.md @@ -2,7 +2,8 @@ Code: `s1a/browser/` (`decision_model.py`, `action_space.py`, `probe_js.py`, `prompts.py`), the decision-model layer in `s1a/decision_models/` -(`docs/decision-models.md`) and the HTTP transport in `s1a/decision_models/wire.py`. Tests: `tests/test_browser_policy.py`. The +(`docs/decision-models.md`) and the HTTP transport in `s1a/decision_models/wire.py`. Tests: `tests/test_browser_policy.py` +and `tests/test_browser_recovery.py`. The harness side (the `DecisionPolicyModel` Protocol, `probe_for_policy` and `activate_page` on the Playwright runtime, the policy path in `create_browser_agent`) is the decision-policy slot pinned in `pyproject.toml`. Measurements: `docs/benchmarks.md`. @@ -38,7 +39,7 @@ browser-use/jev-ultrafast (MIT), whose observe-decide-act tick this policy follo on every page change and discards stale answers, 8 to 11 per run in measurement. 4. Hidden tabs are activated. A hidden tab throttles timers to about 1 Hz and does not render dropdowns. When the probe reports `visibilityState == "hidden"`, the policy calls `activate_page(url)` once and - probes again. + probes again; a failed re-probe is refused like any other failed observation (item 8). 5. The chat model generates typed values in the background before the field is reached. A `TYPE_TEXT` string is generated from the goal, the field, the page text and the history. With `--prefetch on` (`BrowserPolicy.prefetch_values`, the default) that call starts in the background for every editable field @@ -55,10 +56,14 @@ browser-use/jev-ultrafast (MIT), whose observe-decide-act tick this policy follo live in a dataclass; a new goal or a finished run starts a fresh one and cancels the previous run's background tasks. The decision model is `self._decision_model`; `Model.__init__` builds `self._client` as a real telemetry-bearing model client and inherited methods touch it. -8. Failures degrade to `BLOCKED`. A probe failure envelope (`{"ok": False, "error": ..., "elements": []}`) - folds into a control-only action space (WAIT, DONE, BLOCKED); a decisions transport error, HTTP error, - malformed 200 body or invalid distribution ends the turn with a `BLOCKED` summary that still holds URL, - title, steps and page text. Response bodies never enter error messages or logs. +8. A failed observation is refused, not decided. The runtime reports a failed probe as `ok=False` or an + `error` field. The policy checks the snapshot at the observation boundary (`_require_observation`) and + raises `RuntimeError("browser observation failed: ...")` with a bounded one-line reason before the + decision model or the chat answer can read an empty page. This keeps a broken browser (a missing + library such as `libatk-1.0.so.0`, a closed page) from being answered DONE; the recovery refresh and + the in-page action-settle/WAIT probes keep their own bounded handling. A decisions transport error, HTTP + error, malformed 200 body or invalid distribution ends the turn with a `BLOCKED` summary that still holds + URL, title, steps and page text. Response bodies never enter error messages or logs. 9. Answers are validated in the decision-model layer. `decide_many` accepts only a choice among the offered ids whose distribution covers exactly those ids, sums to 1 within 0.02, and peaks at the choice, for every head of the request, including the heads the choice did not select. An unusable answer is re-asked once @@ -113,6 +118,29 @@ browser-use/jev-ultrafast (MIT), whose observe-decide-act tick this policy follo probe's occlusion check sees what covers a control now; a consent banner that returns after every navigation leaves the control looking clickable on each fresh probe. In the Allrecipes batch the run ended on the no-page-change guard after three such clicks. +19. Bounded recovery (`--rethink on`, off by default; `--rethink-attempts`, default 3; `--rethink-timeout`, + default 15 s). A stop signal -- `stall_after` actions without a page change, an A-B-A-B page loop, a fourth + arrival at one URL, an over-long WAIT streak or a WAIT settle that moved nothing -- spends one attempt instead + of ending the run. The attempt re-probes the page read-only through `_raw_probe` (never a navigation; a probe + error or a denied read stops immediately) and, when the `page_key` really did not move, asks the chat model for + a plan through the tool front's `draft_plan`; a changed `page_key` is delayed progress and skips the planner. + The plan rides into the next observation only and never executes a click; the offered tools and candidates are + filtered exactly as an ordinary turn filters them, so recovery cannot widen the tool set or add `unsafe_dev`. + `RecoveryLimits` bounds the attempts and the seconds spent inside the refresh and planner calls, counted from + `time.perf_counter` and cancelled by `asyncio.wait_for`, cumulatively and never reset by progress; the outer + `--max-steps` and `--timeout` still cap the task. Only the detection windows (the history boundary the stall and + oscillation guards read, the WAIT streak and the visit counts) reset after an attempt: history, ticks and the + budget stay, so a plan gets a few actions before the same stall can fire again. A refresh or planner timeout, + error, cancellation or an exhausted budget ends the run BLOCKED with no answer fetched; `finish_decision_model` + reports it as a failure even if the summary carries text. A policy that still answers BLOCKED after one or more + replans is a failure too: the partial answer is kept only as context, the summary carries the block reason and a + next action, and the record is not mislabelled as a failed or exhausted recovery or a failed planner call. + `--model llm` with `--rethink on` is rejected before + the agent or browser is built. `report()` and `decision_ticks.json` keep the recovery events, attempts, seconds + and termination. Failed recovery also includes a next action for the operator; failed and cancelled planner calls + remain in the call count, with unknown token usage kept separate from zero cost. To compare, run the same goal twice, `--rethink off` against `on`, on the same backend; the + on/off difference is only worth quoting once the browser bench is rerun, because the scripted tests pin the loop, + not a task success rate. ## Rejected diff --git a/evals/README.md b/evals/README.md index 2983f5a..e285a45 100644 --- a/evals/README.md +++ b/evals/README.md @@ -14,6 +14,8 @@ folder per episode). `uv run python -m evals.table evals/results` aggregates the | Desktop (Cua Driver) | clicks toward `--goal` until the window shows `--expect` | `s1a run desktop --app Calculator --goal "compute 12 times 7" --expect 84 --execute --model jev --rethink off --episodes 1` | | Ticket router (30 local labelled tickets) | correct routes to five queues | `s1a run ticket_router --model jev --rethink off --episodes 1` | | Injection guard (rail) | precision and recall on a labelled set | `s1a run injection_guard` | +| Bounded recovery on/off (3 local form fixtures) | completion on/off in a real headless Chromium, judged by the fixture server | `python -m evals.recovery --repeat 1` | +| Bounded recovery live (real decision/chat models, 3 local form fixtures) | completion on/off with the real models in a real headless Chromium, judged by the fixture server; cost null unless `CHAT_USD_PER_M_*` is set; optional `--validate-submission` variant | `python -m evals.recovery.live --model jev --repeat 1` | ## The loop @@ -39,6 +41,33 @@ without a score change ends the episode. Blackjack and Millionaire never stall a `act(LEFT)` calls in a row are a legal 2048 line, the harness's own anomaly rail (its bailout ends a run on the third identical call) stays off for these agents. +The browser front and the desktop agent share the switch and the flags: `--rethink on|off`, `--rethink-attempts` +(default 3) and `--rethink-timeout` (default 15 s). On a browser or desktop stall the run spends one bounded attempt +that re-probes the page and asks the chat model for a plan, then the normal decision picks the action (browser: +`docs/browser-front.md` decision 19; desktop: the same `RecoveryLimits`). The attempts and the seconds across all +refreshes and plans are cumulative for the task and never reset by progress; `--max-steps` and `--timeout` still cap +it. Recovery is a loop mechanism, not a task result. `python -m evals.recovery` is a small repeatable on/off subset +for the browser front (issue 16): three local form fixtures (`normal`, `recoverable`, `permanently_blocked`) served +from a loopback server, each played in a real headless Chromium through the production `browse` front with +`create_browser_agent`, the registered `browser_*` tools, their generation ids and the same `RecoveryLimits`; batch +actions and `unsafe_dev` stay off. Success is the fixture server's recorded form POST, never the model's DONE. Each +trial opens its **own fresh fixture** (unique trial id, empty record namespace), so a submit that verified one trial +can never verify a later one; the arms alternate their run order across repeats to blunt startup-order bias. The +decision model and planner are clearly-marked **scripted doubles** (`evals/recovery/scripted.py`) - a controlled +fault that keeps typing until the production plan arrives - so this measures the **mechanism**, not a trained model +and not a real site. Scripted decisions cost 0 by construction; there is no real-backend mode (`--backend` is not an +option) because real-model token usage and pricing are not implemented here. It writes the usual Harbor job folders +under `evals/results/recovery/` and a paired `paired_summary.json`/`.md`. The paired completion counts every planned +trial, errors and timeouts included, and pairs recovery off/on per (task, repeat); it also reports the paired mean +on-minus-off deltas for model calls, elapsed seconds and wasted actions. Model calls are split into `decision_calls`, +`planner_calls` and `chat_calls`. Timing stays separate: `elapsed_s` is the whole trial wall clock (not inference), +`decisions_ms` is the scripted decision time measured with `perf_counter` (near-zero and not a benchmark), and +`recorded_probe_ms` sums recorded probe wall times (already including action settling). It excludes recovery, +final and unticked probes, so it is a diagnostic rather than total environment time. +`--repeat` must be >= 1 and arms and tasks must be valid; the CLI exits non-zero when a planned trial raised (harness error or timeout), while an expected terminal +score such as `BLOCKED` does not fail the run. Native Windows desktop recovery still needs a supported desktop +machine. + Every episode records its ticks (key, confidence, probabilities, latency, tokens, whether a plan was in the state) and its rethink events in `agent/episode.json`; the count of chat-model calls and their token sums go to `result.json`. For `llm` the decisions are the chat calls that produced an `act`; a tick whose key was @@ -49,6 +78,90 @@ raised (a decisions failure, an env error) keeps `error` set: `summary.json` rep the score statistics over the `scored` episodes only, and `evals.table` skips trials whose `result.json` has `exception_info`. + +### Live real-model recovery subset + +`python -m evals.recovery.live --model jev|laya|cua --repeat N --tasks normal,recoverable,permanently_blocked +--arms off,on --results-dir DIR` is the matching real-model run: the same fixtures, production `browse`, +`BrowserPolicy`, registered `browser_*` tools, generation ids and `RecoveryLimits`, but with `build_model` and +`chat_model_from_env` supplying the real decision and chat/planner models. Nothing is injected: no scripted +decisions, no standard plan and no hidden answer, so the run is the model's own. Optional flags: `--timeout`, +`--max-steps`, `--stall-after`, `--recovery-attempts`, `--recovery-timeout`, `--headed`. Success is still the fixture +server's recorded POST, never the model's `DONE`, and each trial opens its own fresh fixture and browser. + +Counting is split and never conflated. Decision calls are counted by a thin proxy that appends its record at the +`_decide` entry and fills the wall time and outcome in a `finally`-safe step, so retries after an unusable answer, +raised calls and cancellations are counted for real (ticks are not a call count) and a failed call's tokens are +unknown, not zero; the proxy preserves the inner model's `name`, `question_types`, `deterministic`, `supports_images`, +`model`, `warm` and `close`. Planner attempts are counted from the production recovery events whose `stage == +"planner"` (a planner call is also inside `chat_calls`, and the two are never summed). Chat call counts and tokens +come from the production `CountingModel` that `browse` wraps around the chat model; when a trial raises before that +summary exists they are taken from the outer `CountingModel`, which records the same real calls. A trial that raises +inside `browse` keeps its decision/chat counters and its fixture oracle: `verified` (the fixture's recorded POST) and +`errored` (the run raised) are independent, and only a trial that failed before the browser existed falls back to a +fabricated record. `cost_usd` is `null` by default, because the live eval never reads a provider catalogue; only an +explicit `CHAT_USD_PER_M_*` configuration values a trial, and that value is labelled an estimate, not a bill. Even +then a trial stays `null` when any chat or decision call did not report usage. (The production browser path's own +`usage_summary` may look up OpenRouter's catalogue; the live eval deliberately ignores it.) Model load time is +measured once and listed separately from the per-trial numbers; a setup failure (a model that will not build, a chat +model with no key) exits non-zero without inventing any trial, and the already-built decision model is still closed. + +Each run writes a directory under `/recovery_live//` (the run id has a unique suffix) with +`manifest.json` (the planned manifest first, then the actual chat `model_config` name and provider, the relevant +`LAYA_*`/`CUA_S1_*` context variables, model ids, platform, config, load times and cost basis), `trials.json` (every +planned trial, errors and timeouts included), `summary.json` and `summary.md` (per-arm and per-task completion, +recovery events/terminations, decision/planner/chat attempt counts and timings, wasted actions, token-known status). +Each finished trial is also written on its own under `trials/`, so a batch scheduler's kill does not erase completed +evidence; there is no resume engine. `--repeat` must be >= 1 and the arms/tasks must be valid; the CLI exits +non-zero when a planned trial raised or could not be set up. The result is honest about its scope: three small +synthetic pages driven by a real model, not an open-task success rate, and recovery may not trigger at all. + +#### Opt-in server-side validation (exploratory follow-up) + +`--validate-submission` (also `LiveConfig(validate_submission=True)`, recorded in the manifest's `config`) is a small +optional variant for a **prospective exploratory follow-up**, not a change to the default fixtures or their reports. +When it is on, `run_live` uses validated equivalents of the supplied tasks: the fixture server answers a real POST +whose value differs from the task's expected value with HTTP 422, re-serving the same retryable form plus an ordinary +visible validation error (the submitted data and message are HTML-escaped), never a success result page. A correct +value still returns the result page and the independent oracle still verifies it. Every real POST, rejected ones +included, is recorded. Nothing is injected: the model still sees only the ordinary page and the goal, with no scripted +decision, plan or hidden answer. The original `DEFAULT_TASKS` and `LiveConfig()` keep `validate_submission=False` and +are immutable; `run_live` replaces only its own local copy, so the default behaviour is unchanged. + +The variant exists because the first valid normal-form pilot submitted an empty value then terminated, which exercises +premature completion rather than a stall; the follow-up asks whether ordinary server-side validation can expose +naturally repeated ineffective actions for the bounded-recovery mechanism. Recovery improvement is **not assumed**. It +is exploratory and not an independent confirmatory benchmark: report its results **separately** from the original +unconstrained forms and never combine or average the two. Each keeps its own full denominator (every planned trial, +errors and timeouts included), its own natural trigger counts, actual planner/decision/chat calls, successes, latency, +wasted actions and termination reasons, even when there is no improvement. + +### Native desktop recovery subset + +`python -m evals.desktop_recovery` runs the matching normal / recoverable / permanently blocked cases on a +small Windows fixture. It compiles the committed C# source with the existing system .NET compiler and drives only +its own verified process through the real `CuaDriver`, `WindowEnv` and tool-agent loop. The decision model and +planner are scripted; the app's own result file is the independent verifier. No model key or GPU is required. +Windows, or WSL with Windows interop and an interactive Windows desktop, is needed; this does not extend the +production desktop CLI's Linux support. + +```bash +python -m evals.desktop_recovery --driver /path/to/cua-driver.exe +S1A_DESKTOP_TESTS=1 S1A_DESKTOP_DRIVER=/path/to/cua-driver.exe \ + pytest -q tests/system/test_recovery_desktop.py +``` + +Each run gets a fresh directory, with a paired summary and Harbor job records. `--resume` explicitly reuses saved +trials from the specified run id; omit it for a new experiment. Errors/timeouts remain in the planned denominator +and cause a nonzero CLI exit; expected blocked tasks do not. The report records actual planner calls, accepted +actions whose observed progress did not change, recovery budgets, and the operator next action recorded by the +runtime. It does not invent an escalation in the evaluator. Trial time includes fixture startup and the agent +episode, but excludes the shared compiler and driver startup. + +The system test is skipped unless explicitly enabled. Driver binaries, generated fixture executables, environments +and result folders are local artifacts, not source contributions. See `tests/fixtures/desktop_recovery/S1AFixture.cs` +and `tests/system/test_recovery_desktop.py` for the fixture and the independent assertions. + ## Setup `uv sync`, plus `--extra blackjack` and `--extra alfworld` for those games (the README lists every extra), then @@ -128,3 +241,5 @@ Jev paged the date picker back and forth and ended BLOCKED after 18 requests (th 7. Stall thresholds per game: 2048 six acts, ALFWorld eight, Blackjack and Millionaire none. 8. One slot model over one decision-model interface: `ToolDecisionModel` holds a decision model, and its `name` is the tick's `source` and the episode's `policy`. + +The [small real-model recovery report](recovery/RESULTS.md) records both the original negative result and the separate exploratory server-validation follow-up, with per-trial data and limitations. diff --git a/evals/desktop_recovery.py b/evals/desktop_recovery.py new file mode 100644 index 0000000..97231a8 --- /dev/null +++ b/evals/desktop_recovery.py @@ -0,0 +1,886 @@ +# coding: utf-8 +"""Native desktop bounded-recovery on/off eval on a self-built, per-run fixture. + +This is the reproducible subset of the 2026-09-27 native Windows experiment: ``normal``, ``recoverable`` and +``permanent`` tasks, each with recovery ``off`` and ``on``. It reuses the repository's real machinery unchanged: +the real ``cua-driver`` over its MCP stdio bridge, the real ``CuaDriver`` adapter, the real ``WindowEnv``, the real +``run_episode`` loop and the real ``RethinkRail`` (bounded branch). Only the decision model and the planner are +scripted, for controlled fault injection: this measures the mechanism wiring and the budget, **not** a trained model +and **not** a real app success rate. There is no API call and the cost is 0 by construction. + +The fixture source lives under ``tests/fixtures/desktop_recovery`` and is compiled by the system .NET compiler; no +installer is run. Every run gets a unique id, so the process image name and the window title are unique and only this +run's window is ever targeted. The fixture never activates its window and the driver clicks in the background, so the +mouse is not moved and no other app is read or clicked. Only the fixture pid this run started is terminated. + +The independent oracle is the fixture's own ``result.json``: success is ``finished`` true, never the model's ``DONE``. +Each trial gets its own fresh fixture and output directory, so no earlier trial's success can verify a later one. +""" + +from __future__ import annotations + +import argparse +import asyncio +import json +import os +import platform +import re +import shutil +import statistics +import subprocess +import sys +import time +from dataclasses import asdict, dataclass, field +from datetime import datetime +from pathlib import Path +from typing import Any, AsyncIterator + +from openjiuwen.core.foundation.llm import AssistantMessage, AssistantMessageChunk, Model + +from s1a.agents import desktop +from s1a.decision_models import RuleModel +from s1a.desktop.driver import CuaDriver, DriverError, opened +from s1a.desktop.env import WindowEnv +from s1a.jobs import RESULTS_DIR, Episode, now_iso, write_job +from s1a.recovery import RecoveryLimits +from s1a.run import started_runner +from s1a.tool import loop +from s1a.tool.models import placeholder_model +from s1a.tool.rethink import progress_digest + +REPO = Path(__file__).resolve().parent.parent +FIXTURE_SOURCE = REPO / "tests" / "fixtures" / "desktop_recovery" / "S1AFixture.cs" +TASKS = ("normal", "recoverable", "permanent") +ARMS = ("off", "on") +SCRIPTED_MODEL = "scripted-rule" # a hand-written rule, not a neural decision model +GOAL = "press unlock then finish so the window title shows finished OK" +PLAN = "Click unlock, then click finish." +DONE_TEXT = "finished OK" +DRIVER_CHANNEL = "not inferred from the executable path" +DEFAULT_DRIVER_VERSION = "unknown" +ORACLE_WAIT_S = 25.0 +WINDOW_WAIT_S = 15.0 # the fixture writes its oracle before its window is on screen; wait for the window +LAUNCH_ATTEMPTS = 3 # WSL interop can refuse a launch with a transient EINVAL; retry a dead attempt only +CSC_CANDIDATES = ( + r"C:\Windows\Microsoft.NET\Framework64\v4.0.30319\csc.exe", + "/mnt/c/Windows/Microsoft.NET/Framework64/v4.0.30319/csc.exe", +) +TASKKILL = "/mnt/c/Windows/System32/taskkill.exe" + +_NEXT_ACTION = re.compile(r'"next_action"\s*:\s*"((?:[^"\\]|\\.)*)"') +_STATUS = re.compile(r'"status"\s*:\s*"([^"]*)"') +_REASON = re.compile(r'"reason"\s*:\s*"((?:[^"\\]|\\.)*)"') + + +@dataclass(frozen=True) +class TrialPlan: + """One planned trial: a task, an arm and a repeat index.""" + + task: str + arm: str + repeat: int = 0 + + +@dataclass(frozen=True) +class DesktopEvalConfig: + """The fixed budgets and binaries one run uses; on/off share every value but the arm itself.""" + + driver_bin: str + csc: str + fixture_source: Path = FIXTURE_SOURCE + max_acts: int = 12 + timeout_s: float = 60.0 + max_recovery_attempts: int = 3 + recovery_timeout_s: float = 15.0 + session_label: str = "s1a-desktop-recovery-eval" + driver_version: str = DEFAULT_DRIVER_VERSION + + +@dataclass +class DesktopTrialRecord: + """One trial's outcome: the independent oracle plus the recovery accounting, all scripted and cost-free.""" + + task: str + arm: str + repeat: int + seed: int + verified: bool + terminal: str + errored: bool + scripted_model: str + driver: dict[str, Any] + platform: str + recovery_attempts: int + recovery_spent_s: float + recovery_failed: bool + recovery_termination: str | None + recovery_next_action: str | None + next_action_source: str | None + wasted_actions: int + noop_clicks: int + steps: int + model_calls: int + decision_calls: int + planner_calls: int + chat_calls: int + elapsed_s: float + cost_usd: float + oracle: dict[str, Any] = field(default_factory=dict) + + def as_json(self) -> dict[str, Any]: + return asdict(self) + + +@dataclass +class DesktopEvalRun: + summary: dict[str, Any] + records: list[DesktopTrialRecord] = field(default_factory=list) + run_dir: Path | None = None + job_dirs: list[Path] = field(default_factory=list) + + +class PlanChat(Model): + """A scripted planner: one useful plan line per stall, counting its calls. No API, no tokens, no cost.""" + + def __init__(self) -> None: + source = placeholder_model() + super().__init__(source.model_client_config, source.model_config) + self.calls = 0 + + async def invoke(self, messages: Any, *, tools: Any = None, **kwargs: Any) -> AssistantMessage: + self.calls += 1 + return AssistantMessage(content=PLAN) + + async def stream(self, messages: Any, *, tools: Any = None, **kwargs: Any) -> AsyncIterator[AssistantMessageChunk]: + self.calls += 1 + yield AssistantMessageChunk(content=PLAN) + + +def default_driver() -> str | None: + return os.getenv("S1A_DESKTOP_DRIVER") or os.getenv("CUA_DRIVER_BIN") or shutil.which("cua-driver") + + +def default_csc() -> str | None: + override = os.getenv("S1A_DESKTOP_CSC") + if override: + return override + for candidate in CSC_CANDIDATES: + if Path(candidate).exists(): + return candidate + return shutil.which("csc") or shutil.which("csc.exe") + + +def detect_driver_version(driver: str) -> str: + """Read the supplied binary's version; do not label an arbitrary binary as the tested release.""" + try: + result = subprocess.run([driver, "--version"], capture_output=True, text=True, timeout=10, check=False) + except (OSError, subprocess.TimeoutExpired): + return "unknown" + match = re.search(r"\b\d+\.\d+\.\d+(?:[-+][\w.-]+)?\b", result.stdout) + return match.group(0) if result.returncode == 0 and match else "unknown" + + +def default_config(*, driver_bin: str | None = None, csc: str | None = None) -> DesktopEvalConfig: + driver = driver_bin or default_driver() + if not driver: + raise RuntimeError("no cua-driver binary; pass --driver / S1A_DESKTOP_DRIVER, or install cua-driver") + compiler = csc or default_csc() + if not compiler: + raise RuntimeError("no C# compiler; pass --csc / S1A_DESKTOP_CSC (the installed system .NET csc.exe)") + return DesktopEvalConfig(driver_bin=driver, csc=compiler, driver_version=detect_driver_version(driver)) + + +def new_run_id() -> str: + return datetime.now().strftime("%Y%m%d-%H%M%S") + "-" + os.urandom(3).hex() + + +def default_trials(repeat: int = 1, tasks: tuple[str, ...] = TASKS, arms: tuple[str, ...] = ARMS) -> list[TrialPlan]: + return [TrialPlan(task, arm, index) for task in tasks for index in range(repeat) for arm in arms] + + +def to_windows_path(path: Path) -> str: + """A path the Windows fixture/compiler can write to; on native Windows the path is already native.""" + if os.name == "nt": + return str(path) + try: + result = subprocess.check_output(["wslpath", "-w", str(path)], text=True).strip() + return result or str(path) + except (OSError, subprocess.CalledProcessError): + return str(path) + + +def fixture_mode(task: str) -> str: + return "permanent" if task == "permanent" else "normal" + + +def scripted_rule(task: str, arm: str) -> RuleModel: + """The controlled fault: a hand-written rule per task; it reads the plan the rail leaves in the observation.""" + state = {"unlocked": False, "go": False} + + def pick(offered: dict[str, str], preferences: list[str]) -> str: + for preference in preferences: + if preference in offered: + return preference + return next(iter(offered)) + + def rule(observation: dict[str, Any], offered: dict[str, str]) -> str: + if task == "normal": + if not state["unlocked"] and "click:unlock" in offered: + state["unlocked"] = True + return "click:unlock" + if "click:finish" in offered: + return "click:finish" + return pick(offered, ["click:unlock", "click:finish", "click:noop"]) + if task == "recoverable": + if observation.get("plan"): + state["go"] = True + if not state["go"]: + return pick(offered, ["click:noop", "click:unlock"]) + if not state["unlocked"] and "click:unlock" in offered: + state["unlocked"] = True + return "click:unlock" + if "click:finish" in offered: + return "click:finish" + return pick(offered, ["click:noop"]) + return pick(offered, ["click:noop", "click:unlock", "click:finish"]) + + return RuleModel(f"scripted-{arm}", rule) + + +def build_fixture(source: Path, out_exe: Path, *, csc: str, log: Path) -> Path: + """Compile the committed fixture source with the system .NET compiler; no installer, no new package.""" + out_exe.parent.mkdir(parents=True, exist_ok=True) + log.parent.mkdir(parents=True, exist_ok=True) + command = [csc, "/nologo", "/target:winexe", f"/out:{to_windows_path(out_exe)}", to_windows_path(source)] + with log.open("wb") as handle: + result = subprocess.run(command, stdout=handle, stderr=subprocess.STDOUT, check=False) + if result.returncode != 0 or not out_exe.exists(): + raise RuntimeError(f"csc failed (exit {result.returncode}); see {log}") + out_exe.chmod(out_exe.stat().st_mode | 0o111) # csc writes from Windows: make it runnable through WSL interop + return out_exe + + +def launch_fixture( + exe: Path, mode: str, run_id: str, outdir: Path, log: Path +) -> tuple[subprocess.Popen, dict[str, Any]]: + """Start this trial's own fixture process and wait for its startup oracle; the caller kills only its pid. + + WSL interop occasionally refuses a launch with a transient ``OSError`` (EINVAL) or the process dies before it + writes the oracle; a dead attempt is retried, and the unique exe name means no stale window can be mistaken for it. + """ + outdir.mkdir(parents=True, exist_ok=True) + log.parent.mkdir(parents=True, exist_ok=True) + oracle_file = outdir / "result.json" + last_error = "no attempt" + for attempt in range(1, LAUNCH_ATTEMPTS + 1): + if oracle_file.exists(): + oracle_file.unlink() + handle = log.open("ab" if attempt > 1 else "wb") + try: + proc = subprocess.Popen( + [str(exe), "--mode", mode, "--id", run_id, "--outdir", to_windows_path(outdir)], + stdout=handle, + stderr=subprocess.STDOUT, + stdin=subprocess.DEVNULL, + start_new_session=True, + ) + except OSError as exc: + handle.close() + last_error = f"launch attempt {attempt}: {exc}" + time.sleep(0.5) + continue + handle.close() + deadline = time.monotonic() + ORACLE_WAIT_S + while time.monotonic() < deadline: + if oracle_file.exists(): + try: + return proc, json.loads(oracle_file.read_text(encoding="utf-8-sig")) + except ValueError: + pass + if proc.poll() is not None: + break # the process died without writing the oracle: the attempt can be retried + time.sleep(0.25) + last_error = f"attempt {attempt} wrote no oracle (returncode={proc.poll()})" + if proc.poll() is None: # alive but never wrote: kill and fail rather than leak a second fixture + proc.terminate() + try: + proc.wait(timeout=5) + except Exception: + pass + handle.close() + raise RuntimeError(f"fixture {mode!r} stayed alive without writing {oracle_file}") + handle.close() + time.sleep(0.5) + raise RuntimeError(f"fixture {mode!r} never wrote {oracle_file}: {last_error}") + + +async def wait_for_window(driver: CuaDriver, app_name: str, pid: int) -> Any: + """Wait for this trial's fixture window to be on screen; the fixture writes its oracle before the window appears.""" + deadline = time.monotonic() + WINDOW_WAIT_S + last_error = "no window seen" + while time.monotonic() < deadline: + try: + window = await asyncio.wait_for( + driver.find_window(app_name), timeout=max(0.01, deadline - time.monotonic()) + ) + except DriverError as exc: + last_error = str(exc) + else: + if window.pid == pid: + return window + last_error = f"found window pid {window.pid} but launched fixture pid {pid}" + await asyncio.sleep(0.25) + raise RuntimeError(f"fixture window for {app_name!r} not ready: {last_error}") + + +def kill_pid(pid: int) -> None: + """Terminate exactly this run's fixture pid; never a generic image name, never another process.""" + if os.name == "nt": + subprocess.run(["taskkill", "/PID", str(pid), "/F"], stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) + elif Path(TASKKILL).exists(): # WSL: the fixture pid is a Windows pid + subprocess.run([TASKKILL, "/PID", str(pid), "/F"], stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) + else: + subprocess.run(["kill", "-9", str(pid)], stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) + + +def read_oracle(outdir: Path) -> dict[str, Any]: + try: + return json.loads((outdir / "result.json").read_text(encoding="utf-8-sig")) + except (OSError, ValueError): + return {} + + +def read_events(outdir: Path) -> list[dict[str, Any]]: + path = outdir / "app_events.jsonl" + events: list[dict[str, Any]] = [] + if path.exists(): + for line in path.read_text(encoding="utf-8-sig").splitlines(): + try: + events.append(json.loads(line)) + except ValueError: + pass + return events + + +def extract_field(output: str, pattern: re.Pattern[str]) -> str | None: + match = pattern.search(output or "") + if match is None: + return None + try: + return json.loads('"' + match.group(1) + '"') + except ValueError: + return match.group(1) + + +def count_wasted_actions(views: list[dict[str, Any]], steps: int) -> int: + """Accepted acts whose window progress digest did not change: the desktop no-op/wasted action count.""" + wasted = 0 + for index in range(min(steps, max(0, len(views) - 1))): + before = progress_digest(views[index].get("state")) + after = progress_digest(views[index + 1].get("state")) + wasted += before == after + return wasted + + +def _recovery_fields(episode: Episode, output: str, *, bounded: bool) -> dict[str, Any]: + events = [event for event in (episode.extra.get("rethinks") or []) if event.get("kind") == "stall"] + attempts = [int(event.get("attempt") or 0) for event in events] + spent = [float(event.get("spent_s") or 0.0) for event in events] + status = extract_field(output, _STATUS) + reason = extract_field(output, _REASON) or episode.error + negative = {"give_up", "error", "timeout", "cancelled"} + failed = bounded and ( + episode.error is not None + or status == "BLOCKED" + or episode.extra.get("result_type") == "timeout" + or any(event.get("termination") in negative for event in events) + ) + source: str | None = None + action = next((event["next_action"] for event in reversed(events) if event.get("next_action")), None) + if action is not None: + source = "event" + else: + terminal_action = extract_field(output, _NEXT_ACTION) + if terminal_action is not None: + action, source = terminal_action, "terminal" + return { + "recovery_attempts": max(attempts or [0]), + "recovery_spent_s": round(max(spent or [0.0]), 3), + "recovery_failed": failed, + "recovery_termination": ( + "task_timeout" + if episode.extra.get("result_type") == "timeout" + else "act_budget" + if failed and reason == "act budget spent" + else events[-1].get("termination") + if events + else None + ), + "recovery_next_action": action, + "next_action_source": source, + } + + +def _episode_to_dict(episode: Episode) -> dict[str, Any]: + data = asdict(episode) + data["frames_dir"] = None if episode.frames_dir is None else str(episode.frames_dir) + return data + + +def _episode_from_dict(data: dict[str, Any]) -> Episode: + frames = data.pop("frames_dir", None) + episode = Episode(**data) + episode.frames_dir = None if frames is None else Path(frames) + return episode + + +async def run_trial( + driver: CuaDriver, + exe: Path, + plan: TrialPlan, + *, + config: DesktopEvalConfig, + run_dir: Path, + run_id: str, +) -> tuple[Episode, DesktopTrialRecord]: + """One real desktop trial over its own fresh fixture; only this trial's pid is ever touched.""" + name = f"{plan.task}-rethink_{plan.arm}-r{plan.repeat}" + outdir = run_dir / "trials" / name + outdir.mkdir(parents=True, exist_ok=True) + log = run_dir / "logs" / f"trial-{name}.out" + expect = f"S1AFixture-{run_id} {DONE_TEXT}" + app_name = exe.name + planner = PlanChat() + rule = scripted_rule(plan.task, plan.arm) + started = time.perf_counter() + proc: subprocess.Popen | None = None + pid: int | None = None + try: + proc, oracle_start = launch_fixture(exe, fixture_mode(plan.task), run_id, outdir, log) + pid = int(oracle_start["pid"]) + await wait_for_window(driver, app_name, pid) # only this trial's verified pid is ever used + env = WindowEnv( + driver, + app_name=app_name, + goal=GOAL, + done_when=lambda snapshot: desktop.shows(snapshot, expect), + execute=True, + clear_labels=(), + ) + episode = await loop.run_episode( + desktop.SPEC, + env, + model_name="rule", + seed=plan.repeat, + chat=planner, + decision_model=rule, + rethink_on=(plan.arm == "on"), + max_acts=config.max_acts, + timeout_s=config.timeout_s, + prices=None, + log=False, + limits=( + RecoveryLimits(max_attempts=config.max_recovery_attempts, timeout_s=config.recovery_timeout_s) + if plan.arm == "on" + else None + ), + ) + elapsed_s = round(time.perf_counter() - started, 3) + # The models here are entirely scripted: no paid call or token usage occurred. + episode.cost_usd = 0.0 + episode.usage_known = True + output = str(episode.extra.get("output") or "") + final_oracle = read_oracle(outdir) + events = read_events(outdir) + event_counts: dict[str, int] = {} + for event in events: + event_counts[event.get("event", "?")] = event_counts.get(event.get("event", "?"), 0) + 1 + status = extract_field(output, _STATUS) + if episode.extra.get("result_type") == "timeout": + terminal = "timeout" + else: + terminal = status or "unknown" + record = DesktopTrialRecord( + task=plan.task, + arm=plan.arm, + repeat=plan.repeat, + seed=plan.repeat, + verified=final_oracle.get("finished") is True, + terminal=terminal, + errored=episode.error is not None or episode.extra.get("result_type") == "timeout", + scripted_model=SCRIPTED_MODEL, + driver={"bin": config.driver_bin, "version": config.driver_version, "channel": DRIVER_CHANNEL}, + platform=platform.platform(), + wasted_actions=count_wasted_actions(episode.views, episode.steps), + noop_clicks=event_counts.get("noop", 0), + steps=episode.steps, + decision_calls=len(episode.decisions), + planner_calls=planner.calls, + chat_calls=max(0, episode.chat_calls - planner.calls), + model_calls=len(episode.decisions) + episode.chat_calls, + elapsed_s=elapsed_s, + cost_usd=0.0, + oracle={ + "start": oracle_start, + "final": final_oracle, + "events": event_counts, + "distinct_pids": sorted({event.get("pid") for event in events if event.get("pid")}), + }, + **_recovery_fields(episode, output, bounded=(plan.arm == "on")), + ) + (outdir / "trial.json").write_text( + json.dumps( + {"record": record.as_json(), "episode": _episode_to_dict(episode)}, ensure_ascii=False, indent=2 + ), + encoding="utf-8", + ) + return episode, record + finally: + if pid is not None: + kill_pid(pid) + if proc is not None: + try: + proc.wait(timeout=5) + except Exception: + pass + + +def _failed_episode( + plan: TrialPlan, config: DesktopEvalConfig, reason: str, elapsed_s: float +) -> tuple[Episode, DesktopTrialRecord]: + """A trial that raised: recorded in the planned denominator, never dropped and never counted as success.""" + stamp = now_iso() + record = DesktopTrialRecord( + task=plan.task, + arm=plan.arm, + repeat=plan.repeat, + seed=plan.repeat, + verified=False, + terminal="timeout" if "timeout" in reason.lower() else f"error: {reason}", + errored=True, + scripted_model=SCRIPTED_MODEL, + driver={"bin": config.driver_bin, "version": config.driver_version, "channel": DRIVER_CHANNEL}, + platform=platform.platform(), + recovery_attempts=0, + recovery_spent_s=0.0, + recovery_failed=False, + recovery_termination=None, + recovery_next_action=None, + next_action_source=None, + wasted_actions=0, + noop_clicks=0, + steps=0, + model_calls=0, + decision_calls=0, + planner_calls=0, + chat_calls=0, + elapsed_s=elapsed_s, + cost_usd=0.0, + oracle={"verified": False, "failed": reason}, + ) + episode = Episode( + env="desktop_recovery", + policy=f"scripted-{plan.arm}", + seed=plan.repeat, + score=0.0, + steps=0, + elapsed_s=elapsed_s, + started_at=stamp, + finished_at=stamp, + final_state={"task": plan.task, "arm": plan.arm, "verified": False, "failed": reason}, + chat_calls=0, + chat_input_tokens=0, + chat_output_tokens=0, + chat_cache_tokens=0, + jev_input_tokens=0, + invalid_keys=0, + cost_usd=0.0, + error=reason, + extra=record.as_json(), + ) + return episode, record + + +def _validate(trials: list[TrialPlan]) -> None: + if not trials: + raise ValueError("at least one trial is required") + seen: set[tuple[str, str, int]] = set() + for plan in trials: + if plan.task not in TASKS: + raise ValueError(f"unknown task {plan.task!r}; expected any of {TASKS}") + if plan.arm not in ARMS: + raise ValueError(f"unknown arm {plan.arm!r}; expected any of {ARMS}") + if plan.repeat < 0: + raise ValueError(f"repeat must be >= 0, got {plan.repeat}") + key = (plan.task, plan.arm, plan.repeat) + if key in seen: + raise ValueError(f"duplicate trial {key}") + seen.add(key) + + +def _load_trial(outdir: Path) -> tuple[Episode, DesktopTrialRecord] | None: + path = outdir / "trial.json" + if not path.exists(): + return None + try: + data = json.loads(path.read_text(encoding="utf-8")) + except ValueError: + return None + record = DesktopTrialRecord(**_without_unknown(DesktopTrialRecord, data["record"])) + return _episode_from_dict(data["episode"]), record + + +def _without_unknown(dataclass_type: Any, payload: dict[str, Any]) -> dict[str, Any]: + names = set(dataclass_type.__dataclass_fields__) + return {key: value for key, value in payload.items() if key in names} + + +async def run_eval( + *, + trials: list[TrialPlan] | None = None, + config: DesktopEvalConfig | None = None, + repeat: int = 1, + run_id: str | None = None, + resume: bool = False, + results_dir: Path = RESULTS_DIR, +) -> DesktopEvalRun: + """Run every planned trial, write the job folders and the paired summary. Fresh by default; ``resume`` reuses.""" + trials = default_trials(repeat) if trials is None else list(trials) + _validate(trials) + config = config or default_config() + if not config.fixture_source.exists(): + raise RuntimeError(f"fixture source not found: {config.fixture_source}") + run_id = run_id or new_run_id() + run_dir = results_dir / "desktop_recovery" / run_id + run_dir.mkdir(parents=True, exist_ok=resume) + previous_workspace = loop.WORKSPACE + loop.WORKSPACE = run_dir / "agent-workspaces" + exe = run_dir / "build" / f"S1AFixture-{run_id}.exe" + pairs: list[tuple[Episode, DesktopTrialRecord]] = [] + try: + async with started_runner(): + driver = CuaDriver(config.driver_bin, session_label=config.session_label) + async with opened(driver): + if not exe.exists(): + build_fixture(config.fixture_source, exe, csc=config.csc, log=run_dir / "logs" / "csc.log") + for plan in trials: + name = f"{plan.task}-rethink_{plan.arm}-r{plan.repeat}" + outdir = run_dir / "trials" / name + cached = _load_trial(outdir) if resume else None + if cached is not None: + print(f" [resume] {name}: reuse {outdir / 'trial.json'}") + pairs.append(cached) + continue + started = time.perf_counter() + try: + pair = await run_trial(driver, exe, plan, config=config, run_dir=run_dir, run_id=run_id) + except Exception as exc: # noqa: BLE001 - a raised trial is a planned failure, not a drop + pair = _failed_episode( + plan, config, f"{type(exc).__name__}: {exc}", time.perf_counter() - started + ) + pairs.append(pair) + print( + f" [done] {name}: verified={pair[1].verified} terminal={pair[1].terminal} " + f"steps={pair[1].steps} planner_calls={pair[1].planner_calls} " + f"next_action={pair[1].next_action_source}" + ) + finally: + loop.WORKSPACE = previous_workspace + records = [record for _, record in pairs] + job_dirs: list[Path] = [] + for arm in ARMS: + episodes = [episode for episode, record in pairs if record.arm == arm] + if episodes: + job_dirs.append(write_job("desktop_recovery", episodes, results_dir=results_dir)) + summary = paired_summary(records) + (run_dir / "summary.json").write_text( + json.dumps(_json_summary(summary, records, run_id, config), ensure_ascii=False, indent=2), encoding="utf-8" + ) + (run_dir / "summary.md").write_text(render_markdown(summary), encoding="utf-8") + return DesktopEvalRun(summary=summary, records=records, run_dir=run_dir, job_dirs=job_dirs) + + +def run_sync(**kwargs: Any) -> DesktopEvalRun: + return asyncio.run(run_eval(**kwargs)) + + +def _rate(count: int, planned: int) -> float | None: + return round(count / planned, 3) if planned else None + + +def _arm_block(records: list[DesktopTrialRecord]) -> dict[str, Any]: + planned = len(records) + return { + "planned": planned, + "verified": sum(record.verified for record in records), + "errored": sum(record.errored for record in records), + "completion_rate": _rate(sum(record.verified for record in records), planned), + "recovery_attempts": sum(record.recovery_attempts for record in records), + "recovery_spent_s": round(sum(record.recovery_spent_s for record in records), 3), + "recovery_failed": sum(record.recovery_failed for record in records), + "wasted_actions": sum(record.wasted_actions for record in records), + "noop_clicks": sum(record.noop_clicks for record in records), + "mean_model_calls": round(statistics.mean(r.model_calls for r in records), 2) if records else None, + "mean_elapsed_s": round(statistics.mean(r.elapsed_s for r in records), 3) if records else None, + "cost_usd": 0.0, + } + + +def paired_summary(records: list[DesktopTrialRecord]) -> dict[str, Any]: + """Every planned trial counts, errors included. Success is each trial's own oracle, never the model verdict.""" + arms = {arm: _arm_block([r for r in records if r.arm == arm]) for arm in ARMS} + tasks = sorted({record.task for record in records}) + by_task = { + task: {arm: _arm_block([r for r in records if r.arm == arm and r.task == task]) for arm in ARMS} + for task in tasks + } + pairs: dict[tuple[str, int], dict[str, DesktopTrialRecord]] = {} + for record in records: + pairs.setdefault((record.task, record.repeat), {})[record.arm] = record + both = off_only = on_only = neither = incomplete = 0 + delta_elapsed: list[float] = [] + delta_wasted: list[int] = [] + for arm_records in pairs.values(): + off, on = arm_records.get("off"), arm_records.get("on") + if off is None or on is None: + incomplete += 1 + continue + delta_elapsed.append(round(on.elapsed_s - off.elapsed_s, 3)) + delta_wasted.append(on.wasted_actions - off.wasted_actions) + if off.verified and on.verified: + both += 1 + elif on.verified: + on_only += 1 + elif off.verified: + off_only += 1 + else: + neither += 1 + off_rate = arms["off"]["verified"] / arms["off"]["planned"] if arms["off"]["planned"] else None + on_rate = arms["on"]["verified"] / arms["on"]["planned"] if arms["on"]["planned"] else None + return { + "planned_trials": len(records), + "arms": arms, + "by_task": by_task, + "paired": { + "pairs": len(pairs) - incomplete, + "incomplete_pairs": incomplete, + "both_verified": both, + "off_only_verified": off_only, + "on_only_verified": on_only, + "neither_verified": neither, + "mean_delta_elapsed_s": round(statistics.mean(delta_elapsed), 3) if delta_elapsed else None, + "mean_delta_wasted_actions": round(statistics.mean(delta_wasted), 2) if delta_wasted else None, + }, + "on_minus_off_completion_rate": ( + round(on_rate - off_rate, 3) if on_rate is not None and off_rate is not None else None + ), + "note": ( + "controlled fault injection with a scripted decision model and a scripted planner on a self-built " + "fixture: it exercises the recovery mechanism and budget, not a trained model and not a real app success " + "rate. Success is the fixture's own result.json, never the model's DONE. Errors count in every " + "denominator; cost is 0 by construction because no paid API is called." + ), + } + + +def render_markdown(summary: dict[str, Any]) -> str: + lines = [ + "# Desktop bounded recovery on/off - scripted fixture subset", + "", + f"Planned trials: {summary['planned_trials']} (errors included).", + "", + "| arm | planned | verified | completion | errored | recovery attempts | recovery failed | wasted actions | mean elapsed s |", + "|---|---|---|---|---|---|---|---|---|", + ] + for arm in ARMS: + block = summary["arms"][arm] + lines.append( + f"| {arm} | {block['planned']} | {block['verified']} | {block['completion_rate']} | {block['errored']} | " + f"{block['recovery_attempts']} | {block['recovery_failed']} | {block['wasted_actions']} | " + f"{block['mean_elapsed_s']} |" + ) + lines += ["", "| task | arm | planned | verified | completion |", "|---|---|---|---|---|"] + for task, arms in summary["by_task"].items(): + for arm in ARMS: + block = arms[arm] + lines.append(f"| {task} | {arm} | {block['planned']} | {block['verified']} | {block['completion_rate']} |") + paired = summary["paired"] + lines += [ + "", + "Paired completion per (task, repeat): " + f"on-only {paired['on_only_verified']}, off-only {paired['off_only_verified']}, " + f"both {paired['both_verified']}, neither {paired['neither_verified']}.", + f"On minus off completion rate: {summary['on_minus_off_completion_rate']}.", + "", + f"Note: {summary['note']}", + "", + ] + return "\n".join(lines) + + +def _json_summary( + summary: dict[str, Any], records: list[DesktopTrialRecord], run_id: str, config: DesktopEvalConfig +) -> dict[str, Any]: + return { + "run_id": run_id, + "summary": summary, + "trials": [record.as_json() for record in records], + "definition": { + "planned_trials": "every (task, arm, repeat) the run planned; errors included", + "verified": "this trial's own fixture result.json has finished true (never the model's DONE)", + "oracle": "per-trial fixture output directory; a fresh fixture is launched for every trial", + "wasted_actions": "accepted acts whose window progress digest did not change (the desktop no-op count)", + "noop_clicks": "the fixture's own logged clicks on its no-op control", + "elapsed_s": "trial wall clock including fixture startup and the agent episode; excludes shared driver and compiler startup", + "model_calls": "decision_calls + chat_calls (all scripted doubles)", + "next_action": "the actionable escalation the terminal carried for a failed bounded recovery", + "next_action_source": "where a recorded next_action came from: a recovery event or the terminal; never generated by the evaluator", + "cost_usd": "0 by construction: scripted decisions and planner, no paid API call", + }, + "budget": { + "max_acts": config.max_acts, + "timeout_s": config.timeout_s, + "max_recovery_attempts": config.max_recovery_attempts, + "recovery_timeout_s": config.recovery_timeout_s, + }, + "driver": {"bin": config.driver_bin, "version": config.driver_version, "channel": DRIVER_CHANNEL}, + "scripted_model": SCRIPTED_MODEL, + } + + +def parser() -> argparse.ArgumentParser: + argument_parser = argparse.ArgumentParser(description="Native desktop bounded-recovery on/off subset eval.") + argument_parser.add_argument("--driver", default=None, help="path to cua-driver (or S1A_DESKTOP_DRIVER)") + argument_parser.add_argument("--csc", default=None, help="path to the system .NET C# compiler (or S1A_DESKTOP_CSC)") + argument_parser.add_argument("--repeat", type=int, default=1, help="repeats per (task, arm); >= 1") + argument_parser.add_argument("--tasks", default=",".join(TASKS), help=f"comma-separated subset of {TASKS}") + argument_parser.add_argument("--arms", default=",".join(ARMS), help=f"comma-separated subset of {ARMS}") + argument_parser.add_argument("--run-id", default=None, help="reuse a run id; with --resume, continue it") + argument_parser.add_argument("--resume", action="store_true", help="reuse existing trial.json files in the run") + argument_parser.add_argument("--results-dir", default=str(RESULTS_DIR), help="where the run directory is created") + return argument_parser + + +def main(argv: list[str] | None = None) -> int: + args = parser().parse_args(argv) + tasks = tuple(part.strip() for part in args.tasks.split(",") if part.strip()) + arms = tuple(part.strip() for part in args.arms.split(",") if part.strip()) + if args.repeat < 1: + raise SystemExit("--repeat must be >= 1") + config = default_config(driver_bin=args.driver, csc=args.csc) + run = run_sync( + trials=default_trials(args.repeat, tasks, arms), + config=config, + run_id=args.run_id, + resume=args.resume, + results_dir=Path(args.results_dir), + ) + print(render_markdown(run.summary)) + print(f"run dir: {run.run_dir}") + for job_dir in run.job_dirs: + print(f"job: {job_dir}") + return 1 if any(record.errored for record in run.records) else 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/evals/recovery/RESULTS.md b/evals/recovery/RESULTS.md new file mode 100644 index 0000000..992b25c --- /dev/null +++ b/evals/recovery/RESULTS.md @@ -0,0 +1,64 @@ +# Small real-model recovery experiment + +This is a fixture experiment, not an open-web benchmark. Each matrix has three synthetic forms, +two recovery arms, and two repeats (12 planned trials). Every planned trial finished; no failure +or timeout was dropped. A fresh server's recorded form POST decides completion independently of +the agent's final text. The [per-trial data](laya-results.json) includes actual submitted values, +calls, timings, failures and source hashes; machine paths and credentials are omitted. + +The decision model was Laya 0.3.5 with the fixed English checkpoint revision listed in the data, +running on four CPU threads. Chat/planning used OpenCode Go DeepSeek v4.1 Flash, temperature 0, +thinking disabled and no SDK retries. The production planner retained its 120-token output cap. +Each task had 180 seconds and 24 steps; recovery had three attempts and 45 seconds of cumulative +refresh/planning time. The second repeat reversed arm order. Model loading was excluded from +per-trial elapsed time. + +## Original forms + +| Arm | Verified | Planner attempts | Decision calls | Chat calls | Mean elapsed (s) | Unchanged-page actions | +| --- | ---: | ---: | ---: | ---: | ---: | ---: | +| off | 0/6 | 0 | 18 | 5 | 30.349 | 8 | +| on | 0/6 | 0 | 18 | 4 | 30.122 | 8 | + +Recovery never triggered. The model submitted an incorrect value or terminated before the stall +threshold. Eight terminal states were DONE despite the independent oracle rejecting completion. +This matrix provides no evidence of a recovery benefit. + +## Exploratory server-validation follow-up + +After observing premature submissions, an optional variant was planned: an incorrect POST returns +HTTP 422 and a retryable form. Default fixtures were left unchanged. Decisions and plans were still +generated by real models, with no injected action outputs. This follow-up is reported separately. + +| Arm | Verified | Planner attempts | Decision calls | Chat calls | Mean elapsed (s) | Unchanged-page actions | +| --- | ---: | ---: | ---: | ---: | ---: | ---: | +| off | 0/6 | 0 | 32 | 9 | 75.623 | 24 | +| on | 2/6 | 6 | 44 | 14 | 93.818 | 30 | + +The two successful on trials were the normal form. The field had already been filled correctly, +but repeated interactions stalled. A refreshed observation led the planner to suggest Enter, and +the decision model selected the ordinary keyboard submission tool. The server then received the +expected value. In a locked-form trial, two plans correctly suggested Enable editing, but the +decision model never selected that button and eventually returned BLOCKED. + +Paired mean overhead was 18.194 seconds, two decision calls, 0.83 chat calls and one additional +unchanged-page action. Planner attempts are already included in chat calls. Recovery increased +work and did not solve every case. With one model and two repeats per task, these observations +cannot establish a general success-rate improvement or reliable population-level timing estimate. + +The experiment also exposed a reporting bug: a nonempty partial answer could make post-replan +BLOCKED appear successful in the browser front. That terminal-reporting correction was made and +regression-tested afterwards. The rows preserve the measured revision; no historical outcomes +were rewritten. The independent oracle already treated those trials as failures. + +## Reproduction + +Use the live runner described in [the eval guide](../README.md). Run the original matrix with +`--repeat 2 --timeout 180 --max-steps 24 --recovery-timeout 45`; add `--validate-submission` for the +separate exploratory matrix. Use the recorded Laya context lengths and checkpoint. The custom +Go chat object was supplied through `run_live(chat=...)` with `extra_body={"thinking":{"type":"disabled"}}`, +rather than changing the production planner's output budget. Match provider configuration when +reproducing; another model or provider is a different experiment. + +The controlled Chromium and Windows suites use scripted fault decisions/plans to validate bounds +and wiring. Their completion counts must not be mixed with this real-model experiment. diff --git a/evals/recovery/__init__.py b/evals/recovery/__init__.py new file mode 100644 index 0000000..5e86480 --- /dev/null +++ b/evals/recovery/__init__.py @@ -0,0 +1,9 @@ +# coding: utf-8 +"""The repeatable bounded-recovery on/off subset: three local form tasks through the real browser front. + +This package is an evaluation, not a runner for production code. It serves three fixture pages from a localhost +server, drives each one through the repository's ``browse`` front with a clearly-marked scripted decision model and +scripted planner, and judges completion only by the fixture server's own record of the real submit POST. Each trial +opens its own fresh fixture, so no earlier trial's submit can verify a later one. It never reads the model's +DONE/answer to decide success. See ``evals/recovery/__main__.py`` and ``evals/README.md``. +""" diff --git a/evals/recovery/__main__.py b/evals/recovery/__main__.py new file mode 100644 index 0000000..5fec1ed --- /dev/null +++ b/evals/recovery/__main__.py @@ -0,0 +1,63 @@ +# coding: utf-8 +"""``python -m evals.recovery [--repeat 1] [--arms off on] [--tasks normal recoverable permanently_blocked]``. + +Runs the repeatable bounded-recovery on/off subset against the local fixture pages in a real headless Chromium, +writes the Harbor job folders and a paired JSON + Markdown summary, and prints the summary. Defaults to one repeat; +raise ``--repeat`` for more pairs, which are then averaged. Needs Node/``npx`` and the Playwright MCP package, the +same dependency ``tests/system`` uses. The decision model and planner are scripted doubles, so this is a controlled +mechanism test, not a real model benchmark. Exits non-zero when any planned trial raised (a harness error or timeout); +expected terminal scores such as BLOCKED do not make the process fail. +""" + +from __future__ import annotations + +import s1a.entry # noqa: F401 # routes the harness logs to files before anything imports openjiuwen +import argparse +from pathlib import Path + +from s1a.jobs import RESULTS_DIR + +from evals.recovery.fixture import DEFAULT_TASKS +from evals.recovery.runner import EvalConfig, format_run, run_sync + +TASKS = {task.name: task for task in DEFAULT_TASKS} + + +def parser() -> argparse.ArgumentParser: + build = argparse.ArgumentParser(prog="python -m evals.recovery", description=__doc__) + build.add_argument("--repeat", type=int, default=1, help="repeats per (task, arm), paired; must be >= 1") + build.add_argument("--tasks", nargs="+", choices=sorted(TASKS), default=sorted(TASKS), help="which fixture tasks") + build.add_argument("--arms", nargs="+", choices=("off", "on"), default=["off", "on"], help="recovery arms") + build.add_argument("--headed", action="store_true", help="show the browser (default: headless)") + build.add_argument("--timeout", type=float, default=120.0, help="seconds per trial") + build.add_argument("--max-steps", type=int, default=24, help="the subagent's iteration cap") + build.add_argument("--stall-after", type=int, default=3, help="no-op actions before a stall") + build.add_argument("--recovery-attempts", type=int, default=3, help="bounded recovery attempts per trial") + build.add_argument("--recovery-timeout", type=float, default=15.0, help="bounded recovery active seconds per trial") + build.add_argument("--results-dir", type=Path, default=RESULTS_DIR, help="root of the Harbor job folders") + return build + + +def main() -> int: + args = parser().parse_args() + config = EvalConfig( + timeout_s=args.timeout, + max_steps=args.max_steps, + stall_after=args.stall_after, + max_recovery_attempts=args.recovery_attempts, + recovery_timeout_s=args.recovery_timeout, + headless=not args.headed, + ) + run = run_sync( + repeat=args.repeat, + tasks=tuple(TASKS[name] for name in args.tasks), + arms=tuple(args.arms), + config=config, + results_dir=args.results_dir, + ) + print(format_run(run)) + return 1 if any(record.errored for record in run.records) else 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/evals/recovery/fixture.py b/evals/recovery/fixture.py new file mode 100644 index 0000000..fb938f0 --- /dev/null +++ b/evals/recovery/fixture.py @@ -0,0 +1,278 @@ +# coding: utf-8 +"""The local form fixtures and the independent submit oracle for the recovery eval. + +Three tasks, one page each, served from 127.0.0.1 on a free port from a daemon thread (the same shape as +``s1a.tool.hands.serve_static``): + +- ``normal``: the field keeps what it is typed; a direct type-and-submit completes it. +- ``recoverable``: an ``input`` handler restores the initial value until the ``Enable editing`` button is clicked, + which flips a page flag (and removes the button, changing the page key). Typing into the restored field is a + no-op at the page level, so it stalls the policy; the unlock needs a different action. +- ``permanently_blocked``: the field always restores its initial value and the page offers no unlock control, so the + policy should stop clearly after its recovery budget is spent. + +The server records every real ``POST /submit/`` with the form fields it received. That record is the only +success criterion the eval uses; a model's DONE or answer is never read. The pages themselves are the environment, +and the recorded POST is the independent oracle. + +One fixture server is one trial's oracle: ``start_fixture`` mints a unique ``trial_id`` and starts an empty record +namespace, so a submit from an earlier trial can never verify a later one. The runner starts a fresh fixture inside +``run_trial`` for exactly this reason, and every recorded submit carries the trial id it belongs to. +""" + +from __future__ import annotations + +import dataclasses +import html +import http.server +import itertools +import json +import socketserver +import threading +from contextlib import contextmanager +from dataclasses import dataclass +from typing import Any, Iterator +from urllib.parse import parse_qs, urlencode, urlparse + +_TRIAL_IDS = itertools.count(1) + + +@dataclass(frozen=True) +class FormTask: + """One fixture task: its page, the field value a real person would submit, and whether a page unlock exists. + + ``validate_submission`` is off by default, so the original default tasks keep their exact original behaviour. An + opt-in task rejects a submitted value that differs from ``expected_value`` at the server (HTTP 422) and re-serves + the same retryable form with a visible validation error; a correct value still returns the result page. The oracle + records every real POST either way. + """ + + name: str + path: str # the URL path the page is served at + title: str + expected_value: str # the value the oracle looks for in the recorded submit + kind: str # normal | recoverable | blocked: how the page reacts to typing + validate_submission: bool = False # opt-in server-side validation; the default tasks leave it off + + def __post_init__(self) -> None: + if self.kind not in ("normal", "recoverable", "blocked"): + raise ValueError(f"task {self.name}: kind must be normal, recoverable or blocked, got {self.kind!r}") + if not self.path.startswith("/") or not self.expected_value.strip(): + raise ValueError(f"task {self.name}: path must start with / and expected_value must not be blank") + + +NORMAL = FormTask("normal", "/normal", "Recovery normal", "hello world", "normal") +RECOVERABLE = FormTask("recoverable", "/recoverable", "Recovery recoverable", "hello world", "recoverable") +BLOCKED = FormTask("permanently_blocked", "/blocked", "Recovery permanently blocked", "hello world", "blocked") +# The DEFAULT_TASKS are retained unchanged: every field is as before and validate_submission stays False. +DEFAULT_TASKS: tuple[FormTask, ...] = (NORMAL, RECOVERABLE, BLOCKED) + + +def validated_tasks(tasks: tuple[FormTask, ...]) -> tuple[FormTask, ...]: + """The same tasks with server-side validation on; ``FormTask`` is frozen, so the originals are never mutated.""" + return tuple(dataclasses.replace(task, validate_submission=True) for task in tasks) + + +def validation_error(task: FormTask, submitted: str) -> str: + """The ordinary visible message a rejected form shows; it echoes the submitted value, never the expected one.""" + return f"The value '{submitted}' is not valid for {task.name}. Fix the field and submit again." + + +_REVERT_SCRIPT = """ +const input = document.getElementById('value'); +window.__editing = false; +input.addEventListener('input', function () { if (!window.__editing) { input.value = 'original'; } }); +""" +_ENABLE_SCRIPT = """ +document.getElementById('enable').addEventListener('click', function () { + window.__editing = true; + document.getElementById('status').textContent = 'Editing enabled'; + this.remove(); +}); +""" + + +def page_html(task: FormTask, *, error: str | None = None) -> str: + """The fixture page: a labelled field, a Submit button, and the task's own revert/unlock behaviour. + + ``error`` re-renders the retryable form with a visible validation message (HTML escaped). With no error the page + is byte-for-byte the original one, so the default fixture is unchanged. + """ + initial = "" if task.kind == "normal" else "original" + status = '

Locked

\n' if task.kind != "normal" else "" + enable = '\n' if task.kind == "recoverable" else "" + banner = f'\n' if error else "" + script = "" + if task.kind == "blocked": + script = f"" + elif task.kind == "recoverable": + script = f"" + return ( + "" + f"{task.title}\n" + f"

{task.title}

\n" + f"{status}" + f"{banner}" + f"
\n" + f"\n" + f"\n" + f"{enable}" + "\n" + "
\n" + f"{script}\n" + "" + ) + + +def result_html(task_name: str, value: str) -> str: + return ( + "" + f"Submitted {task_name}\n" + f"

Submitted

value={value}

\n" + "" + ) + + +class _Handler(http.server.BaseHTTPRequestHandler): + """Serve the fixture pages and record the real form POSTs; everything else is a 404.""" + + server: "_Server" # set by the server factory; the type checker cannot see BaseHTTPRequestHandler's server + + def log_message(self, format: str, *args: Any) -> None: + return None + + def _send(self, status: int, body: str, content_type: str = "text/html; charset=utf-8") -> None: + encoded = body.encode("utf-8") + self.send_response(status) + self.send_header("Content-Type", content_type) + self.send_header("Content-Length", str(len(encoded))) + self.end_headers() + self.wfile.write(encoded) + + def do_GET(self) -> None: # noqa: N802 - http.server's spelling + path = urlparse(self.path).path + task = self.server.tasks.get(path) + if task is not None: + self._send(200, page_html(task)) + return + if path == "/": + rows = "".join(f"
  • {t.name}
  • " for t in self.server.tasks.values()) + self._send(200, f"
      {rows}
    ") + return + if path == "/oracle": + self._send(200, json.dumps(self.server.submission_records(), ensure_ascii=False), "application/json") + return + self._send(404, "

    404

    ") + + def do_POST(self) -> None: # noqa: N802 - http.server's spelling + path = urlparse(self.path).path + if not path.startswith("/submit/"): + self._send(404, "

    404

    ") + return + task_name = path[len("/submit/") :] + length = int(self.headers.get("Content-Length") or 0) + raw = self.rfile.read(length).decode("utf-8", errors="replace") if length else "" + fields = {key: values[-1] for key, values in parse_qs(raw).items()} + self.server.record_submission(task_name, fields) # every real POST is recorded, valid or not + task = self.server.tasks_by_name.get(task_name) + value = fields.get("value", "") + if task is not None and task.validate_submission and value != task.expected_value: + # A rejected submit stays on the retryable form: no success result page and no oracle completion. + self._send(422, page_html(task, error=validation_error(task, value))) + return + self._send(200, result_html(task_name, value)) + + +class _Server(socketserver.ThreadingTCPServer): + allow_reuse_address = True + daemon_threads = True + + def __init__(self, address: tuple[str, int], handler: Any, tasks: dict[str, FormTask], trial_id: str) -> None: + super().__init__(address, handler) + self.tasks = tasks + self.tasks_by_name = {task.name: task for task in tasks.values()} + self.trial_id = trial_id + self._submissions: list[dict[str, Any]] = [] + self._lock = threading.Lock() + + def record_submission(self, task_name: str, fields: dict[str, str]) -> None: + with self._lock: + self._submissions.append({"trial": self.trial_id, "task": task_name, "fields": fields}) + + def submission_records(self) -> list[dict[str, Any]]: + with self._lock: + return [dict(record) for record in self._submissions] + + +class Fixture: + """A running fixture server for exactly one trial: its base URL, its own empty submit namespace and the oracle.""" + + def __init__(self, server: _Server) -> None: + self._server = server + self.trial_id = server.trial_id + self.base_url = f"http://127.0.0.1:{server.server_address[1]}" + + def url(self, path: str) -> str: + return f"{self.base_url}{path}" + + def task_url(self, task: FormTask) -> str: + return self.url(task.path) + + def submissions(self, task: FormTask | None = None) -> list[dict[str, Any]]: + """Copies of this trial's recorded real submits, optionally just the named task's.""" + records = self._server.submission_records() + if task is not None: + records = [record for record in records if record["task"] == task.name] + return records + + def verified(self, task: FormTask) -> bool: + """The independent oracle: the server saw a real submit carrying the task's expected field value.""" + return any(record["fields"].get("value") == task.expected_value for record in self.submissions(task)) + + def oracle_state(self, task: FormTask) -> dict[str, Any]: + """What this trial's oracle holds: its id, the recorded submits, and whether any verifies completion.""" + return { + "trial": self.trial_id, + "task": task.name, + "verified": self.verified(task), + "submissions": self.submissions(task), + } + + +@contextmanager +def start_fixture(tasks: tuple[FormTask, ...] = DEFAULT_TASKS) -> Iterator[Fixture]: + """Yield a one-trial Fixture: a fresh, empty oracle serving ``tasks`` on a free loopback port. + + The unique ``trial_id`` names this trial and every submit it records, so a later fixture starts with no history + and a previous trial's success can never verify a later one. The runner opens one of these per trial. + """ + table = {task.path: task for task in tasks} + if len(table) != len(tasks): + raise ValueError("every fixture task needs a distinct path") + trial_id = f"trial-{next(_TRIAL_IDS):04d}" + server = _Server(("127.0.0.1", 0), _Handler, table, trial_id) + threading.Thread(target=server.serve_forever, daemon=True).start() + try: + yield Fixture(server) + finally: + server.shutdown() + server.server_close() + + +def post_submit(fixture: Fixture, task: FormTask, value: str) -> tuple[int, str]: + """A plain HTTP form POST against a fixture; returns the status and body for tests of the oracle without a browser. + + A 422 (a validated task rejecting a mismatched value) is a normal response here, not an exception. + """ + from urllib.error import HTTPError + from urllib.request import Request, urlopen + + body = urlencode({"value": value}).encode() + request = Request( + fixture.url(f"/submit/{task.name}"), data=body, headers={"Content-Type": "application/x-www-form-urlencoded"} + ) + try: + with urlopen(request, timeout=10) as response: # noqa: S310 - a loopback URL built by this module + return response.status, response.read().decode("utf-8", errors="replace") + except HTTPError as exc: # a rejected submit is an expected status, so read its body and return it + return exc.code, exc.read().decode("utf-8", errors="replace") diff --git a/evals/recovery/laya-results.json b/evals/recovery/laya-results.json new file mode 100644 index 0000000..1646492 --- /dev/null +++ b/evals/recovery/laya-results.json @@ -0,0 +1,763 @@ +{ + "date": "2026-09-28", + "kind": "small synthetic browser fixture experiment", + "decision": { + "library": "laya", + "version": "0.3.5", + "checkpoint": "convaiinnovations/laya", + "revision": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851", + "device": "cpu", + "threads": 4, + "max_len": 4096, + "head_max_len": 1536 + }, + "chat": { + "provider": "OpenCode Go", + "model": "deepseek-v4.1-flash", + "temperature": 0, + "thinking": "disabled", + "sdk_retries": 0, + "planner_max_tokens": 120 + }, + "protocol": { + "repeat": 2, + "timeout_s": 180, + "max_steps": 24, + "stall_after": 3, + "recovery_attempts_limit": 3, + "recovery_active_time_limit_s": 45, + "order": "off/on in repeat 0; on/off in repeat 1; fresh browser and independent server oracle per trial" + }, + "source": { + "base_git": "34fc6a69263e05119358fb152d3f39d2d79d91dd", + "original_bundle_sha256": "192c58cae58f58ec700041fdc1518e0086134bfcd9b89f43b58cf97673364075", + "validated_delta_sha256": "a4e7b546cd80b0ae515d4f185e2b068cf1b481aa51a4da9f44f1be142090262b", + "note": "Uncommitted experimental source. The final post-replan BLOCKED verdict correction was subsequently regression-tested; it changes terminal reporting, not the recorded actions, calls or budgets. Historical rows are unchanged." + }, + "original": [ + { + "task": "normal", + "arm": "off", + "repeat": 0, + "verified": false, + "terminal": "DONE", + "errored": false, + "error": null, + "elapsed_s": 13.263, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 2, + "decision_failures": 0, + "decision_ms_total": 7782, + "chat_calls": 2, + "chat_failures": 0, + "chat_ms_total": 4339, + "chat_input_tokens": 309, + "chat_output_tokens": 28, + "usage_known": true, + "wasted_actions": 0, + "steps": 1, + "cost_usd": null, + "submitted_values": [ + "" + ] + }, + { + "task": "normal", + "arm": "on", + "repeat": 0, + "verified": false, + "terminal": "DONE", + "errored": false, + "error": null, + "elapsed_s": 12.612, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 2, + "decision_failures": 0, + "decision_ms_total": 7761, + "chat_calls": 1, + "chat_failures": 0, + "chat_ms_total": 1521, + "chat_input_tokens": 109, + "chat_output_tokens": 27, + "usage_known": true, + "wasted_actions": 0, + "steps": 1, + "cost_usd": null, + "submitted_values": [ + "" + ] + }, + { + "task": "normal", + "arm": "on", + "repeat": 1, + "verified": false, + "terminal": "DONE", + "errored": false, + "error": null, + "elapsed_s": 12.821, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 2, + "decision_failures": 0, + "decision_ms_total": 7748, + "chat_calls": 1, + "chat_failures": 0, + "chat_ms_total": 1793, + "chat_input_tokens": 109, + "chat_output_tokens": 26, + "usage_known": true, + "wasted_actions": 0, + "steps": 1, + "cost_usd": null, + "submitted_values": [ + "" + ] + }, + { + "task": "normal", + "arm": "off", + "repeat": 1, + "verified": false, + "terminal": "DONE", + "errored": false, + "error": null, + "elapsed_s": 12.647, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 2, + "decision_failures": 0, + "decision_ms_total": 7768, + "chat_calls": 1, + "chat_failures": 0, + "chat_ms_total": 1600, + "chat_input_tokens": 109, + "chat_output_tokens": 25, + "usage_known": true, + "wasted_actions": 0, + "steps": 1, + "cost_usd": null, + "submitted_values": [ + "" + ] + }, + { + "task": "recoverable", + "arm": "off", + "repeat": 0, + "verified": false, + "terminal": "DONE", + "errored": false, + "error": null, + "elapsed_s": 42.534, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 4, + "decision_failures": 0, + "decision_ms_total": 30019, + "chat_calls": 1, + "chat_failures": 0, + "chat_ms_total": 1683, + "chat_input_tokens": 111, + "chat_output_tokens": 25, + "usage_known": true, + "wasted_actions": 2, + "steps": 3, + "cost_usd": null, + "submitted_values": [ + "original" + ] + }, + { + "task": "recoverable", + "arm": "on", + "repeat": 0, + "verified": false, + "terminal": "DONE", + "errored": false, + "error": null, + "elapsed_s": 42.301, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 4, + "decision_failures": 0, + "decision_ms_total": 29871, + "chat_calls": 1, + "chat_failures": 0, + "chat_ms_total": 1610, + "chat_input_tokens": 111, + "chat_output_tokens": 29, + "usage_known": true, + "wasted_actions": 2, + "steps": 3, + "cost_usd": null, + "submitted_values": [ + "original" + ] + }, + { + "task": "recoverable", + "arm": "on", + "repeat": 1, + "verified": false, + "terminal": "DONE", + "errored": false, + "error": null, + "elapsed_s": 42.418, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 4, + "decision_failures": 0, + "decision_ms_total": 30086, + "chat_calls": 1, + "chat_failures": 0, + "chat_ms_total": 1479, + "chat_input_tokens": 111, + "chat_output_tokens": 29, + "usage_known": true, + "wasted_actions": 2, + "steps": 3, + "cost_usd": null, + "submitted_values": [ + "original" + ] + }, + { + "task": "recoverable", + "arm": "off", + "repeat": 1, + "verified": false, + "terminal": "DONE", + "errored": false, + "error": null, + "elapsed_s": 42.989, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 4, + "decision_failures": 0, + "decision_ms_total": 30358, + "chat_calls": 1, + "chat_failures": 0, + "chat_ms_total": 1635, + "chat_input_tokens": 111, + "chat_output_tokens": 29, + "usage_known": true, + "wasted_actions": 2, + "steps": 3, + "cost_usd": null, + "submitted_values": [ + "original" + ] + }, + { + "task": "permanently_blocked", + "arm": "off", + "repeat": 0, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 35.268, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 3, + "decision_failures": 0, + "decision_ms_total": 26210, + "chat_calls": 0, + "chat_failures": 0, + "chat_ms_total": 0, + "chat_input_tokens": 0, + "chat_output_tokens": 0, + "usage_known": true, + "wasted_actions": 2, + "steps": 2, + "cost_usd": null, + "submitted_values": [] + }, + { + "task": "permanently_blocked", + "arm": "on", + "repeat": 0, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 35.298, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 3, + "decision_failures": 0, + "decision_ms_total": 26207, + "chat_calls": 0, + "chat_failures": 0, + "chat_ms_total": 0, + "chat_input_tokens": 0, + "chat_output_tokens": 0, + "usage_known": true, + "wasted_actions": 2, + "steps": 2, + "cost_usd": null, + "submitted_values": [] + }, + { + "task": "permanently_blocked", + "arm": "on", + "repeat": 1, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 35.28, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 3, + "decision_failures": 0, + "decision_ms_total": 26136, + "chat_calls": 0, + "chat_failures": 0, + "chat_ms_total": 0, + "chat_input_tokens": 0, + "chat_output_tokens": 0, + "usage_known": true, + "wasted_actions": 2, + "steps": 2, + "cost_usd": null, + "submitted_values": [] + }, + { + "task": "permanently_blocked", + "arm": "off", + "repeat": 1, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 35.392, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 3, + "decision_failures": 0, + "decision_ms_total": 26238, + "chat_calls": 0, + "chat_failures": 0, + "chat_ms_total": 0, + "chat_input_tokens": 0, + "chat_output_tokens": 0, + "usage_known": true, + "wasted_actions": 2, + "steps": 2, + "cost_usd": null, + "submitted_values": [] + } + ], + "validated_exploratory": [ + { + "task": "normal", + "arm": "off", + "repeat": 0, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 85.418, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 7, + "decision_failures": 0, + "decision_ms_total": 56596, + "chat_calls": 4, + "chat_failures": 0, + "chat_ms_total": 9230, + "chat_input_tokens": 922, + "chat_output_tokens": 56, + "usage_known": true, + "wasted_actions": 5, + "steps": 7, + "cost_usd": null, + "submitted_values": [ + "", + "", + "" + ] + }, + { + "task": "normal", + "arm": "on", + "repeat": 0, + "verified": true, + "terminal": "verified", + "errored": false, + "error": null, + "elapsed_s": 97.097, + "recovery_attempts": 1, + "recovery_spent_s": 2.66, + "recovery_termination": "planned", + "planner_attempts": 1, + "planner_failures": 0, + "decision_calls": 9, + "decision_failures": 0, + "decision_ms_total": 66754, + "chat_calls": 4, + "chat_failures": 0, + "chat_ms_total": 6754, + "chat_input_tokens": 1590, + "chat_output_tokens": 70, + "usage_known": true, + "wasted_actions": 5, + "steps": 8, + "cost_usd": null, + "submitted_values": [ + "", + "", + "", + "hello world" + ] + }, + { + "task": "normal", + "arm": "on", + "repeat": 1, + "verified": true, + "terminal": "verified", + "errored": false, + "error": null, + "elapsed_s": 96.871, + "recovery_attempts": 1, + "recovery_spent_s": 2.443, + "recovery_termination": "planned", + "planner_attempts": 1, + "planner_failures": 0, + "decision_calls": 9, + "decision_failures": 0, + "decision_ms_total": 66681, + "chat_calls": 4, + "chat_failures": 0, + "chat_ms_total": 6614, + "chat_input_tokens": 1590, + "chat_output_tokens": 66, + "usage_known": true, + "wasted_actions": 5, + "steps": 8, + "cost_usd": null, + "submitted_values": [ + "", + "", + "", + "hello world" + ] + }, + { + "task": "normal", + "arm": "off", + "repeat": 1, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 85.35, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 7, + "decision_failures": 0, + "decision_ms_total": 56397, + "chat_calls": 3, + "chat_failures": 0, + "chat_ms_total": 6631, + "chat_input_tokens": 722, + "chat_output_tokens": 57, + "usage_known": true, + "wasted_actions": 5, + "steps": 7, + "cost_usd": null, + "submitted_values": [ + "", + "", + "" + ] + }, + { + "task": "recoverable", + "arm": "off", + "repeat": 0, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 127.243, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 6, + "decision_failures": 0, + "decision_ms_total": 59806, + "chat_calls": 1, + "chat_failures": 1, + "chat_ms_total": 45390, + "chat_input_tokens": 0, + "chat_output_tokens": 0, + "usage_known": false, + "wasted_actions": 5, + "steps": 6, + "cost_usd": null, + "submitted_values": [ + "original" + ] + }, + { + "task": "recoverable", + "arm": "on", + "repeat": 0, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 149.847, + "recovery_attempts": 2, + "recovery_spent_s": 5.732, + "recovery_termination": "planned", + "planner_attempts": 2, + "planner_failures": 0, + "decision_calls": 10, + "decision_failures": 0, + "decision_ms_total": 108385, + "chat_calls": 3, + "chat_failures": 0, + "chat_ms_total": 6982, + "chat_input_tokens": 2047, + "chat_output_tokens": 127, + "usage_known": true, + "wasted_actions": 8, + "steps": 9, + "cost_usd": null, + "submitted_values": [ + "original" + ] + }, + { + "task": "recoverable", + "arm": "on", + "repeat": 1, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 148.147, + "recovery_attempts": 2, + "recovery_spent_s": 5.198, + "recovery_termination": "planned", + "planner_attempts": 2, + "planner_failures": 0, + "decision_calls": 10, + "decision_failures": 0, + "decision_ms_total": 107864, + "chat_calls": 3, + "chat_failures": 0, + "chat_ms_total": 5836, + "chat_input_tokens": 2047, + "chat_output_tokens": 126, + "usage_known": true, + "wasted_actions": 8, + "steps": 9, + "cost_usd": null, + "submitted_values": [ + "original" + ] + }, + { + "task": "recoverable", + "arm": "off", + "repeat": 1, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 84.617, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 6, + "decision_failures": 0, + "decision_ms_total": 60242, + "chat_calls": 1, + "chat_failures": 0, + "chat_ms_total": 2298, + "chat_input_tokens": 140, + "chat_output_tokens": 50, + "usage_known": true, + "wasted_actions": 5, + "steps": 6, + "cost_usd": null, + "submitted_values": [ + "original" + ] + }, + { + "task": "permanently_blocked", + "arm": "off", + "repeat": 0, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 35.499, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 3, + "decision_failures": 0, + "decision_ms_total": 26413, + "chat_calls": 0, + "chat_failures": 0, + "chat_ms_total": 0, + "chat_input_tokens": 0, + "chat_output_tokens": 0, + "usage_known": true, + "wasted_actions": 2, + "steps": 2, + "cost_usd": null, + "submitted_values": [] + }, + { + "task": "permanently_blocked", + "arm": "on", + "repeat": 0, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 35.466, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 3, + "decision_failures": 0, + "decision_ms_total": 26414, + "chat_calls": 0, + "chat_failures": 0, + "chat_ms_total": 0, + "chat_input_tokens": 0, + "chat_output_tokens": 0, + "usage_known": true, + "wasted_actions": 2, + "steps": 2, + "cost_usd": null, + "submitted_values": [] + }, + { + "task": "permanently_blocked", + "arm": "on", + "repeat": 1, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 35.478, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 3, + "decision_failures": 0, + "decision_ms_total": 26395, + "chat_calls": 0, + "chat_failures": 0, + "chat_ms_total": 0, + "chat_input_tokens": 0, + "chat_output_tokens": 0, + "usage_known": true, + "wasted_actions": 2, + "steps": 2, + "cost_usd": null, + "submitted_values": [] + }, + { + "task": "permanently_blocked", + "arm": "off", + "repeat": 1, + "verified": false, + "terminal": "BLOCKED", + "errored": false, + "error": null, + "elapsed_s": 35.613, + "recovery_attempts": 0, + "recovery_spent_s": 0.0, + "recovery_termination": null, + "planner_attempts": 0, + "planner_failures": 0, + "decision_calls": 3, + "decision_failures": 0, + "decision_ms_total": 26514, + "chat_calls": 0, + "chat_failures": 0, + "chat_ms_total": 0, + "chat_input_tokens": 0, + "chat_output_tokens": 0, + "usage_known": true, + "wasted_actions": 2, + "steps": 2, + "cost_usd": null, + "submitted_values": [] + } + ], + "limitations": [ + "One decision model, three synthetic forms, two repeats per arm/task; no general benchmark claim.", + "The server-validation variant was planned after observing premature submissions in the original forms. It is exploratory, reported separately.", + "Chat calls include planner attempts; do not add these counts.", + "Unchanged-page actions are a proxy, not a complete measure of wasted effort.", + "Cost is unknown; no provider invoice was verified.", + "Earlier environment/pilot failures are separate setup records, not silently replaced trials in either 12-trial matrix." + ] +} diff --git a/evals/recovery/live.py b/evals/recovery/live.py new file mode 100644 index 0000000..95c96b8 --- /dev/null +++ b/evals/recovery/live.py @@ -0,0 +1,939 @@ +# coding: utf-8 +"""The live real-model recovery eval: one browser trial per (task, arm, repeat) through the production ``browse``. + +Unlike ``evals.recovery`` (scripted doubles), the decision model and the chat/planner model here are the real ones: +``build_model`` (``jev`` over HTTP, ``laya`` or ``cua`` in process) fills the model slot and ``chat_model_from_env`` +supplies the chat/planner/answer model. Nothing is injected: no scripted decisions, no standard plan and no hidden +answer reach the run, so a downstream ``DONE`` is only ever the model's own. The browser, the policy, the registered +``browser_*`` tools, the generation ids, the recovery budget and the terminal handling are all the production ones. + +Success is the fixture server's recorded form POST carrying the task's expected value, never the model's answer. Each +trial opens its own fresh fixture (unique trial id, empty submit namespace) and its own browser inside ``browse``, so +an earlier trial's submit can never verify a later one. Every planned trial is recorded, errors and timeouts included. + +Counting is deliberately split, and never conflated: + +- Decision calls are counted by a thin proxy that appends its record at the ``_decide`` entry and fills the wall time + and outcome in a ``finally``-safe step, so retries (``decide_many`` re-asking after an unusable answer), failures and + cancellations are counted for real and consistently. A failed call leaves its usage unknown, not zero. The proxy + preserves the inner model's ``name``, ``question_types``, ``deterministic``, ``supports_images``, ``model``, ``warm`` + and ``close``, so the production policy sees an identical model. Ticks are not used as a call count. +- Planner attempts are counted from the production recovery events (``stage == "planner"``), including planner calls + that timed out or failed. A planner call is a chat call too, so it is inside ``chat_calls``; the two are never summed. +- Chat call counts and tokens come from the production ``CountingModel`` that ``browse`` wraps around the chat model + (``answer["usage"]``); when a trial raises before that summary exists they are taken from the outer ``CountingModel``, + which records the same calls' wall time, tokens and failure status. + +A trial's ``verified`` (the fixture's recorded POST) and ``errored`` (the run raised) are independent: a trial that +posted the right value and then raised is both. Only a trial that failed before the browser existed (a fixture or +setup failure) falls back to a fabricated record. + +Cost is ``null`` by default: the live eval never reads a provider catalogue, and only an explicit ``CHAT_USD_PER_M_*`` +configuration values a trial as an estimate, never a bill. Even then the estimate stays ``null`` when any chat or +decision call did not report usage. (The production browser path's own ``usage_summary`` may consult OpenRouter's +catalogue for a price; the live eval deliberately ignores that and keeps its ``cost_usd`` explicit-config or ``null``.) +Model load time is measured once and listed separately from the per-trial numbers; a setup failure (a model that will +not build, a chat model with no key) exits non-zero without inventing any trial. +""" + +from __future__ import annotations + +import s1a.entry # noqa: F401 # routes the harness logs to files before anything imports openjiuwen + +import argparse +import asyncio +import json +import platform +import sys +import time +from dataclasses import asdict, dataclass, field +from datetime import datetime +from pathlib import Path +from typing import Any +from uuid import uuid4 + +from s1a.browser import browse, prompts +from s1a.browser.decision_model import BrowserPolicy +from s1a.config import HOME, chat_model_from_env, first_env +from s1a.counting_model import CountingModel +from s1a.decision_models import Decision, DecisionModel, Observation, Question, Reply, build_model +from s1a.jobs import RESULTS_DIR, now_iso +from s1a.pricing import ChatPrices, cost_usd, env_prices +from s1a.recovery import RecoveryLimits +from s1a.run import started_runner +from s1a.spec import BrowserAgentSpec, Budget + +from evals.recovery.fixture import DEFAULT_TASKS, Fixture, FormTask, start_fixture, validated_tasks + +GOAL_TEMPLATE = ( + "Open {url} and complete the form: type '{value}' into the field labelled Value and click Submit. " + "Stop once the form has been submitted." +) +LIVE_DIRNAME = "recovery_live" +ARMS = ("off", "on") +# A failed or cancelled planner/refresh call, as the production recovery event records it. +PLANNER_FAILURES = ("timeout", "error", "cancelled") +# The decision-model context variables worth recording: weights, subfolder, window and device. Never keys or URLs. +CONTEXT_ENV = ( + "LAYA_MODEL", + "LAYA_SUBFOLDER", + "LAYA_MAX_LEN", + "LAYA_HEAD_MAX_LEN", + "LAYA_DEVICE", + "CUA_S1_CHECKPOINT", + "CUA_S1_SUBFOLDER", + "CUA_S1_DEVICE", +) + + +def _context_env() -> dict[str, str]: + """The named decision-model context variables that are set; never the whole environment, never keys or URLs.""" + return {name: value for name in CONTEXT_ENV if (value := first_env(name))} + + +def _chat_identity(chat: Any) -> dict[str, str | None]: + """The chat model's safe identity: its ``model_config``'s model name and its client provider. No key, no base URL. + + The whole config objects are not serializable (the client config carries the api key), so only the safe fields are + read; a caller-supplied chat model is identified the same way. + """ + request = getattr(chat, "model_config", None) + client = getattr(chat, "model_client_config", None) + return { + "model_name": getattr(request, "model_name", None), + "provider": getattr(client, "client_provider", None), + } + + +@dataclass(frozen=True) +class LiveConfig: + """The budgets a live trial runs under; defaults keep one repeat cheap for development. + + ``validate_submission`` is off by default, so the original unconstrained fixtures and their reports are unchanged. + When on, ``run_live`` uses validated equivalents of the supplied tasks: a mismatched POST gets HTTP 422 and the + same retryable form, never a success page. This is a server-side variant, not an injected plan or action. + """ + + timeout_s: float = 180.0 + max_steps: int = 24 + stall_after: int = 3 + max_recovery_attempts: int = 3 + recovery_timeout_s: float = 15.0 + headless: bool = True + validate_submission: bool = False + + +class CountingDecisionModel(DecisionModel): + """A proxy that counts every ``_decide`` call, retries and failures included, and changes nothing else. + + ``decide_many`` is the production policy's entry point; the retry loop inside the base class calls ``_decide`` + again on an unusable answer, so counting at ``_decide`` counts the backend calls that actually happened. The + record is appended before the call runs and updated in a ``finally``-safe step, with wall time measured + consistently for every outcome, so a failed or cancelled call is never lost. A failed call's tokens are unknown, + not zero. The inner model's interface is preserved, so ``BrowserDecisionModel`` behaves exactly as with the raw + model. + """ + + def __init__(self, inner: DecisionModel) -> None: + self._inner = inner + self.name = inner.name + self.question_types = inner.question_types + self.deterministic = inner.deterministic + self.supports_images = inner.supports_images + self.calls: list[dict[str, Any]] = [] + self.decide_many_calls = 0 + + @property + def model(self) -> str: + return self._inner.model + + @property + def usage_known(self) -> bool: + """Whether every counted call reported its usage: one failed call makes the total unknown, not zero.""" + return all(bool(call.get("usage_known", True)) for call in self.calls) + + async def _decide(self, observation: Observation, questions: dict[str, Question]) -> Reply: + started = time.perf_counter() + record: dict[str, Any] = { + "ms": 0, + "status": "running", + "input_tokens": 0, + "output_tokens": 0, + "usage_known": False, # no reply yet: the tokens are unknown, not zero + } + self.calls.append(record) + try: + reply = await self._inner._decide(observation, questions) + except asyncio.CancelledError: + record["ms"] = round((time.perf_counter() - started) * 1000) + record["status"] = "cancelled" + raise + except BaseException as exc: # noqa: BLE001 - the failure is counted, then re-raised unchanged + record["ms"] = round((time.perf_counter() - started) * 1000) + record["status"] = "error" + record["error"] = type(exc).__name__ # the type only: a provider error can carry headers or keys + raise + record["ms"] = round((time.perf_counter() - started) * 1000) + record["status"] = "ok" + record["input_tokens"] = int(reply.usage.input_tokens) + record["output_tokens"] = int(reply.usage.output_tokens) + record["usage_known"] = True + return reply + + async def decide_many( + self, observation: Observation, questions: dict[str, Question], *, attempts: int = 1 + ) -> Decision: + self.decide_many_calls += 1 + return await super().decide_many(observation, questions, attempts=attempts) + + async def warm(self) -> None: + await self._inner.warm() + + async def close(self) -> None: + await self._inner.close() + + +@dataclass +class LiveTrial: + """One planned trial's outcome, whether it verified, failed, timed out or never reached the browser.""" + + task: str + arm: str + repeat: int + verified: bool + terminal: str + errored: bool + error: str | None + model: str + platform: str + elapsed_s: float + recovery_attempts: int = 0 + recovery_spent_s: float = 0.0 + recovery_failed: bool = False + recovery_termination: str | None = None + recovery_events: list[dict[str, Any]] = field(default_factory=list) + planner_attempts: int = 0 + planner_failures: int = 0 + decision_calls: int = 0 + decision_failures: int = 0 + decision_retries: int = 0 + decision_ms_total: int = 0 + decision_ms_mean: float | None = None + decision_input_tokens: int = 0 + decision_usage_known: bool = True + unknown_decision_calls: int = 0 + chat_calls: int = 0 + chat_calls_observed: int = 0 + chat_failures: int = 0 + chat_ms_total: int = 0 + chat_input_tokens: int = 0 + chat_output_tokens: int = 0 + chat_cache_tokens: int = 0 + usage_known: bool = False + unknown_chat_calls: int = 0 + jev_input_tokens: int = 0 + cost_usd: float | None = None + cost_basis: str = "" + wasted_actions: int = 0 + steps: int = 0 + oracle: dict[str, Any] = field(default_factory=dict) + + def as_json(self) -> dict[str, Any]: + return asdict(self) + + +@dataclass +class LiveRun: + """A finished run: its manifest, every trial, the summary, and where they were written.""" + + manifest: dict[str, Any] + trials: list[LiveTrial] + summary: dict[str, Any] + run_dir: Path + + +def _csv(value: str) -> list[str]: + return [part.strip() for part in value.split(",") if part.strip()] + + +def _planned_trials(repeat: int, tasks: tuple[FormTask, ...], arms: tuple[str, ...]) -> list[tuple[FormTask, str, int]]: + """Every (task, arm, repeat) the run plans, with the arm order alternating by repeat to blunt startup bias.""" + planned: list[tuple[FormTask, str, int]] = [] + for task in tasks: + for index in range(repeat): + order = list(arms) if index % 2 == 0 else list(reversed(arms)) + for arm in order: + planned.append((task, arm, index)) + return planned + + +def _validate_run(repeat: int, tasks: tuple[FormTask, ...], arms: tuple[str, ...]) -> None: + """Reject arguments that would otherwise fabricate a plausible-looking but meaningless report.""" + if repeat < 1: + raise ValueError(f"repeat must be >= 1, got {repeat}") + if not tasks: + raise ValueError("at least one fixture task is required") + if not arms: + raise ValueError("at least one arm is required") + if len(set(arms)) != len(arms): + raise ValueError(f"arms must be distinct, got {arms}") + unknown = [arm for arm in arms if arm not in ARMS] + if unknown: + raise ValueError(f"unknown arms {unknown}; expected any of {ARMS}") + names = [task.name for task in tasks] + if len(set(names)) != len(names): + raise ValueError(f"fixture task names must be distinct, got {names}") + + +def _cost( + *, + usage_known: bool, + decision_usage_known: bool, + jev_input_tokens: int, + chat_input_tokens: int, + chat_output_tokens: int, + chat_cache_tokens: int, + prices: ChatPrices | None, +) -> tuple[float | None, str]: + """A trial's cost: ``None`` unless the user explicitly configured prices, then an estimate, never a bill. + + The estimate stays ``None`` when any call's usage is incomplete: a failed chat or decision call spent tokens that + cannot be summed, and a partial sum must not be reported as if it were the whole bill. + """ + if prices is None: + return None, "unknown: live provider pricing is not verified and CHAT_USD_PER_M_* was not set" + if not usage_known: + return None, "unknown: at least one chat call did not report usage" + if not decision_usage_known: + return None, "unknown: at least one decision call did not report usage" + value = cost_usd(jev_input_tokens, chat_input_tokens, chat_output_tokens, chat_cache_tokens, prices) + if value is None: + return None, "unknown: chat tokens were spent but the configured price could not value them" + return value, "estimated from explicit CHAT_USD_PER_M_* config; not a provider bill" + + +def _recovery_events(recovery: dict[str, Any]) -> list[dict[str, Any]]: + """The compact recovery events kept in a trial: the trigger, the stage and how it ended, never the page.""" + return [ + { + "trigger": event.get("trigger"), + "stage": event.get("stage"), + "termination": event.get("termination"), + "tick": event.get("tick"), + "attempt": event.get("attempt"), + "spent_s": event.get("spent_s"), + "error": event.get("error"), + } + for event in (recovery.get("events") or []) + if isinstance(event, dict) + ] + + +async def play_trial( + fixture: Fixture, + task: FormTask, + arm: str, + repeat: int, + *, + config: LiveConfig, + logs_dir: Path, + tasks: tuple[FormTask, ...], + model_name: str, + chat: Any, + decision_model: DecisionModel, + prices: ChatPrices | None, +) -> LiveTrial: + """Play one live trial over an already-open fixture; the caller owns the per-trial fixture and browser.""" + wall_started = time.perf_counter() + proxy = CountingDecisionModel(decision_model) + outer_calls: list[dict[str, Any]] = [] + counting_chat = CountingModel(chat, outer_calls) + spec = BrowserAgentSpec( + name="recovery_live", + description=f"Live recovery eval fixture: {task.kind} form.", + rules=prompts.OPERATION_RULES["en"], + language="en", + budget=Budget(max_steps=config.max_steps, timeout_s=config.timeout_s, stall_after=config.stall_after), + goal=None, + ) + policy = BrowserPolicy( + prefetch_values=False, + batch_actions=False, + goal_value_cache=False, + rethink_on=arm == "on", + recovery_limits=RecoveryLimits(max_attempts=config.max_recovery_attempts, timeout_s=config.recovery_timeout_s), + ) + goal = GOAL_TEMPLATE.format(url=fixture.task_url(task), value=task.expected_value) + answer: dict[str, Any] = {} + error: str | None = None + try: + answer = await browse.browse( + spec, + policy, + model_name=model_name, + goal=goal, + timeout_s=config.timeout_s, + max_steps=config.max_steps, + logs_dir=logs_dir, + headless=config.headless, + chat=counting_chat, + decision_model=proxy, + ) + except asyncio.CancelledError: + raise + except Exception as exc: # noqa: BLE001 - the trial keeps its real counters and oracle; the type is the summary + error = type(exc).__name__ # the type only: a provider error can carry headers or keys in its message + elapsed_s = round(time.perf_counter() - wall_started, 3) + report = answer.get("report") or {} + recovery = report.get("recovery") or {} + usage = answer.get("usage") or {} + history = report.get("history") or [] + events = _recovery_events(recovery) + + calls = proxy.calls + decision_calls = len(calls) + decision_ms_total = sum(int(call.get("ms") or 0) for call in calls) + decision_input_tokens = sum(int(call.get("input_tokens") or 0) for call in calls) + unknown_decision_calls = sum(1 for call in calls if not call.get("usage_known", True)) + usage_known = bool(usage.get("usage_known")) + if usage: + chat_calls = int(usage.get("chat_calls") or 0) + chat_input_tokens = int(usage.get("chat_input_tokens") or 0) + chat_output_tokens = int(usage.get("chat_output_tokens") or 0) + chat_cache_tokens = int(usage.get("chat_cache_tokens") or 0) + unknown_chat_calls = int(usage.get("unknown_calls") or 0) + else: # the browser raised before producing its own usage summary: use the outer counter's real records + chat_calls = len(outer_calls) + chat_input_tokens = sum(int(call.get("input_tokens") or 0) for call in outer_calls) + chat_output_tokens = sum(int(call.get("output_tokens") or 0) for call in outer_calls) + chat_cache_tokens = sum(int(call.get("cache_tokens") or 0) for call in outer_calls) + usage_known = bool(outer_calls) and all(call.get("usage_known", True) for call in outer_calls) + unknown_chat_calls = sum(1 for call in outer_calls if not call.get("usage_known", True)) + # The report's jev ticks miss validation retries; the proxy saw every actual backend reply, retries included. + jev_input_tokens = max(int(usage.get("jev_input_tokens") or 0), decision_input_tokens if proxy.name == "jev" else 0) + cost, basis = _cost( + usage_known=usage_known, + decision_usage_known=proxy.usage_known, + jev_input_tokens=jev_input_tokens, + chat_input_tokens=chat_input_tokens, + chat_output_tokens=chat_output_tokens, + chat_cache_tokens=chat_cache_tokens, + prices=prices, + ) + verified = fixture.verified(task) + status = answer.get("status") + if verified: + terminal = "verified" + elif status is not None: + terminal = str(status) + else: + reason = str(answer.get("error") or error or "no terminal verdict") + terminal = "timeout" if "timeout" in reason.lower() else f"error: {reason}" + if error is None: + error = reason + return LiveTrial( + task=task.name, + arm=arm, + repeat=repeat, + verified=verified, + terminal=terminal, + errored=error is not None, + error=error, + model=decision_model.model, + platform=platform.platform(), + elapsed_s=elapsed_s, + recovery_attempts=int(recovery.get("attempts", 0) or 0), + recovery_spent_s=float(recovery.get("spent_s", 0.0) or 0.0), + recovery_failed=bool(recovery.get("failed")), + recovery_termination=recovery.get("termination"), + recovery_events=events, + planner_attempts=sum(1 for event in events if event.get("stage") == "planner"), + planner_failures=sum( + 1 for event in events if event.get("stage") == "planner" and event.get("termination") in PLANNER_FAILURES + ), + decision_calls=decision_calls, + decision_failures=sum(1 for call in calls if call.get("status") != "ok"), + decision_retries=max(0, decision_calls - proxy.decide_many_calls), + decision_ms_total=decision_ms_total, + decision_ms_mean=round(decision_ms_total / decision_calls, 1) if decision_calls else None, + decision_input_tokens=decision_input_tokens, + decision_usage_known=proxy.usage_known, + unknown_decision_calls=unknown_decision_calls, + chat_calls=chat_calls, + chat_calls_observed=len(outer_calls), + chat_failures=sum(1 for call in outer_calls if call.get("status") != "ok"), + chat_ms_total=sum(int(call.get("ms") or 0) for call in outer_calls), + chat_input_tokens=chat_input_tokens, + chat_output_tokens=chat_output_tokens, + chat_cache_tokens=chat_cache_tokens, + usage_known=usage_known, + unknown_chat_calls=unknown_chat_calls, + jev_input_tokens=jev_input_tokens, + cost_usd=cost, + cost_basis=basis, + wasted_actions=sum( + 1 for entry in history if entry.get("kind") != "wait" and entry.get("page_changed") is False + ), + steps=len([entry for entry in history if entry.get("kind") != "wait"]), + oracle=fixture.oracle_state(task), + ) + + +async def run_trial( + task: FormTask, + arm: str, + repeat: int, + *, + config: LiveConfig, + logs_dir: Path, + tasks: tuple[FormTask, ...], + model_name: str, + chat: Any, + decision_model: DecisionModel, + prices: ChatPrices | None, +) -> LiveTrial: + """One live trial with its own fresh fixture/oracle; needs a started Runner.""" + with start_fixture(tasks) as fixture: + return await play_trial( + fixture, + task, + arm, + repeat, + config=config, + logs_dir=logs_dir, + tasks=tasks, + model_name=model_name, + chat=chat, + decision_model=decision_model, + prices=prices, + ) + + +def _failed_trial( + task: FormTask, arm: str, repeat: int, exc: BaseException, *, model_id: str, elapsed_s: float +) -> LiveTrial: + """A trial that raised before any oracle existed: recorded with its real reason in the planned denominator.""" + reason = f"{type(exc).__name__}: {exc}" + terminal = "timeout" if isinstance(exc, asyncio.TimeoutError) else f"error: {reason}" + return LiveTrial( + task=task.name, + arm=arm, + repeat=repeat, + verified=False, + terminal=terminal, + errored=True, + error=reason, + model=model_id, + platform=platform.platform(), + elapsed_s=round(elapsed_s, 3), + cost_usd=None, + cost_basis="unknown: the trial never produced usage", + oracle={"task": task.name, "verified": False, "submissions": []}, + ) + + +def _rate(count: int, planned: int) -> float | None: + return round(count / planned, 3) if planned else None + + +def _counts(values: Any) -> dict[str, int]: + counts: dict[str, int] = {} + for value in values: + if value is None: + continue + counts[str(value)] = counts.get(str(value), 0) + 1 + return counts + + +def _mean(values: list[float], digits: int) -> float | None: + return round(sum(values) / len(values), digits) if values else None + + +def _arm_block(records: list[LiveTrial]) -> dict[str, Any]: + planned = len(records) + verified = sum(record.verified for record in records) + costs = [record.cost_usd for record in records] + decision_means = [record.decision_ms_mean for record in records if record.decision_ms_mean is not None] + if not costs: + cost_total: float | None = 0.0 + elif any(cost is None for cost in costs): + cost_total = None # one unknown trial makes the arm's cost unknown, not zero + else: + cost_total = round(sum(cost for cost in costs if cost is not None), 6) + return { + "planned": planned, + "verified": verified, + "errored": sum(record.errored for record in records), + "completion_rate": _rate(verified, planned), + "terminals": _counts(record.terminal for record in records), + "recovery_attempts": sum(record.recovery_attempts for record in records), + "recovery_failed": sum(record.recovery_failed for record in records), + "recovery_terminations": _counts(record.recovery_termination for record in records), + "planner_attempts": sum(record.planner_attempts for record in records), + "planner_failures": sum(record.planner_failures for record in records), + "decision_calls": sum(record.decision_calls for record in records), + "decision_failures": sum(record.decision_failures for record in records), + "decision_retries": sum(record.decision_retries for record in records), + "unknown_decision_calls": sum(record.unknown_decision_calls for record in records), + "decision_usage_known": all(record.decision_usage_known for record in records) if records else None, + "chat_calls": sum(record.chat_calls for record in records), + "chat_failures": sum(record.chat_failures for record in records), + "wasted_actions": sum(record.wasted_actions for record in records), + "mean_elapsed_s": _mean([record.elapsed_s for record in records], 3), + "mean_decision_ms": _mean(decision_means, 1), + "mean_recovery_attempts": _mean([float(record.recovery_attempts) for record in records], 2), + "chat_input_tokens": sum(record.chat_input_tokens for record in records), + "chat_output_tokens": sum(record.chat_output_tokens for record in records), + "chat_cache_tokens": sum(record.chat_cache_tokens for record in records), + "jev_input_tokens": sum(record.jev_input_tokens for record in records), + "usage_known": all(record.usage_known for record in records) if records else None, + "unknown_chat_calls": sum(record.unknown_chat_calls for record in records), + "cost_usd": cost_total, + } + + +def summarize(trials: list[LiveTrial], arms: tuple[str, ...], *, model_id: str, model_name: str) -> dict[str, Any]: + """Every planned trial counts. Per arm, per task, and the paired off/on outcomes per (task, repeat).""" + arms_block = {arm: _arm_block([record for record in trials if record.arm == arm]) for arm in arms} + tasks = sorted({record.task for record in trials}) + by_task = { + task: {arm: _arm_block([r for r in trials if r.arm == arm and r.task == task]) for arm in arms} + for task in tasks + } + pairs: dict[tuple[str, int], dict[str, LiveTrial]] = {} + for record in trials: + pairs.setdefault((record.task, record.repeat), {})[record.arm] = record + both = off_only = on_only = neither = incomplete = 0 + delta_decision_calls: list[int] = [] + delta_chat_calls: list[int] = [] + delta_elapsed_s: list[float] = [] + delta_wasted_actions: list[int] = [] + for arm_records in pairs.values(): + off, on = arm_records.get("off"), arm_records.get("on") + if off is None or on is None: + incomplete += 1 + continue + delta_decision_calls.append(on.decision_calls - off.decision_calls) + delta_chat_calls.append(on.chat_calls - off.chat_calls) + delta_elapsed_s.append(round(on.elapsed_s - off.elapsed_s, 3)) + delta_wasted_actions.append(on.wasted_actions - off.wasted_actions) + if off.verified and on.verified: + both += 1 + elif on.verified: + on_only += 1 + elif off.verified: + off_only += 1 + else: + neither += 1 + off = arms_block.get("off") or {} + on = arms_block.get("on") or {} + off_rate = (off.get("verified") or 0) / off["planned"] if off.get("planned") else None + on_rate = (on.get("verified") or 0) / on["planned"] if on.get("planned") else None + return { + "planned_trials": len(trials), + "model": {"requested": model_name, "resident": model_id}, + "arms": arms_block, + "by_task": by_task, + "paired": { + "pairs": len(pairs) - incomplete, + "incomplete_pairs": incomplete, + "both_verified": both, + "off_only_verified": off_only, + "on_only_verified": on_only, + "neither_verified": neither, + "mean_delta_decision_calls": _mean([float(v) for v in delta_decision_calls], 2), + "mean_delta_chat_calls": _mean([float(v) for v in delta_chat_calls], 2), + "mean_delta_elapsed_s": _mean(delta_elapsed_s, 3), + "mean_delta_wasted_actions": _mean([float(v) for v in delta_wasted_actions], 2), + }, + "on_minus_off_completion_rate": ( + round(on_rate - off_rate, 3) if on_rate is not None and off_rate is not None else None + ), + "note": ( + "real decision and chat models on three small synthetic form pages; success is the fixture's recorded " + "POST, not the model's DONE. This is not an open-task success rate and recovery may not trigger. " + "chat_calls/tokens are the production CountingModel's; planner attempts are the production recovery " + "events with stage == 'planner' and are also inside chat_calls (never summed). cost_usd is null unless " + "CHAT_USD_PER_M_* was explicitly configured and every chat and decision call reported usage; then it is " + "that config's estimate, not a provider bill." + ), + } + + +def render_markdown(summary: dict[str, Any]) -> str: + """A short human-readable companion to the JSON.""" + arms = list(summary["arms"]) + lines: list[str] = [ + "# Bounded recovery on/off - live real-model fixture subset", + "", + f"Model: {summary['model']['requested']} (resident id: {summary['model']['resident']}).", + f"Planned trials: {summary['planned_trials']} (errors and timeouts included, never dropped).", + "", + "| arm | planned | verified | completion | errored | recovery attempts | planner attempts | " + "decision calls | chat calls | mean elapsed s | wasted actions | cost usd |", + "|---|---|---|---|---|---|---|---|---|---|---|---|", + ] + for arm in arms: + block = summary["arms"][arm] + lines.append( + f"| {arm} | {block['planned']} | {block['verified']} | {block['completion_rate']} | {block['errored']} | " + f"{block['recovery_attempts']} | {block['planner_attempts']} | {block['decision_calls']} | " + f"{block['chat_calls']} | {block['mean_elapsed_s']} | {block['wasted_actions']} | {block['cost_usd']} |" + ) + lines += ["", "| task | arm | planned | verified | completion |", "|---|---|---|---|---|"] + for task, task_arms in summary["by_task"].items(): + for arm in arms: + block = task_arms[arm] + lines.append(f"| {task} | {arm} | {block['planned']} | {block['verified']} | {block['completion_rate']} |") + paired = summary["paired"] + lines += [ + "", + "Paired completion per (task, repeat): " + f"on-only {paired['on_only_verified']}, off-only {paired['off_only_verified']}, " + f"both {paired['both_verified']}, neither {paired['neither_verified']}.", + f"On minus off completion rate: {summary['on_minus_off_completion_rate']}.", + "Paired mean on-minus-off deltas: " + f"decision calls {paired['mean_delta_decision_calls']}, chat calls {paired['mean_delta_chat_calls']}, " + f"elapsed s {paired['mean_delta_elapsed_s']}, wasted actions {paired['mean_delta_wasted_actions']}.", + "", + f"Note: {summary['note']}", + "", + ] + return "\n".join(lines) + + +def _write_json(path: Path, data: Any) -> None: + """Write one JSON artifact, creating its parent: a finished trial survives a batch kill that stops the run.""" + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8") + + +async def run_live( + *, + model_name: str, + repeat: int = 1, + tasks: tuple[FormTask, ...] = DEFAULT_TASKS, + arms: tuple[str, ...] = ARMS, + config: LiveConfig | None = None, + results_dir: Path = RESULTS_DIR, + chat: Any = None, +) -> LiveRun: + """Build the real models once, run every planned trial, and write the manifest, trials and summary. + + The decision and chat models are loaded once and their load time is reported separately from the trials. The + Runner is held in one process; each trial opens its own fixture and ``browse`` tears down its own browser. A trial + that raises inside ``browse`` is recorded with its real counters and oracle in the planned denominator, never + dropped. A setup failure (no key, no weights, a half-set price) propagates before any trial, so the caller can + exit non-zero without a fake run. ``chat`` lets the caller supply the chat/planner model (for example a real model + with custom headers); when it is ``None`` the environment's chat model is built. + + A unique run directory is created and the planned manifest written before the first trial, and each finished trial + is written on its own, so a batch scheduler's kill does not erase completed evidence. There is no resume engine. + """ + _validate_run(repeat, tasks, arms) + config = config or LiveConfig() + if config.validate_submission: + # The supplied tasks are immutable and stay untouched; the run uses validated equivalents of them. + tasks = validated_tasks(tasks) + results_dir = Path(results_dir) + run_id = f"{datetime.now():%Y-%m-%d__%H-%M-%S}-{uuid4().hex[:6]}" + run_dir = results_dir / LIVE_DIRNAME / run_id + trials_dir = run_dir / "trials" + logs_root = HOME / "runs" / "recovery_live" / run_id + prices = env_prices() # explicit user config only; a half-set pair raises here, before any trial + planned = _planned_trials(repeat, tasks, arms) + + started = now_iso() + manifest: dict[str, Any] = { + "run_id": run_id, + "kind": "recovery_live", + "platform": platform.platform(), + "python": sys.version.split()[0], + "started_at": started, + "status": "planned", + "tasks": [task.name for task in tasks], + "arms": list(arms), + "repeat": repeat, + "planned_trials": len(planned), + "plan": [{"task": task.name, "arm": arm, "repeat": index} for task, arm, index in planned], + "config": asdict(config), + "results_dir": str(results_dir), + } + _write_json(run_dir / "manifest.json", manifest) # the run id exists and the plan is on disk before any model + + decision_started = time.perf_counter() + decision_model = build_model(model_name) + decision_load_ms = round((time.perf_counter() - decision_started) * 1000) + model_id = decision_model.model + trials: list[LiveTrial] = [] + try: + chat_started = time.perf_counter() + if chat is None: + chat = chat_model_from_env() + chat_load_ms = round((time.perf_counter() - chat_started) * 1000) + chat_load_note = "one-time load for the whole run; excluded from every per-trial number" + else: + chat_load_ms = 0 + chat_load_note = "chat model supplied by the caller; its load time is not measured here" + manifest.update( + { + "status": "running", + "model": {"requested": model_name, "resident": model_id}, + "chat": _chat_identity(chat), + "context": _context_env(), + "model_load_ms": { + "decision": decision_load_ms, + "chat": chat_load_ms, + "total": decision_load_ms + chat_load_ms, + "note": chat_load_note, + }, + } + ) + _write_json(run_dir / "manifest.json", manifest) + async with started_runner(): + for ordinal, (task, arm, index) in enumerate(planned): + logs_dir = logs_root / f"{task.name}__{arm}__r{index}" + trial_started = time.perf_counter() + try: + trial = await run_trial( + task, + arm, + index, + config=config, + logs_dir=logs_dir, + tasks=tasks, + model_name=model_name, + chat=chat, + decision_model=decision_model, + prices=prices, + ) + except Exception as exc: # noqa: BLE001 - a raised trial is a planned failure, not a drop + trial = _failed_trial( + task, + arm, + index, + exc, + model_id=model_id, + elapsed_s=time.perf_counter() - trial_started, + ) + trials.append(trial) + _write_json( + trials_dir / f"{ordinal:02d}__{task.name}__{arm}__r{index}.json", trial.as_json() + ) # a finished trial is on disk before the next one starts + finally: + await decision_model.close() + + summary = summarize(trials, arms, model_id=model_id, model_name=model_name) + manifest["status"] = "finished" + manifest["finished_at"] = now_iso() + manifest["completed_trials"] = len(trials) + manifest["cost"] = { + "default": "null unless CHAT_USD_PER_M_* is explicitly set", + "configured": prices is not None, + "basis": ( + "estimate from explicit CHAT_USD_PER_M_* config; not a provider bill" + if prices is not None + else "unknown: live provider pricing is not verified here" + ), + "note": ( + "the live eval never reads a provider catalogue; even with explicit prices a trial is null when any " + "call's usage is incomplete" + ), + } + manifest["note"] = summary["note"] + _write_json(run_dir / "manifest.json", manifest) + _write_json(run_dir / "trials.json", [trial.as_json() for trial in trials]) + _write_json(run_dir / "summary.json", summary) + (run_dir / "summary.md").write_text(render_markdown(summary), encoding="utf-8") + return LiveRun(manifest=manifest, trials=trials, summary=summary, run_dir=run_dir) + + +def format_run(run: LiveRun) -> str: + return "\n".join( + [ + render_markdown(run.summary), + "Outputs:", + f"- run dir: {run.run_dir}", + f"- manifest: {run.run_dir / 'manifest.json'}", + f"- trials: {run.run_dir / 'trials.json'}", + f"- per-trial: {run.run_dir / 'trials'}", + f"- summary: {run.run_dir / 'summary.json'}", + ] + ) + + +def parser() -> argparse.ArgumentParser: + build = argparse.ArgumentParser(prog="python -m evals.recovery.live", description=__doc__) + build.add_argument( + "--model", + choices=("jev", "laya", "cua"), + required=True, + help="the real decision model: jev over HTTP, laya or cua in process", + ) + build.add_argument("--repeat", type=int, default=1, help="repeats per (task, arm), paired; must be >= 1") + build.add_argument( + "--tasks", + default=",".join(task.name for task in DEFAULT_TASKS), + help="comma-separated fixture tasks (normal,recoverable,permanently_blocked)", + ) + build.add_argument("--arms", default=",".join(ARMS), help="comma-separated recovery arms (off,on)") + build.add_argument("--results-dir", type=Path, default=RESULTS_DIR, help="root of the live run folders") + build.add_argument("--timeout", type=float, default=LiveConfig.timeout_s, help="seconds per trial") + build.add_argument("--max-steps", type=int, default=LiveConfig.max_steps, help="the subagent's iteration cap") + build.add_argument("--stall-after", type=int, default=LiveConfig.stall_after, help="actions without a page change") + build.add_argument( + "--recovery-attempts", type=int, default=LiveConfig.max_recovery_attempts, help="bounded recovery attempts" + ) + build.add_argument( + "--recovery-timeout", + type=float, + default=LiveConfig.recovery_timeout_s, + help="bounded recovery active seconds per trial", + ) + build.add_argument("--headed", action="store_true", help="show the browser (default: headless)") + build.add_argument( + "--validate-submission", + action="store_true", + help="reject a mismatched form value with HTTP 422 and the same retryable form (default: off, unchanged)", + ) + return build + + +def main(argv: list[str] | None = None) -> int: + parser_ = parser() + args = parser_.parse_args(argv) + task_table = {task.name: task for task in DEFAULT_TASKS} + try: + tasks = tuple(task_table[name] for name in _csv(args.tasks)) + except KeyError as exc: + parser_.error(f"unknown task {exc}; expected any of {sorted(task_table)}") + arms = tuple(_csv(args.arms)) + config = LiveConfig( + timeout_s=args.timeout, + max_steps=args.max_steps, + stall_after=args.stall_after, + max_recovery_attempts=args.recovery_attempts, + recovery_timeout_s=args.recovery_timeout, + headless=not args.headed, + validate_submission=args.validate_submission, + ) + try: + run = asyncio.run( + run_live( + model_name=args.model, + repeat=args.repeat, + tasks=tasks, + arms=arms, + config=config, + results_dir=args.results_dir, + ) + ) + except Exception as exc: # noqa: BLE001 - a setup failure is reported as such and exits non-zero + print(f"live recovery eval setup failed: {type(exc).__name__}: {exc}", file=sys.stderr) + return 2 + print(format_run(run)) + return 1 if any(trial.errored for trial in run.trials) else 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/evals/recovery/runner.py b/evals/recovery/runner.py new file mode 100644 index 0000000..b3e7be9 --- /dev/null +++ b/evals/recovery/runner.py @@ -0,0 +1,382 @@ +# coding: utf-8 +"""The recovery eval runner: one real browser trial per (task, arm, repeat) through the repository's ``browse``. + +The browser, the model slot (``BrowserDecisionModel``), the registered ``browser_*`` tools and their generation ids, +the recovery budget and the terminal handling are all the production ones; ``create_browser_agent`` supplies the same +runtime. The decision model and the fallback chat model are the scripted doubles (``evals.recovery.scripted``); there +is no real-backend mode here. Batch actions and ``unsafe_dev`` stay off. The fixture server's recorded POST is the +independent oracle, and every trial gets its own fresh fixture (see ``run_trial``) so no earlier submit can verify it. + +Time is reported as separate numbers and never conflated with inference: ``elapsed_s`` is the whole trial wall clock +(it includes harness and browser startup), ``decisions_ms`` is the scripted decision time measured with +``perf_counter`` (near-zero by construction and not a benchmark). ``recorded_probe_ms`` sums recorded probe wall times, +already including action settling. It excludes recovery, final and unticked probes, so it is not total environment +time. Cost is 0 by construction because no paid API is called. +""" + +from __future__ import annotations + +import asyncio +import json +import platform +import time +from dataclasses import dataclass, field +from datetime import datetime +from pathlib import Path +from typing import Any + +from s1a.browser import browse, prompts +from s1a.browser.decision_model import BrowserPolicy +from s1a.config import HOME +from s1a.decision_models import DecisionModel +from s1a.jobs import RESULTS_DIR, Episode, now_iso, write_job +from s1a.recovery import RecoveryLimits +from s1a.run import started_runner +from s1a.spec import BrowserAgentSpec, Budget + +from evals.recovery.fixture import DEFAULT_TASKS, Fixture, FormTask, start_fixture +from evals.recovery.scripted import SCRIPTED_MODEL_ID, ScriptedChatModel, ScriptedRecoveryModel +from evals.recovery.summary import ARMS, TrialRecord, paired_summary, render_markdown + +PLAN_HINT = ( + "The field keeps reverting after typing. Click the control labelled 'Enable editing' first, then type the value " + "and submit the form." +) +GOAL_TEMPLATE = ( + "Open {url} and complete the form: type '{value}' into the field labelled Value and click Submit. " + "Stop once the form has been submitted." +) + + +@dataclass(frozen=True) +class EvalConfig: + """The budgets a trial runs under; defaults keep one repeat cheap for development.""" + + timeout_s: float = 120.0 + max_steps: int = 24 + stall_after: int = 3 + max_recovery_attempts: int = 3 + recovery_timeout_s: float = 15.0 + headless: bool = True + + +@dataclass +class EvalRun: + summary: dict[str, Any] + records: list[TrialRecord] = field(default_factory=list) + job_dirs: list[Path] = field(default_factory=list) + paired_dir: Path | None = None + + +def _scripted_models(task: FormTask) -> tuple[DecisionModel, Any]: + """The controlled doubles: a fault-injecting decision model and a planner/answer chat model, both offline.""" + return ScriptedRecoveryModel(task), ScriptedChatModel(value=task.expected_value, plan=PLAN_HINT) + + +def _trial_metrics(answer: dict[str, Any], decision_model: DecisionModel, chat: Any) -> dict[str, Any]: + report = answer.get("report") or {} + history = report.get("history") or [] + ticks = answer.get("ticks") or [] + decision_calls = int(getattr(decision_model, "decide_calls", 0) or 0) + planner_calls = int(getattr(chat, "planner_calls", 0) or 0) + chat_calls = int(getattr(chat, "invoke_calls", 0) or 0) - planner_calls + return { + "wasted_actions": sum( + 1 for entry in history if entry.get("kind") != "wait" and entry.get("page_changed") is False + ), + # probe_ms already includes action settling; adding the settle fields would double-count. + # This is a diagnostic over recorded ticks, not total environment time. + "recorded_probe_ms": sum(int(tick.get("probe_ms", 0) or 0) for tick in ticks), + "decisions_ms": sum(int(tick.get("decision_ms", 0) or 0) for tick in ticks), + "decision_calls": decision_calls, + "planner_calls": planner_calls, + "chat_calls": chat_calls, + "model_calls": decision_calls + chat_calls + planner_calls, + } + + +def _decisions(ticks: list[dict[str, Any]]) -> list[dict[str, Any]]: + return [ + { + "step": tick.get("tick"), + "key": tick.get("operation"), + "target": tick.get("target"), + "ms": tick.get("decision_ms", 0), + "source": SCRIPTED_MODEL_ID, + } + for tick in ticks + ] + + +async def run_trial( + task: FormTask, + arm: str, + repeat: int, + *, + config: EvalConfig, + logs_dir: Path, + tasks: tuple[FormTask, ...] = DEFAULT_TASKS, +) -> tuple[Episode, TrialRecord]: + """One real browser trial with its own fresh fixture/oracle; needs a started Runner. + + A new fixture server is opened for this trial alone, so its submit records start empty and no earlier trial's + success can carry over. Both the CLI and the system test call this function and nothing else. + """ + with start_fixture(tasks) as fixture: + return await play_trial(fixture, task, arm, repeat, config=config, logs_dir=logs_dir) + + +async def play_trial( + fixture: Fixture, + task: FormTask, + arm: str, + repeat: int, + *, + config: EvalConfig, + logs_dir: Path, +) -> tuple[Episode, TrialRecord]: + """Play one trial over an already-open fixture; ``run_trial`` owns the per-trial fixture, tests may pass one.""" + started_at = now_iso() + wall_started = time.perf_counter() + decision_model, chat = _scripted_models(task) + spec = BrowserAgentSpec( + name="recovery_eval", + description=f"Recovery eval fixture: {task.kind} form.", + rules=prompts.OPERATION_RULES["en"], + language="en", + budget=Budget(max_steps=config.max_steps, timeout_s=config.timeout_s, stall_after=config.stall_after), + goal=None, + ) + policy = BrowserPolicy( + prefetch_values=False, + batch_actions=False, + goal_value_cache=False, + rethink_on=arm == "on", + recovery_limits=RecoveryLimits(max_attempts=config.max_recovery_attempts, timeout_s=config.recovery_timeout_s), + ) + goal = GOAL_TEMPLATE.format(url=fixture.task_url(task), value=task.expected_value) + # The scripted model is handed to browse's decision-model slot; that branch is keyed "jev" but no Jev client + # exists here, so no API is called. + answer = await browse.browse( + spec, + policy, + model_name="jev", + goal=goal, + timeout_s=config.timeout_s, + max_steps=config.max_steps, + logs_dir=logs_dir, + headless=config.headless, + chat=chat, + decision_model=decision_model, + ) + elapsed_s = round(time.perf_counter() - wall_started, 3) + report = answer.get("report") or {} + terminal = answer.get("terminal") or {} + recovery = report.get("recovery") or {} + ticks = answer.get("ticks") or [] + verified = fixture.verified(task) + status = answer.get("status") + if verified: + terminal_name = "verified" + error = None + elif status is not None: + terminal_name = str(status) # a played episode that did not submit (e.g. BLOCKED): score 0, not an exception + error = None + else: + reason = str(answer.get("error") or "no terminal verdict") # the harness could not play the episode at all + terminal_name = "timeout" if "timeout" in reason.lower() else f"error: {reason}" + error = reason + metrics = _trial_metrics(answer, decision_model, chat) + record = TrialRecord( + task=task.name, + arm=arm, + repeat=repeat, + verified=verified, + terminal=terminal_name, + errored=error is not None, + scripted_model=SCRIPTED_MODEL_ID, + platform=platform.platform(), + recovery_attempts=int(recovery.get("attempts", 0) or 0), + recovery_spent_s=float(recovery.get("spent_s", 0.0) or 0.0), + recovery_failed=bool(recovery.get("failed")), + recovery_termination=recovery.get("termination"), + wasted_actions=metrics["wasted_actions"], + model_calls=metrics["model_calls"], + decision_calls=metrics["decision_calls"], + planner_calls=metrics["planner_calls"], + chat_calls=metrics["chat_calls"], + elapsed_s=elapsed_s, + decisions_ms=metrics["decisions_ms"], + recorded_probe_ms=metrics["recorded_probe_ms"], + cost_usd=0.0, + oracle=fixture.oracle_state(task), + ) + episode = Episode( + env="recovery", + policy=f"recovery_{arm}", + seed=repeat, + score=1.0 if verified else 0.0, + steps=len([entry for entry in (report.get("history") or []) if entry.get("kind") != "wait"]), + elapsed_s=elapsed_s, + started_at=started_at, + finished_at=now_iso(), + final_state={ + "task": task.name, + "arm": arm, + "verified": verified, + "terminal": terminal, + "recovery": recovery, + "oracle": fixture.oracle_state(task), + }, + chat_calls=int(getattr(chat, "invoke_calls", 0)), + chat_input_tokens=0, + chat_output_tokens=0, + chat_cache_tokens=0, + jev_input_tokens=int(report.get("jev_input_tokens", 0) or 0), + invalid_keys=0, + cost_usd=0.0, + error=error, + decisions=_decisions(ticks), + extra=record.as_json(), + ) + return episode, record + + +def _failed_episode( + task: FormTask, arm: str, repeat: int, exc: BaseException, elapsed_s: float +) -> tuple[Episode, TrialRecord]: + """A trial that raised: recorded with its real reason in the planned denominator, never dropped and never 'ok'.""" + stamp = now_iso() + reason = f"{type(exc).__name__}: {exc}" + terminal = "timeout" if isinstance(exc, (TimeoutError, asyncio.TimeoutError)) else f"error: {reason}" + record = TrialRecord( + task=task.name, + arm=arm, + repeat=repeat, + verified=False, + terminal=terminal, + errored=True, + scripted_model=SCRIPTED_MODEL_ID, + platform=platform.platform(), + elapsed_s=round(elapsed_s, 3), + cost_usd=0.0, + oracle={"task": task.name, "verified": False, "submissions": []}, + ) + episode = Episode( + env="recovery", + policy=f"recovery_{arm}", + seed=repeat, + score=0.0, + steps=0, + elapsed_s=round(elapsed_s, 3), + started_at=stamp, + finished_at=stamp, + final_state={"task": task.name, "arm": arm, "verified": False, "failed": reason}, + chat_calls=0, + chat_input_tokens=0, + chat_output_tokens=0, + chat_cache_tokens=0, + jev_input_tokens=0, + invalid_keys=0, + cost_usd=0.0, + error=reason, + extra=record.as_json(), + ) + return episode, record + + +def _validate_run(repeat: int, tasks: tuple[FormTask, ...], arms: tuple[str, ...]) -> None: + """Reject arguments that would otherwise fabricate a plausible-looking but meaningless report.""" + if repeat < 1: + raise ValueError(f"repeat must be >= 1, got {repeat}") + if not tasks: + raise ValueError("at least one fixture task is required") + if not arms: + raise ValueError("at least one arm is required") + if len(set(arms)) != len(arms): + raise ValueError(f"arms must be distinct, got {arms}") + unknown = [arm for arm in arms if arm not in ARMS] + if unknown: + raise ValueError(f"unknown arms {unknown}; expected any of {ARMS}") + names = [task.name for task in tasks] + if len(set(names)) != len(names): + raise ValueError(f"fixture task names must be distinct, got {names}") + + +async def run_eval( + *, + repeat: int = 1, + tasks: tuple[FormTask, ...] = DEFAULT_TASKS, + config: EvalConfig | None = None, + results_dir: Path = RESULTS_DIR, + arms: tuple[str, ...] = ARMS, +) -> EvalRun: + """Run every (task, repeat, arm) once, write the Harbor jobs and the paired summary, and return them. + + The Runner is held in one process; each trial opens its own fixture and tears down its own browser inside + ``browse``. The arm order alternates with the repeat so the first arm run in a pair is not always the same + startup; a trial that raises is recorded with its real reason in the planned denominator, never dropped. + """ + _validate_run(repeat, tasks, arms) + config = config or EvalConfig() + run_id = f"{datetime.now():%Y-%m-%d__%H-%M-%S}" + logs_root = HOME / "runs" / "recovery" / run_id + episodes_by_arm: dict[str, list[Episode]] = {arm: [] for arm in arms} + records: list[TrialRecord] = [] + async with started_runner(): + for task in tasks: + for index in range(repeat): + order = list(arms) if index % 2 == 0 else list(reversed(arms)) + for arm in order: + logs_dir = logs_root / f"{task.name}__{arm}__r{index}" + started = time.perf_counter() + try: + episode, record = await run_trial(task, arm, index, config=config, logs_dir=logs_dir) + except Exception as exc: # noqa: BLE001 - a trial that raised is a planned failure, not a drop + episode, record = _failed_episode(task, arm, index, exc, time.perf_counter() - started) + episodes_by_arm[arm].append(episode) + records.append(record) + job_dirs: list[Path] = [] + for arm in arms: + if episodes_by_arm[arm]: + job_dirs.append(write_job("recovery", episodes_by_arm[arm], results_dir=results_dir)) + summary = paired_summary(records) + paired_dir = results_dir / "recovery_paired" / run_id + paired_dir.mkdir(parents=True, exist_ok=True) + (paired_dir / "paired_summary.json").write_text(_json_summary(summary, records), encoding="utf-8") + (paired_dir / "paired_summary.md").write_text(render_markdown(summary), encoding="utf-8") + return EvalRun(summary=summary, records=records, job_dirs=job_dirs, paired_dir=paired_dir) + + +def _json_summary(summary: dict[str, Any], records: list[TrialRecord]) -> str: + return json.dumps( + { + "summary": summary, + "trials": [record.as_json() for record in records], + "definition": { + "planned_trials": "every (task, repeat, arm) the run planned; errors and timeouts included", + "verified": "this trial's own fixture server saw a real submit carrying the task's expected value", + "oracle": "per-trial fixture namespace; a fresh empty fixture is opened for every trial", + "wasted_actions": "policy history entries with kind != wait and page_changed == False", + "elapsed_s": "whole trial wall clock, includes harness and browser startup (not inference)", + "decisions_ms": "scripted decision time measured with perf_counter, summed over ticks (near-zero, not a benchmark)", + "recorded_probe_ms": "sum of recorded probe wall times, including action settling; excludes recovery, final and unticked probes; not total environment time", + "model_calls": "decision_calls + chat_calls + planner_calls (all scripted doubles)", + "cost_usd": "0 by construction: scripted decisions and planner, no paid API call", + }, + }, + ensure_ascii=False, + indent=2, + ) + + +def format_run(run: EvalRun) -> str: + lines = [render_markdown(run.summary), "Outputs:"] + lines += [f"- job: {path}" for path in run.job_dirs] + if run.paired_dir is not None: + lines.append(f"- paired summary: {run.paired_dir}") + return "\n".join(lines) + + +def run_sync(**kwargs: Any) -> EvalRun: + return asyncio.run(run_eval(**kwargs)) diff --git a/evals/recovery/scripted.py b/evals/recovery/scripted.py new file mode 100644 index 0000000..26d58fc --- /dev/null +++ b/evals/recovery/scripted.py @@ -0,0 +1,172 @@ +# coding: utf-8 +"""Clearly-marked scripted doubles for the recovery eval: a faulted decision model and a planner/answer model. + +These are **not** a trained decision model and **not** a real benchmark. They are controlled doubles that only +exercise the recovery wiring: + +- ``ScriptedRecoveryModel`` is a deterministic state machine that reads the real observation and picks only offered + candidates. On the ``recoverable`` and ``blocked`` tasks it *deliberately keeps typing* (the injected fault) until + the production ``draft_plan`` puts a plan in the observation; only then may it choose the page's unlock control. + The same instance logic runs for the recovery on and off arms: with recovery off no plan ever appears, so it never + leaves the fault. It can never complete a task by itself -- every non-terminal turn returns one real ``browser_*`` + tool call, and completion is judged by the fixture server, never by the model. +- ``ScriptedChatModel`` answers only the fallback calls the browser policy makes: the replan prompt (``draft_plan``), + value generation, goal-value extraction and the final answer. It makes no network call. +""" + +from __future__ import annotations + +import json +import time +from typing import Any, AsyncIterator + +from openjiuwen.core.foundation.llm import AssistantMessage, AssistantMessageChunk, Model, ModelClientConfig + +from s1a.browser import prompts +from s1a.decision_models import DecisionModel, Observation, Question, Reply, Usage +from s1a.tool.rethink import REPLAN_PROMPT + +from evals.recovery.fixture import FormTask + +SCRIPTED_MODEL_ID = "scripted-fault-v1" +SCRIPTED_NOTE = "scripted double: controlled fault injection, not a trained model and not a real API call" + +_UNUSED_CLIENT = ModelClientConfig( + client_provider="OpenAI", + api_key="scripted-not-a-secret", + api_base="http://127.0.0.1:1/scripted", + verify_ssl=False, +) + + +def _one_hot(key: str, ids: list[str]) -> dict[str, Any]: + """A valid choice answer over ``ids``: the chosen key peaks, the distribution sums to one.""" + chosen = key if key in ids else ids[0] + return { + "choice": chosen, + "probabilities": {candidate: (1.0 if candidate == chosen else 0.0) for candidate in ids}, + "confidence": 1.0, + } + + +class ScriptedRecoveryModel(DecisionModel): + """The faulted form-filler: type until a plan arrives, then act on the offered unlock control, then submit.""" + + name = "scripted" + supports_images = False + deterministic = True + + def __init__(self, task: FormTask) -> None: + self._task = task + self._unlocked = False # set only after a production plan offered the page's unlock control + self.decide_calls = 0 + self.operations: list[str] = [] + + @property + def model(self) -> str: + return SCRIPTED_MODEL_ID + + async def _decide(self, observation: Observation, questions: dict[str, Question]) -> Reply: + started = time.perf_counter() + self.decide_calls += 1 + state = observation.state if isinstance(observation.state, dict) else {} + plan = str(state.get("plan") or "") + elements = [row for row in (state.get("elements") or []) if isinstance(row, dict)] + operation, target = self._choose(plan, elements) + self.operations.append(operation) + latency_ms = int(round((time.perf_counter() - started) * 1000)) # measured, not a fabricated constant + return Reply( + answers=self._answers(operation, target, questions), + latency_ms=latency_ms, + usage=Usage(), + model=SCRIPTED_MODEL_ID, + ) + + # -- the faulted strategy ------------------------------------------------- + + @staticmethod + def _row(elements: list[dict[str, Any]], *, operation: str, label: str | None = None) -> dict[str, Any] | None: + for row in elements: + if operation not in (row.get("operations") or []): + continue + if label is not None and label not in str(row.get("label") or "").lower(): + continue + return row + return None + + def _choose(self, plan: str, elements: list[dict[str, Any]]) -> tuple[str, str | None]: + """The operation and target index for this turn; only offered candidates, DONE only when the page shows it.""" + input_row = self._row(elements, operation="TYPE_TEXT") + submit = self._row(elements, operation="CLICK", label="submit") + enable = self._row(elements, operation="CLICK", label="enable") + expected = self._task.expected_value + if self._task.kind == "normal": + if input_row is not None and str(input_row.get("value") or "") != expected: + return "TYPE_TEXT", str(input_row["index"]) + if submit is not None: + return "CLICK", str(submit["index"]) + return "DONE", None + # recoverable/blocked: the fault keeps typing until the production plan names the unlock. + if plan and enable is not None and not self._unlocked: + self._unlocked = True + return "CLICK", str(enable["index"]) + if plan or self._unlocked: + if input_row is not None and str(input_row.get("value") or "") != expected: + return "TYPE_TEXT", str(input_row["index"]) + if submit is not None: + return "CLICK", str(submit["index"]) + return "DONE", None + if input_row is not None: # the injected fault: keep filling the field the page restores + return "TYPE_TEXT", str(input_row["index"]) + return "DONE", None + + @staticmethod + def _answers(operation: str, target: str | None, questions: dict[str, Question]) -> dict[str, Any]: + """Answer every asked head with an offered key; the chosen operation's target head carries the pick.""" + answers: dict[str, Any] = {} + for name, question in questions.items(): + ids = list(question.ids) + if name == "operation": + key = operation + elif target is not None and name == f"{operation.lower()}_target": + key = target + else: + key = ids[0] + answers[name] = _one_hot(key, ids) + return answers + + +class ScriptedChatModel(Model): + """The fallback chat model as a scripted double: replan, values and the final answer, with no network call.""" + + def __init__(self, *, value: str, plan: str, answer: str = "done") -> None: + super().__init__(_UNUSED_CLIENT, None) + self._value = value + self._plan = plan + self._answer = answer + self.invoke_calls = 0 + self.planner_calls = 0 # the replan prompt's share of invoke_calls, the rest is chat fallback + + async def invoke(self, messages: Any, *, tools: Any = None, **kwargs: Any) -> AssistantMessage: + self.invoke_calls += 1 + return AssistantMessage(content=self._reply(messages), finish_reason="stop") + + async def stream(self, messages: Any, *, tools: Any = None, **kwargs: Any) -> AsyncIterator[AssistantMessageChunk]: + self.invoke_calls += 1 + yield AssistantMessageChunk(content=self._reply(messages), finish_reason="stop") + + def _reply(self, messages: Any) -> str: + system = "" + if messages: + first = messages[0] + system = first.get("content", "") if isinstance(first, dict) else str(getattr(first, "content", "")) + if system == REPLAN_PROMPT: + self.planner_calls += 1 + return self._plan + if system in prompts.VALUE_GENERATION.values(): + return json.dumps({"text": self._value}) + if system in prompts.VALUE_EXTRACTION.values(): + return json.dumps({"values": [self._value]}) + if system in prompts.ANSWER_RULES.values(): + return self._answer + return self._answer diff --git a/evals/recovery/summary.py b/evals/recovery/summary.py new file mode 100644 index 0000000..1a357e4 --- /dev/null +++ b/evals/recovery/summary.py @@ -0,0 +1,176 @@ +# coding: utf-8 +"""The trial record and the paired on/off summary for the recovery eval. + +``jobs.summarize`` scores only the episodes that did not error, so its ``mean_score`` is not a completion rate over +every planned trial. ``paired_summary`` is the authoritative denominator: it counts every planned trial, errors and +timeouts included, and pairs the off and on arms per task and repeat. It also reports the paired mean on-minus-off +deltas for model calls, elapsed seconds and wasted actions. +""" + +from __future__ import annotations + +import statistics +from dataclasses import asdict, dataclass, field +from typing import Any + +ARMS = ("off", "on") + + +@dataclass +class TrialRecord: + """One planned trial's outcome, whether it verified, failed, timed out or was never supported.""" + + task: str + arm: str + repeat: int + verified: bool + terminal: str + errored: bool + scripted_model: str + platform: str + recovery_attempts: int = 0 + recovery_spent_s: float = 0.0 + recovery_failed: bool = False + recovery_termination: str | None = None + wasted_actions: int = 0 + model_calls: int = 0 + decision_calls: int = 0 + planner_calls: int = 0 + chat_calls: int = 0 + elapsed_s: float = 0.0 + decisions_ms: int = 0 + recorded_probe_ms: int = 0 + cost_usd: float = 0.0 + oracle: dict[str, Any] = field(default_factory=dict) + + def as_json(self) -> dict[str, Any]: + return asdict(self) + + +def _rate(count: int, planned: int) -> float | None: + return round(count / planned, 3) if planned else None + + +def _arm_block(records: list[TrialRecord]) -> dict[str, Any]: + planned = len(records) + verified = sum(record.verified for record in records) + costs = [record.cost_usd for record in records] + return { + "planned": planned, + "verified": verified, + "errored": sum(record.errored for record in records), + "completion_rate": _rate(verified, planned), + "recovery_attempts": sum(record.recovery_attempts for record in records), + "recovery_spent_s": round(sum(record.recovery_spent_s for record in records), 3), + "recovery_failed": sum(record.recovery_failed for record in records), + "wasted_actions": sum(record.wasted_actions for record in records), + "model_calls": sum(record.model_calls for record in records), + "decision_calls": sum(record.decision_calls for record in records), + "planner_calls": sum(record.planner_calls for record in records), + "chat_calls": sum(record.chat_calls for record in records), + "mean_model_calls": round(statistics.mean(r.model_calls for r in records), 2) if records else None, + "mean_elapsed_s": round(statistics.mean(r.elapsed_s for r in records), 3) if records else None, + "mean_wasted_actions": round(statistics.mean(r.wasted_actions for r in records), 2) if records else None, + "decisions_ms": sum(record.decisions_ms for record in records), + "recorded_probe_ms": sum(record.recorded_probe_ms for record in records), + "cost_usd": round(sum(costs), 6) if costs else 0.0, + } + + +def paired_summary(records: list[TrialRecord]) -> dict[str, Any]: + """Every planned trial counts. Per arm, per task, and the paired off/on outcomes per (task, repeat).""" + arms = {arm: _arm_block([r for r in records if r.arm == arm]) for arm in ARMS} + tasks = sorted({record.task for record in records}) + by_task = { + task: {arm: _arm_block([r for r in records if r.arm == arm and r.task == task]) for arm in ARMS} + for task in tasks + } + pairs: dict[tuple[str, int], dict[str, TrialRecord]] = {} + for record in records: + pairs.setdefault((record.task, record.repeat), {})[record.arm] = record + both = off_only = on_only = neither = incomplete = 0 + delta_model_calls: list[int] = [] + delta_elapsed_s: list[float] = [] + delta_wasted_actions: list[int] = [] + for arm_records in pairs.values(): + off, on = arm_records.get("off"), arm_records.get("on") + if off is None or on is None: + incomplete += 1 + continue + delta_model_calls.append(on.model_calls - off.model_calls) + delta_elapsed_s.append(round(on.elapsed_s - off.elapsed_s, 3)) + delta_wasted_actions.append(on.wasted_actions - off.wasted_actions) + if off.verified and on.verified: + both += 1 + elif on.verified: + on_only += 1 + elif off.verified: + off_only += 1 + else: + neither += 1 + off_rate = arms["off"]["verified"] / arms["off"]["planned"] if arms["off"]["planned"] else None + on_rate = arms["on"]["verified"] / arms["on"]["planned"] if arms["on"]["planned"] else None + return { + "planned_trials": len(records), + "arms": arms, + "by_task": by_task, + "paired": { + "pairs": len(pairs) - incomplete, + "incomplete_pairs": incomplete, + "both_verified": both, + "off_only_verified": off_only, + "on_only_verified": on_only, + "neither_verified": neither, + "mean_delta_model_calls": round(statistics.mean(delta_model_calls), 2) if delta_model_calls else None, + "mean_delta_elapsed_s": round(statistics.mean(delta_elapsed_s), 3) if delta_elapsed_s else None, + "mean_delta_wasted_actions": ( + round(statistics.mean(delta_wasted_actions), 2) if delta_wasted_actions else None + ), + }, + "on_minus_off_completion_rate": ( + round(on_rate - off_rate, 3) if on_rate is not None and off_rate is not None else None + ), + "note": ( + "scripted doubles: controlled fault injection to exercise recovery, not a trained model; " + "cost is 0 by construction, not a real API bill; starts/timeouts/failures count in every denominator." + ), + } + + +def render_markdown(summary: dict[str, Any]) -> str: + """A short human-readable companion to the JSON.""" + lines: list[str] = [ + "# Bounded recovery on/off - scripted fixture subset", + "", + f"Planned trials: {summary['planned_trials']} (errors and timeouts included, never dropped).", + "", + "| arm | planned | verified | completion | errored | recovery attempts | mean model calls | mean elapsed s | wasted actions |", + "|---|---|---|---|---|---|---|---|---|", + ] + for arm in ARMS: + block = summary["arms"][arm] + lines.append( + f"| {arm} | {block['planned']} | {block['verified']} | {block['completion_rate']} | {block['errored']} | " + f"{block['recovery_attempts']} | {block['mean_model_calls']} | {block['mean_elapsed_s']} | " + f"{block['wasted_actions']} |" + ) + lines += ["", "| task | arm | planned | verified | completion |", "|---|---|---|---|---|"] + for task, arms in summary["by_task"].items(): + for arm in ARMS: + block = arms[arm] + lines.append(f"| {task} | {arm} | {block['planned']} | {block['verified']} | {block['completion_rate']} |") + paired = summary["paired"] + lines += [ + "", + "Paired completion per (task, repeat): " + f"on-only {paired['on_only_verified']}, off-only {paired['off_only_verified']}, " + f"both {paired['both_verified']}, neither {paired['neither_verified']}.", + f"On minus off completion rate: {summary['on_minus_off_completion_rate']}.", + "Paired mean on-minus-off deltas: " + f"model calls {paired['mean_delta_model_calls']}, elapsed s {paired['mean_delta_elapsed_s']}, " + f"wasted actions {paired['mean_delta_wasted_actions']}.", + "", + f"Note: {summary['note']}", + "", + ] + return "\n".join(lines) diff --git a/s1a/agents/desktop.py b/s1a/agents/desktop.py index 1727a03..546b8b8 100644 --- a/s1a/agents/desktop.py +++ b/s1a/agents/desktop.py @@ -24,7 +24,7 @@ from s1a.desktop.driver import CuaDriver, Snapshot, driver_from_env, opened from s1a.desktop.env import ABSTAIN, DONE, WindowEnv, clickable -from s1a.spec import Budget, Series, ToolAgentSpec +from s1a.spec import Budget, Series, ToolAgentSpec, positive_float, positive_int RULES = ( "A desktop app window. goal says what to do; elements lists the window's controls with their labels and values; " @@ -103,13 +103,25 @@ def flags(parser: argparse.ArgumentParser) -> None: parser.add_argument("--execute", action="store_true", help="click for real; without it one decision is planned") parser.add_argument("--plan", default="", help="the rule baseline: button labels in order, | between variants") parser.add_argument("--clear", default="", help="button labels pressed on reset when the window has one") + parser.add_argument( + "--rethink-attempts", + type=positive_int, + default=3, + help="bounded rethink: stalls handled by a refresh and a plan before the episode gives up", + ) + parser.add_argument( + "--rethink-timeout", + type=positive_float, + default=15.0, + help="bounded rethink: seconds across all refreshes and plans in one episode", + ) SPEC = ToolAgentSpec( name="desktop", description="A Windows or macOS app window through Cua Driver: click controls toward --goal until --expect appears.", rules=RULES, - budget=Budget(max_steps=12, timeout_s=90, stall_after=0), + budget=Budget(max_steps=12, timeout_s=90, stall_after=3), flags=flags, series=make_series, ) diff --git a/s1a/browser/browse.py b/s1a/browser/browse.py index b9be304..84160c5 100644 --- a/s1a/browser/browse.py +++ b/s1a/browser/browse.py @@ -28,8 +28,9 @@ from s1a.browser.decision_model import URL_RE, BrowserDecisionModel, BrowserPolicy from s1a.decision_models import DecisionModel, build_model from s1a.config import HOME, browser_launch_args, chat_model_from_env, first_env -from s1a.counting_model import CountingModel +from s1a.counting_model import CountingModel, usage_known from s1a.pricing import chat_prices, cost_usd +from s1a.recovery import RecoveryLimits from s1a.browser.profiler import BrowserProfiler, render from s1a.spec import BrowserAgentSpec, positive_float, positive_int @@ -75,14 +76,33 @@ def finish_llm(answer: Answer) -> Answer: def finish_decision_model(answer: Answer, *, model_name: str) -> Answer: """The policy's answer from the page it reached (on DONE, or on BLOCKED after progress) is the result; none fails. - ``model_name`` names the decision model in the error.""" + + A bounded recovery that timed out, errored or ran out of attempts is a failure even when the summary carries an + answer: the run did not settle the task, it stopped trying. A policy that still answers BLOCKED after one or more + replans is a failure too, even with a partial answer: the task remains blocked, so the partial text is kept only + as context, with the real reason and an actionable next step. ``model_name`` names the decision model in the error. + """ summary = terminal_summary(answer["final"]) if summary is None: return answer status = str(summary.get("status") or "") + recovery = summary.get("recovery") or {} answer["status"] = status answer["terminal"] = summary answer["final"] = str(summary.get("answer") or "") + if recovery.get("failed"): + answer["ok"] = False + detail = recovery.get("error") or summary.get("reason") or "no reason given" + answer["error"] = f"{model_name} recovery {recovery.get('termination') or 'failed'}: {detail}" + # The actionable escalation rides with the failure; it never widens the tool set or retries on its own. + answer["next_action"] = recovery.get("next_action") or summary.get("next_action") + return answer + if recovery.get("blocked_after_recovery"): + answer["ok"] = False + detail = recovery.get("reason") or summary.get("reason") or "no reason given" + answer["error"] = f"{model_name} BLOCKED after recovery: {detail}" + answer["next_action"] = recovery.get("next_action") or summary.get("next_action") + return answer answer["ok"] = bool(answer["final"]) if answer["ok"]: answer["error"] = None # run_task flagged the harness's own verdict; the policy's answer is the one that counts @@ -94,10 +114,16 @@ def finish_decision_model(answer: Answer, *, model_name: str) -> Answer: def usage_summary(calls: list[dict[str, Any]], *, jev_input_tokens: int, decisions: int) -> dict[str, Any]: - """Counts and dollars for one task: decisions, the chat model's calls and tokens, Jev's input tokens, the sum in USD.""" + """Counts and dollars for one task: decisions, the chat model's calls and tokens, Jev's input tokens, the sum in USD. + + A failed or cancelled call is in ``calls`` with unknown usage: its tokens cannot be summed, so the task's + ``cost_usd`` is None (unknown), never zero, and ``usage_known`` is False with the count in ``unknown_calls``. + """ chat_in = sum(int(call["input_tokens"]) for call in calls) chat_out = sum(int(call["output_tokens"]) for call in calls) chat_cached = sum(int(call.get("cache_tokens") or 0) for call in calls) + complete = usage_known(calls) + unknown_calls = len([call for call in calls if not call.get("usage_known", True)]) prices = chat_prices(first_env("MODEL_NAME")) if chat_in + chat_out else None return { "decisions": decisions, @@ -106,7 +132,9 @@ def usage_summary(calls: list[dict[str, Any]], *, jev_input_tokens: int, decisio "chat_output_tokens": chat_out, "chat_cache_tokens": chat_cached, "jev_input_tokens": jev_input_tokens, - "cost_usd": cost_usd(jev_input_tokens, chat_in, chat_out, chat_cached, prices), + "usage_known": complete, + "unknown_calls": unknown_calls, + "cost_usd": None if not complete else cost_usd(jev_input_tokens, chat_in, chat_out, chat_cached, prices), } @@ -156,8 +184,13 @@ async def browse( A decision model writes ``decision_ticks.json`` under ``logs_dir`` and returns the ticks and the policy's report with the answer; ``llm`` writes ``chat_calls.json``. ``decision_model`` is required by every other name and unused - by ``llm``. + by ``llm``. ``policy.rethink_on`` needs a decision model: the plain chat model cannot read a plan, so the pair is + rejected here, before any agent or browser is built, instead of silently ignoring the flag. """ + if policy.rethink_on and model_name == "llm": + raise RuntimeError("--rethink on needs a decision model; --model llm cannot use it") + if policy.rethink_on and decision_model is None: + raise RuntimeError(f"--rethink on needs a decision model; --model {model_name} was given none") logs_dir.mkdir(parents=True, exist_ok=True) calls: list[dict[str, Any]] = [] counted = CountingModel(chat, calls) @@ -257,6 +290,27 @@ def parser(spec: BrowserAgentSpec) -> argparse.ArgumentParser: default="off", help="decision models only: offer values extracted from the goal as a choice head", ) + build.add_argument( + "--rethink", + choices=("on", "off"), + default="off", + help=( + "decision models only: on a stall re-probe the page and ask the chat model for a plan under a bounded " + "budget (off keeps the legacy BLOCKED guard; run the same goal both ways to compare)" + ), + ) + build.add_argument( + "--rethink-attempts", + type=positive_int, + default=3, + help="bounded rethink: stalls handled by a refresh and a plan before the task gives up", + ) + build.add_argument( + "--rethink-timeout", + type=positive_float, + default=15.0, + help="bounded rethink: seconds across all refreshes and plans in one task, never reset by progress", + ) build.add_argument( "--logs-dir", type=Path, @@ -273,10 +327,13 @@ def parser(spec: BrowserAgentSpec) -> argparse.ArgumentParser: def policy_from_args(args: argparse.Namespace) -> BrowserPolicy: + """The policy from the parsed flags. ``RecoveryLimits`` rejects a non-finite or non-positive budget here.""" return BrowserPolicy( prefetch_values=args.prefetch == "on", batch_actions=args.batch == "on", goal_value_cache=args.goal_values == "on", + rethink_on=args.rethink == "on", + recovery_limits=RecoveryLimits(max_attempts=args.rethink_attempts, timeout_s=float(args.rethink_timeout)), ) diff --git a/s1a/browser/decision_model.py b/s1a/browser/decision_model.py index 001478b..7c3ea57 100644 --- a/s1a/browser/decision_model.py +++ b/s1a/browser/decision_model.py @@ -19,6 +19,7 @@ import re import statistics import time +from collections.abc import Awaitable from dataclasses import dataclass, field from typing import Any, AsyncIterator from urllib.parse import urlparse @@ -42,7 +43,9 @@ ) from s1a.browser.probe_js import POLICY_PROBE_JS, STAMP_ATTRIBUTE from s1a.decision_models import DecisionModel +from s1a.recovery import RecoveryBudget, RecoveryExhausted, RecoveryLimits, recovery_next_action from s1a.spec import BrowserAgentSpec +from s1a.tool.rethink import draft_plan BROWSER_TURN_TOOL = "browser_click" # The runtime registers MCP tools under a server prefix (mcp__browser_click); match by suffix. @@ -75,6 +78,8 @@ ) MAX_URL_VISITS = 3 # the fourth arrival at one URL ends the run: the policy is circling MAX_PREFETCHED_FIELDS = 8 +OBSERVATION_ERROR_CHARS = 200 # a failed probe is refused with a short reason, never a whole browser log +RECENT_ACTIONS = 12 # the actions and candidates a bounded recovery's plan is written over PROBE_SETTLE_MS = 500 # must equal probe_js.py's JS default so an un-escalated probe's timing is unchanged PROBE_QUIET_MS = 60 # must equal probe_js.py's JS default DOM-quiet window ACTION_SETTLE_START_MS = 250 # first in-page wait for an action whose effect has not shown yet @@ -86,11 +91,18 @@ @dataclass(frozen=True) class BrowserPolicy: - """The run-time switches of the Jev browser policy, set per run by flags or the environment, never by the spec.""" + """The run-time switches of the Jev browser policy, set per run by flags or the environment, never by the spec. + + ``rethink_on`` turns the bounded recovery on: each stall (an unchanged page, an A-B-A-B loop, a visited page, an + over-long WAIT) spends one refresh and, when the page really did not move, one plan under ``recovery_limits``. + The fields are last so a three-argument ``BrowserPolicy`` keeps working. + """ prefetch_values: bool # generate a typed value for every editable field as soon as a probe shows it batch_actions: bool # each action and the next probe in one browser_run_code_unsafe call; needs unsafe_dev goal_value_cache: bool # offer values extracted from the goal to Jev as a choice head + rethink_on: bool = False # bounded recovery: refresh plus plan under a per-task budget, never a wider tool set + recovery_limits: RecoveryLimits = field(default_factory=RecoveryLimits) # attempts and active seconds, never reset @dataclass @@ -117,6 +129,20 @@ class _Run: url_visits: dict[str, int] = field(default_factory=dict) last_url: str = "" # the URL of the latest good probe; a probe landing elsewhere counts a visit to its URL token: str = field(default_factory=lambda: uuid4().hex[:6]) # keeps tool-call ids unique across runs + # -- bounded recovery: one per-run budget, its events, and the one-turn plan they may produce -------------- + recovery: RecoveryBudget | None = None + recovery_events: list[dict[str, Any]] = field(default_factory=list) + plan: str = "" # the last plan, attached to the next observation and cleared there; it never executes a click + history_boundary: int = 0 # detection reads history from here, so a plan gets a few actions before a re-stall + recovery_stopped: bool = False # the run ended because recovery failed: no answer is fetched for it + recovery_error: str | None = None + recovery_snapshot: dict[str, Any] | None = None # the freshest observation a stopped recovery reached + # The policy still answered BLOCKED after a replan: the task did not settle, but recovery itself did not fail. + recovery_blocked: bool = False + recovery_blocked_reason: str | None = None + terminal_message: AssistantMessage | None = ( + None # re-served to a same-task re-call, so a terminal run never revives + ) class BrowserDecisionModel(Model): @@ -146,8 +172,12 @@ def __init__( self._goal_value_cache = policy.goal_value_cache self._prefetch_enabled = policy.prefetch_values self._batch_actions = policy.batch_actions + self._rethink_on = policy.rethink_on + self._recovery_limits = policy.recovery_limits self._language = spec.language self._decision_model = decision_model + # the chat model writes the bounded recovery's plan, as the tool front's RethinkRail does + self._planner: Model | None = fallback self._runtime: Any = None self._tool_names: dict[str, str] = {} self._run: _Run | None = None @@ -217,10 +247,19 @@ async def _decide_message(self, messages: Any) -> AssistantMessage: ) goal = self._goal_from(messages) run = self._run + if run is not None and run.recovery_stopped and goal == run.goal: + if run.terminal_message is None: + # Cancellation can unwind before _final. A repeated call still cannot resume this task. + return await self._final( + run, "BLOCKED", run.recovery_snapshot or {}, run.recovery_error or "recovery stopped" + ) + return run.terminal_message if run is None or run.finished or goal != run.goal: if run is not None: self._cancel_run_tasks(run) run = _Run(goal=goal, started_at=time.perf_counter()) + if self._rethink_on: + run.recovery = RecoveryBudget(self._recovery_limits) self._run = run if self._goal_value_cache: run.values_task = asyncio.create_task(self._extract_values(goal)) @@ -234,21 +273,18 @@ async def _decide_message(self, messages: Any) -> AssistantMessage: snapshot, probe_ms = await (self._batched_probe(run, batched) if batched else self._probe(run)) settle_probes, settle_ms = 0, 0 while True: # a settle probe can set the guards too, so they are read on every pass - if snapshot.get("stalled"): - reason = f"{self._spec.budget.stall_after} actions without page change" - return await self._final(run, "BLOCKED", snapshot, reason) - if snapshot.get("oscillating"): - return await self._final(run, "BLOCKED", snapshot, "oscillating between two pages") - if snapshot.get("revisiting"): - return await self._final(run, "BLOCKED", snapshot, "reached the same page four times") + self._require_observation(snapshot) + guard = self._guard(snapshot) + if guard is not None: + trigger, reason = guard + recovered = await self._recover(run, snapshot, trigger=trigger, reason=reason) + if recovered is None: + return await self._final(run, "BLOCKED", snapshot, reason) + snapshot, probe_ms = recovered + settle_probes, settle_ms = 0, 0 + continue space = build_action_space(snapshot, run.history) - for operation, tool in _OPERATION_TOOLS.items(): - if tool not in self._tool_names: - space.operations = [op for op in space.operations if op != operation] - space.heads.pop(operation, None) - acted = next((h for h in reversed(run.history) if h["kind"] != "wait"), None) - if acted is not None and acted["kind"] == "scroll" and acted["page_changed"] is False: - space.operations = [op for op in space.operations if op != acted["action"]] # the page end was reached + self._filter_operations(space, run) values = run.values if run.values_task is None or run.values_task.done() else [] if run.values_task is not None and run.values_task.done() and not run.values: try: @@ -258,6 +294,10 @@ async def _decide_message(self, messages: Any) -> AssistantMessage: run.values = [] values = run.values observation = build_observation(space, snapshot, run.history) + if run.plan: + if isinstance(observation.state, dict): + observation.state["plan"] = run.plan # rides one turn; the decision still picks the action + run.plan = "" questions = build_questions( space, goal=run.goal, values=values, rules=self._spec.rules, language=self._language ) @@ -292,12 +332,26 @@ async def _decide_message(self, messages: Any) -> AssistantMessage: run.consecutive_waits += 1 run.history.append({"action": "wait", "kind": "wait", "text": None, "page_changed": None}) if run.consecutive_waits > MAX_CONSECUTIVE_WAITS: - return await self._final(run, "BLOCKED", snapshot, "waited without progress") + recovered = await self._recover( + run, snapshot, trigger="wait_streak", reason="waited without progress" + ) + if recovered is None: + return await self._final(run, "BLOCKED", snapshot, "waited without progress") + snapshot, probe_ms = recovered + settle_probes, settle_ms = 0, 0 + continue snapshot, settle_probes, settle_ms, progressed = await self._settle_wait(run, snapshot) if not progressed: record["settle_probes"] += settle_probes record["settle_ms"] += settle_ms - return await self._final(run, "BLOCKED", snapshot, "waited without progress") + recovered = await self._recover( + run, snapshot, trigger="wait_settle", reason="waited without progress" + ) + if recovered is None: + return await self._final(run, "BLOCKED", snapshot, "waited without progress") + snapshot, probe_ms = recovered + settle_probes, settle_ms = 0, 0 + continue probe_ms = settle_ms continue run.consecutive_waits = 0 @@ -309,6 +363,210 @@ async def _decide_message(self, messages: Any) -> AssistantMessage: return await self._final(run, "BLOCKED", snapshot, "") return await self._act(run, move, snapshot, record) + # -- bounded recovery --------------------------------------------------- + + @staticmethod + def _require_observation(snapshot: dict[str, Any]) -> None: + """Refuse a failed probe before the decision model or a chat answer can read it as an empty page. + + The runtime reports a failed probe as ``ok=False`` or an ``error`` field; either is a hard observation + failure (a missing browser library, a closed page), not a page with zero elements. Raising here keeps a + broken browser from being answered DONE, and the bounded reason keeps a full browser log out of the logs. + The recovery refresh and the in-page settle probes keep their own handling; only the observation fed to a + decision is refused. + """ + if _probe_failed(snapshot): + raise RuntimeError(f"browser observation failed: {_probe_error_reason(snapshot)}") + + def _guard(self, snapshot: dict[str, Any]) -> tuple[str, str] | None: + """The page-level stop signals as ``(trigger, reason)``; the caller recovers or ends the run.""" + if snapshot.get("stalled"): + return "stalled", f"{self._spec.budget.stall_after} actions without page change" + if snapshot.get("oscillating"): + return "oscillating", "oscillating between two pages" + if snapshot.get("revisiting"): + return "revisiting", "reached the same page four times" + return None + + def _filter_operations(self, space: Any, run: _Run) -> None: + """Drop the operations whose tool this turn was not offered: the offer is the run's, not the plan's.""" + for operation, tool in _OPERATION_TOOLS.items(): + if tool not in self._tool_names: + space.operations = [op for op in space.operations if op != operation] + space.heads.pop(operation, None) + acted = next((h for h in reversed(run.history) if h["kind"] != "wait"), None) + if acted is not None and acted["kind"] == "scroll" and acted["page_changed"] is False: + space.operations = [op for op in space.operations if op != acted["action"]] + + async def _recover( + self, run: _Run, snapshot: dict[str, Any], *, trigger: str, reason: str + ) -> tuple[dict[str, Any], int] | None: + """One stall's bounded attempt: a read-only refresh, then a plan for the next turn; ``None`` means stop. + + The event is appended and the attempt charged before the awaits, so a timeout, a cancellation or an error + still leaves the trigger, the fresh observation and the stage in the record. The fresh page is read through + the read-only probe, never a navigation; the plan only ever rides into the next observation, and the offered + tools and candidates are filtered exactly as the normal turn filters them, so recovery cannot widen what the + run may do. Only the detection windows reset after an attempt: history, ticks and the global budget stay. + """ + budget = run.recovery + if budget is None: + return None + event: dict[str, Any] = { + "kind": "recovery", + "trigger": trigger, + "reason": reason, + "tick": run.tick, + "attempt": budget.attempts, + "recent_actions": [ + {key: entry.get(key) for key in ("action", "kind", "page_changed")} + for entry in run.history[-RECENT_ACTIONS:] + ], + "stage": "refresh", + "fresh_obs": None, + "plan": "", + "spent_s": round(budget.spent_s, 3), + "termination": None, + } + run.recovery_events.append(event) + try: + budget.begin_attempt() # a stall starts one attempt; a spent budget gives up, never a success + except RecoveryExhausted as exc: + self._stop_recovery(run, event, stage="exhausted", termination="give_up", error=str(exc)) + return None + event["attempt"] = budget.attempts + try: + fresh = await budget.call(self._raw_probe(settle_ms=PROBE_SETTLE_MS, quiet_ms=PROBE_QUIET_MS, after=None)) + except asyncio.CancelledError: + self._stop_recovery(run, event, stage="refresh", termination="cancelled", error="recovery cancelled") + raise + except asyncio.TimeoutError: + self._stop_recovery(run, event, stage="refresh", termination="timeout", error="recovery refresh timed out") + return None + except Exception as exc: # noqa: BLE001 - any refresh failure is a clean terminal, never a retry loop + self._stop_recovery( + run, event, stage="refresh", termination="error", error=f"recovery refresh failed: {exc}" + ) + return None + finally: + event["spent_s"] = round(budget.spent_s, 3) + event["fresh_obs"] = fresh + run.recovery_snapshot = fresh + if _probe_failed(fresh): + self._stop_recovery( + run, + event, + stage="refresh", + termination="error", + error=f"recovery probe failed: {_probe_error_reason(fresh)}", + ) + return None + self._reset_detection(run, fresh) + if snapshot.get("page_key") != fresh.get("page_key"): + event["termination"] = "delayed_progress" # the page moved after all: no plan to write + return fresh, 0 + if self._planner is None: + event["termination"] = "no_planner" + return fresh, 0 + event["stage"] = "planner" + try: + plan = await budget.call(self._draft_plan(run, fresh)) + except asyncio.CancelledError: + self._stop_recovery(run, event, stage="planner", termination="cancelled", error="recovery cancelled") + raise + except asyncio.TimeoutError: + self._stop_recovery(run, event, stage="planner", termination="timeout", error="recovery planner timed out") + return None + except Exception as exc: # noqa: BLE001 - any planner failure is a clean terminal, never a retry loop + self._stop_recovery( + run, event, stage="planner", termination="error", error=f"recovery planner failed: {exc}" + ) + return None + finally: + event["spent_s"] = round(budget.spent_s, 3) + run.plan = plan # the next observation offers it; the plan never executes a click itself + event.update(stage="planner", plan=plan, termination="planned") + return fresh, 0 + + @staticmethod + def _reset_detection(run: _Run, fresh: dict[str, Any]) -> None: + """After an attempt only the detection windows reset; history, ticks and the global budget stay.""" + run.history_boundary = len(run.history) + run.consecutive_waits = 0 + run.settle_spent_ms = 0 + run.url_visits.clear() + run.last_url = str(fresh.get("url") or "") + + @staticmethod + def _stop_recovery(run: _Run, event: dict[str, Any], *, stage: str, termination: str, error: str) -> None: + if run.recovery is not None: + event["spent_s"] = round(run.recovery.spent_s, 3) + event.update(stage=stage, termination=termination, error=error) + event["next_action"] = recovery_next_action(termination=termination, stage=stage, error=error) + run.recovery_stopped = True + run.recovery_error = error + + def _draft_plan(self, run: _Run, fresh: dict[str, Any]) -> Awaitable[str]: + """The replan request over the fresh page, offering exactly the tools and candidates the turn will.""" + space = build_action_space(fresh, run.history) + self._filter_operations(space, run) + recent = [ + {"action": entry.get("action"), "kind": entry.get("kind"), "page_changed": entry.get("page_changed")} + for entry in run.history[-RECENT_ACTIONS:] + ] + state = { + "goal": run.goal, + "page": {"url": fresh.get("url", ""), "title": fresh.get("title", ""), "text": fresh.get("text", "")}, + "elements": space.elements, + } + candidates = { + "operations": list(space.operations), + "targets": { + operation: {key: candidate.item.get("label", "") for key, candidate in head.items()} + for operation, head in space.heads.items() + }, + } + assert self._planner is not None + return draft_plan( + self._planner, rules=self._spec.rules, recent_steps=recent, state=state, candidates=candidates + ) + + def _recovery_summary(self, run: _Run) -> dict[str, Any]: + """The recovery record for the report and the terminal summary: counts, active time and the last termination. + + On a failure it also carries the specific reason and one short ``next_action``: the escalation the task says + an operator should take, never a wider tool set or an automatic retry. A model that still answers BLOCKED + after a replan is recorded separately (``blocked_after_recovery``): the task did not settle, so it is not a + success, but recovery itself neither failed nor exhausted its budget. + """ + budget = run.recovery + last = run.recovery_events[-1] if run.recovery_events else {} + failed = run.recovery_stopped + blocked = run.recovery_blocked + if failed: + reason = run.recovery_error + next_action = recovery_next_action( + termination=last.get("termination"), stage=last.get("stage"), error=run.recovery_error + ) + elif blocked: + reason = run.recovery_blocked_reason + next_action = recovery_next_action(termination="blocked", error=run.recovery_blocked_reason) + else: + reason = next_action = None + return { + "attempts": budget.attempts if budget is not None else 0, + "spent_s": round(budget.spent_s, 3) if budget is not None else 0.0, + "exhausted": budget.exhausted if budget is not None else False, + "failed": failed, + "blocked_after_recovery": blocked, + "termination": last.get("termination"), + "stage": last.get("stage"), + "error": run.recovery_error, + "reason": reason, + "next_action": next_action, + "events": run.recovery_events, + } + async def _settle_wait(self, run: _Run, asked_snapshot: dict[str, Any]) -> tuple[dict[str, Any], int, int, bool]: """Spend a WAIT verdict in-page instead of paying another decisions request. @@ -372,14 +630,14 @@ async def _finish_probe( snapshot, changed = await self._settle_action(run, snapshot, pending["page_key"]) pending["entry"]["page_changed"] = changed limit = self._spec.budget.stall_after - recent = run.history[-limit:] if limit else [] + recent = run.history[run.history_boundary :][-limit:] if limit else [] if ( limit and len(recent) == limit and all(h["page_changed"] is False and h["kind"] != "wait" for h in recent) ): snapshot["stalled"] = True - acted = [h for h in run.history if h["kind"] != "wait"][-4:] + acted = [h for h in run.history[run.history_boundary :] if h["kind"] != "wait"][-4:] labels = [h["action"] for h in acted] if ( len(labels) == 4 @@ -717,12 +975,33 @@ def _tool_message(self, run: _Run, name: str, args: dict[str, Any], *, label: st async def _final(self, run: _Run, status: str, snapshot: dict[str, Any], reason: str) -> AssistantMessage: self._cancel_run_tasks(run) # no value call outlives its run - if status == "BLOCKED" and not run.answer and any(h.get("page_changed") for h in run.history): + if run.recovery_stopped and run.recovery_snapshot is not None: + snapshot = run.recovery_snapshot # the freshest page a stopped recovery reached + if ( + status == "BLOCKED" + and not run.answer + and not run.recovery_stopped # a failed recovery is a clear BLOCKED, not another answer request + and not _probe_failed(snapshot) # a failed observation holds no page a chat answer may read + and any(h.get("page_changed") for h in run.history) + ): run.answer = await self._answer(run, snapshot) # the page the run reached may already hold the answer run.finished = True + if status == "BLOCKED" and not run.recovery_stopped and run.recovery is not None and run.recovery.attempts > 0: + # The policy answered BLOCKED after a replan: the task is not settled, but this is not a failed or + # exhausted recovery either. Keep the partial answer as context and name the real reason and next step. + run.recovery_blocked = True + run.recovery_blocked_reason = ( + reason or f"model reported BLOCKED after {run.recovery.attempts} recovery attempt(s)" + ) summary = { "status": status, - "reason": reason, + "reason": ( + run.recovery_error + if run.recovery_stopped + else run.recovery_blocked_reason + if run.recovery_blocked + else reason + ), "url": snapshot.get("url"), "title": snapshot.get("title"), "steps": len([h for h in run.history if h["kind"] != "wait"]), @@ -730,7 +1009,15 @@ async def _final(self, run: _Run, status: str, snapshot: dict[str, Any], reason: "page_text": str(snapshot.get("text", ""))[:SUMMARY_TEXT_CHARS], "answer": run.answer, } - return AssistantMessage(content=json.dumps(summary, ensure_ascii=False), finish_reason="stop") + if run.recovery is not None: + # Full observations belong in report()/decision_ticks.json; the harness caps terminal text. + recovery = self._recovery_summary(run) + summary["recovery"] = {key: value for key, value in recovery.items() if key != "events"} + if recovery["next_action"] and (recovery["failed"] or recovery["blocked_after_recovery"]): + summary["next_action"] = recovery["next_action"] # the actionable escalation, next to the reason + message = AssistantMessage(content=json.dumps(summary, ensure_ascii=False), finish_reason="stop") + run.terminal_message = message + return message @staticmethod def _goal_from(messages: Any) -> str: @@ -765,6 +1052,19 @@ def report(self) -> dict[str, Any]: "settle_ms": 0, "values": {"cache": 0, "prefetch": 0, "llm": 0}, "history": [], + "recovery": { + "attempts": 0, + "spent_s": 0.0, + "exhausted": False, + "failed": False, + "blocked_after_recovery": False, + "termination": None, + "stage": None, + "error": None, + "reason": None, + "next_action": None, + "events": [], + }, } jev = [t["decision_ms"] for t in run.ticks] return { @@ -785,12 +1085,24 @@ def report(self) -> dict[str, Any]: for source in ("cache", "prefetch", "llm") }, "history": run.history, + "recovery": self._recovery_summary(run), } +def _probe_failed(snapshot: dict[str, Any]) -> bool: + """A probe that failed is not a page: the runtime reports it as ``ok=False`` or an ``error`` field.""" + return snapshot.get("ok") is False or bool(snapshot.get("error")) + + +def _probe_error_reason(snapshot: dict[str, Any]) -> str: + """The bounded, single-line reason for a failed probe; a browser log never floods the terminal or a log.""" + reason = " ".join(str(snapshot.get("error") or "policy probe returned no result").split()) + return reason[:OBSERVATION_ERROR_CHARS] + + def _progressed(snapshot: dict[str, Any], before_key: Any) -> bool: """A page moved on only when a probe that succeeded shows a different ``page_key``.""" - return not snapshot.get("error") and snapshot.get("page_key") != before_key + return not _probe_failed(snapshot) and snapshot.get("page_key") != before_key def _json_field(content: Any, key: str) -> Any: diff --git a/s1a/counting_model.py b/s1a/counting_model.py index 819f6e1..eb38449 100644 --- a/s1a/counting_model.py +++ b/s1a/counting_model.py @@ -1,30 +1,98 @@ # coding: utf-8 -"""``CountingModel``: a chat-model wrapper that records every call's latency, tokens and tool calls.""" +"""``CountingModel``: a chat-model wrapper that records every call's latency, tokens and tool calls. + +The record is appended before the call runs and updated when it ends, so a call that failed, timed out or was +cancelled is counted too, not lost. The record's ``error`` summary keeps only the error's type -- never its message, +the prompt or a secret -- and the compact ``call_digests`` keep no raw tool arguments. The full in-memory record does +retain ``tool_args`` for internal routing, kept for compatibility with earlier consumers. A call only marks its usage +known when it ends normally with ``usage_metadata``; a reply without it, or a stream that errored, was cancelled or +interrupted after a partial usage, leaves the cost unknown (``usage_known`` False), never a confirmed zero, so a +consumer can tell "no tokens" from "we do not know how many tokens". The caller's exception is always re-raised. +""" from __future__ import annotations +import asyncio import time from typing import Any, AsyncIterator from openjiuwen.core.foundation.llm import AssistantMessage, AssistantMessageChunk, Model +RECORD_KEYS = ("ms", "status", "error", "usage_known", "input_tokens", "output_tokens", "cache_tokens", "tool_calls") + -def call_record(reply: AssistantMessage, started: float) -> dict[str, Any]: - usage = reply.usage_metadata - calls = list(reply.tool_calls or []) +def start_record(started: float) -> dict[str, Any]: + """The record appended before the call runs: a failed or cancelled call is still visible in ``calls``.""" return { - "ms": round((time.perf_counter() - started) * 1000), - "input_tokens": usage.input_tokens if usage is not None else 0, - "output_tokens": usage.output_tokens if usage is not None else 0, - "cache_tokens": int(usage.cache_tokens or 0) if usage is not None else 0, - "tool_calls": [str(call.name) for call in calls], - "tool_args": [str(call.arguments) for call in calls], - "content_chars": len(reply.content) if isinstance(reply.content, str) else 0, + "ms": 0, + "status": "running", + "input_tokens": 0, + "output_tokens": 0, + "cache_tokens": 0, + "usage_known": False, # no reply yet: the tokens are unknown, not zero + "tool_calls": [], + "tool_args": [], + "content_chars": 0, } +def finish_record( + record: dict[str, Any], reply: Any, started: float, status: str, *, error: BaseException | None = None +) -> None: + """Update the record in place when the call ends: elapsed ms, status, and the reply's tokens when it has usage. + + ``reply`` is the ``AssistantMessage`` (or the merged ``AssistantMessageChunk``) on success and None otherwise. + A partial reply keeps whatever tokens it reported, but ``usage_known`` is only True when the call ended normally + (``status == "ok"``) with usage metadata: a stream that errored, was cancelled or interrupted after a partial + usage keeps its token counts while still marking the total unknown, never a confirmed zero. The ``error`` summary + keeps only the error's type, never its message, the prompt or a secret. + """ + record["ms"] = round((time.perf_counter() - started) * 1000) + record["status"] = status + if error is not None: + record["error"] = type(error).__name__ + if reply is None: + return + usage = getattr(reply, "usage_metadata", None) + calls = list(getattr(reply, "tool_calls", None) or []) + content = getattr(reply, "content", None) + record["input_tokens"] = usage.input_tokens if usage is not None else 0 + record["output_tokens"] = usage.output_tokens if usage is not None else 0 + record["cache_tokens"] = int(usage.cache_tokens or 0) if usage is not None else 0 + record["usage_known"] = usage is not None and status == "ok" + record["tool_calls"] = [str(call.name) for call in calls] + record["tool_args"] = [str(call.arguments) for call in calls] + record["content_chars"] = len(content) if isinstance(content, str) else 0 + + +def usage_known(calls: list[dict[str, Any]]) -> bool: + """Whether every recorded call reported its usage: one unknown call makes the whole total unknown, not zero. + + A record without the field predates the flag and was only ever appended on success, so it counts as known. + """ + return all(bool(call.get("usage_known", True)) for call in calls) + + +def call_digests(calls: list[dict[str, Any]]) -> list[dict[str, Any]]: + """The compact records worth keeping for a failed or unpriced call: no prompt, no raw tool arguments. + + Only the calls whose usage is unknown or whose status is not ``ok`` are kept, so an episode's ``extra`` stays + small while a failed or cancelled planner call remains traceable. + """ + return [ + {key: call.get(key) for key in RECORD_KEYS} + for call in calls + if call.get("status") != "ok" or not call.get("usage_known", True) + ] + + class CountingModel(Model): - """Forwards ``invoke`` and ``stream`` to ``inner``; appends one record per call to ``calls``.""" + """Forwards ``invoke`` and ``stream`` to ``inner``; appends one record per call to ``calls``. + + The record is appended before the call runs and updated in a ``finally``-safe way when it ends, so a failed, + timed-out or cancelled call is counted with its status and elapsed time. A cancellation or exception is never + swallowed: it is recorded and re-raised unchanged. + """ def __init__(self, inner: Model, calls: list[dict[str, Any]]) -> None: super().__init__(inner.model_client_config, inner.model_config) @@ -33,15 +101,37 @@ def __init__(self, inner: Model, calls: list[dict[str, Any]]) -> None: async def invoke(self, messages: Any, *, tools: Any = None, **kwargs: Any) -> AssistantMessage: started = time.perf_counter() - reply = await self._inner.invoke(messages, tools=tools, **kwargs) - self._calls.append(call_record(reply, started)) + record = start_record(started) + self._calls.append(record) + try: + reply = await self._inner.invoke(messages, tools=tools, **kwargs) + except asyncio.CancelledError: + finish_record(record, None, started, "cancelled") + raise + except BaseException as exc: # noqa: BLE001 - the caller's exception is recorded and re-raised unchanged + finish_record(record, None, started, "error", error=exc) + raise + finish_record(record, reply, started, "ok") return reply async def stream(self, messages: Any, *, tools: Any = None, **kwargs: Any) -> AsyncIterator[AssistantMessageChunk]: started = time.perf_counter() + record = start_record(started) + self._calls.append(record) merged: AssistantMessageChunk | None = None - async for chunk in self._inner.stream(messages, tools=tools, **kwargs): - merged = chunk if merged is None else merged + chunk - yield chunk - if merged is not None: - self._calls.append(call_record(merged, started)) + try: + async for chunk in self._inner.stream(messages, tools=tools, **kwargs): + merged = chunk if merged is None else merged + chunk + yield chunk + except asyncio.CancelledError: + finish_record(record, merged, started, "cancelled") + raise + except GeneratorExit: + # The consumer stopped early or closed the stream: keep the partial usage and let the close finish. + finish_record(record, merged, started, "interrupted") + raise + except BaseException as exc: # noqa: BLE001 - a failed stream is recorded, then re-raised unchanged + finish_record(record, merged, started, "error", error=exc) + raise + else: + finish_record(record, merged, started, "ok") diff --git a/s1a/desktop/env.py b/s1a/desktop/env.py index 6328a92..f7a6734 100644 --- a/s1a/desktop/env.py +++ b/s1a/desktop/env.py @@ -67,6 +67,12 @@ async def reset(self) -> None: "the window already shows --expect before any action; pass --clear