What this enables
The desktop agent can use a screenshot to choose between configured click targets when accessibility data does not identify the controls. A local Cua-S1 4B model receives the current image and selects an offered action.
Features: #14 for the Cua-S1 4B adapter, and #15 for screenshot-based desktop target selection. Native text, keyboard, scrolling, and save-dialog actions are covered by RFC #26.
Implemented design
s1a/decision_models/cua_four_b.py wraps the official 4B scorer through the existing decision-model interface. It supports the text and multimodal adapters, maps scores by option key, and uses the shared reply validation. The model is an optional cua-four-b extra, selected with CUA_S1_VARIANT=4b.
- The desktop driver captures an image with its capture ID and dimensions. The environment passes that image through
Observation.images to the multimodal model. Image bytes stay out of serialized observation state and run logs.
--pixel-target name=x,y defines candidates as fractions of the current screenshot. The model chooses among those candidates. The driver executes the selected point against the same capture, with window binding and bounds checks.
- Multimodal requests require one current screenshot. Pixel-target tasks require a model that accepts images. The scorer supports up to 26 choices.
visual_fixture.swift draws Save and Cancel tiles without accessibility children. Reset swaps their positions. The app records every selected tile in a temporary file; the run checks all four selections independently.
Implementation status
Implementation: #32. It includes the Cua-S1 4B adapter, desktop screenshot flow, native test app, automated tests, and run instructions.
This RFC records the adapter and desktop screenshot integration within #14 and #15. The wider model comparison remains tracked in #14.
Verification
Added test_decision_models_cua_four_b.py and test_desktop_visual.py, and extended the model factory tests. They check option mapping, invalid outputs, image delivery and cleanup, screenshot metadata, capture-bound clicks, dry runs, and rejection of missing images or text-only models for pixel targets.
The implementation passes 507 tests, with 41 skipped. Ruff, ty, lock consistency, package build, shell syntax, and CLI/MCP smoke checks pass.
Tested on macOS 26.2 arm64 with Cua Driver 0.30.1 and local Cua-S1 4B 0.2 multimodal on MPS with bfloat16:
| Task |
Passed |
Total time |
Mean time |
Clicks per trial |
| Select Save from the screenshot as the tiles swap positions |
4/4 |
44.7 s |
11.2 s |
1 |
| Fixed correct-point check |
1/1 |
3.0 s |
3.0 s |
1 |
The model selected right, left, right, left as Save moved. All four selections were checked against the app's output file. Median model decision time was 7.369 s. Episode times exclude model loading. The model chose the actions without a fixed plan or hosted model API calls.
An intentionally wrong-point check was rejected and recorded Cancel in the output file.
How to run
Requires macOS, uv, Swift command-line tools, and Cua Driver with Accessibility and Screen Recording permissions. Keep the Mac unlocked and close any older S1A visual fixture. Run from the current project directory:
S1A_PROJECT_DIR="$(pwd -P)"
uv sync --project "$S1A_PROJECT_DIR" --extra cua-four-b
bash "$S1A_PROJECT_DIR/evals/desktop/build_visual_fixture.sh" /tmp/S1AVisualFixture.app
export CUA_DRIVER_BIN=/Applications/CuaDriver.app/Contents/MacOS/cua-driver
export CUA_S1_VARIANT=4b CUA_S1_MODALITY=multimodal
export CUA_S1_DEVICE=mps CUA_S1_DTYPE=bfloat16 HF_DEACTIVATE_ASYNC_LOAD=1
export PYTORCH_ENABLE_MPS_FALLBACK=1
: > /tmp/s1a-visual-target.txt
uv run --project "$S1A_PROJECT_DIR" --no-sync s1a run desktop \
--model cua --rethink off --episodes 4 --max-steps 1 \
--app S1AVisualFixture --app-path /tmp/S1AVisualFixture.app \
--window-title "S1A Visual Fixture" --expect Saved \
--goal "Click the tile labelled Save in the screenshot" \
--pixel-target left=0.27,0.51 --pixel-target right=0.73,0.51 \
--clear Reset --execute
uv run --project "$S1A_PROJECT_DIR" --no-sync python -c \
'from pathlib import Path; rows = Path("/tmp/s1a-visual-target.txt").read_text().splitlines(); assert rows == ["Save selected"] * 4, rows; print("Verified all four clicks")'
The base model and adapter download on first use. CUA_S1_BASE_MODEL and CUA_S1_CHECKPOINT can point to downloaded model directories.
What this enables
The desktop agent can use a screenshot to choose between configured click targets when accessibility data does not identify the controls. A local Cua-S1 4B model receives the current image and selects an offered action.
Features: #14 for the Cua-S1 4B adapter, and #15 for screenshot-based desktop target selection. Native text, keyboard, scrolling, and save-dialog actions are covered by RFC #26.
Implemented design
s1a/decision_models/cua_four_b.pywraps the official 4B scorer through the existing decision-model interface. It supports the text and multimodal adapters, maps scores by option key, and uses the shared reply validation. The model is an optionalcua-four-bextra, selected withCUA_S1_VARIANT=4b.Observation.imagesto the multimodal model. Image bytes stay out of serialized observation state and run logs.--pixel-target name=x,ydefines candidates as fractions of the current screenshot. The model chooses among those candidates. The driver executes the selected point against the same capture, with window binding and bounds checks.visual_fixture.swiftdraws Save and Cancel tiles without accessibility children. Reset swaps their positions. The app records every selected tile in a temporary file; the run checks all four selections independently.Implementation status
Implementation: #32. It includes the Cua-S1 4B adapter, desktop screenshot flow, native test app, automated tests, and run instructions.
This RFC records the adapter and desktop screenshot integration within #14 and #15. The wider model comparison remains tracked in #14.
Verification
Added
test_decision_models_cua_four_b.pyandtest_desktop_visual.py, and extended the model factory tests. They check option mapping, invalid outputs, image delivery and cleanup, screenshot metadata, capture-bound clicks, dry runs, and rejection of missing images or text-only models for pixel targets.The implementation passes 507 tests, with 41 skipped. Ruff, ty, lock consistency, package build, shell syntax, and CLI/MCP smoke checks pass.
Tested on macOS 26.2 arm64 with Cua Driver 0.30.1 and local Cua-S1 4B 0.2 multimodal on MPS with bfloat16:
The model selected right, left, right, left as Save moved. All four selections were checked against the app's output file. Median model decision time was 7.369 s. Episode times exclude model loading. The model chose the actions without a fixed plan or hosted model API calls.
An intentionally wrong-point check was rejected and recorded Cancel in the output file.
How to run
Requires macOS,
uv, Swift command-line tools, and Cua Driver with Accessibility and Screen Recording permissions. Keep the Mac unlocked and close any older S1A visual fixture. Run from the current project directory:The base model and adapter download on first use.
CUA_S1_BASE_MODELandCUA_S1_CHECKPOINTcan point to downloaded model directories.