diff --git a/README.md b/README.md index 4db7a97..ec3e0d2 100644 --- a/README.md +++ b/README.md @@ -1,265 +1,172 @@ # openadapt-privacy -> [!IMPORTANT] -> **Status: Experimental.** The API is published on the 1.x version line, but -> the PHI/PII detector is backed by synthetic regression evidence rather than -> clinical validation. Scrubbing is one control in a reviewed egress process, -> not a guarantee that an artifact is free of protected data. -> -> The OpenAdapt product is the demonstration compiler, -> [`openadapt-flow`](https://github.com/OpenAdaptAI/openadapt-flow), installed -> via the [`OpenAdapt`](https://github.com/OpenAdaptAI/OpenAdapt) launcher -> (`pip install openadapt`): it compiles a demonstrated GUI workflow into a -> deterministic, locally executable program. Healthy runs make no model calls, -> and it halts instead of guessing when verification fails. Lifecycle labels for -> every repository are in the -> [repository lifecycle registry](https://github.com/OpenAdaptAI/.github/blob/main/REPOSITORY_LIFECYCLE.md). - [![Build Status](https://github.com/OpenAdaptAI/openadapt-privacy/actions/workflows/test.yml/badge.svg?branch=main)](https://github.com/OpenAdaptAI/openadapt-privacy/actions) [![PyPI version](https://img.shields.io/pypi/v/openadapt-privacy.svg)](https://pypi.org/project/openadapt-privacy/) -[![Downloads](https://img.shields.io/pypi/dm/openadapt-privacy.svg)](https://pypi.org/project/openadapt-privacy/) -[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT) [![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue)](https://www.python.org/downloads/) +[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT) -PHI/PII detection and redaction for GUI automation data. Presidio-backed -scrubbing of the text, structured element trees, and screenshots that a -recording captures. - -This is the trust and auditability asset of the OpenAdapt stack. A demonstration -recording can contain everything visible on screen and everything typed, so -`openadapt-privacy` gives the compiler and its operators a reviewable control for -removing protected data before an artifact is stored, shared, or exported. It -detects a fixed set of entity types, replaces them with typed placeholders, and -fails loud when its model is unavailable rather than scrubbing silently weaker. - -## The OpenAdapt stack - -OpenAdapt is a governed demonstration compiler: record a workflow once, compile -the recording into a deterministic program, and replay that program with zero -model calls on the healthy path. When the live screen does not match what was -demonstrated it halts instead of guessing, using identity gates and independent -effect verification. Every substrate is first-class: web and desktop recording -are validated, RDP and Windows replay are early, and Citrix is exploratory. - -| Package | Role | -| --- | --- | -| [`openadapt`](https://github.com/OpenAdaptAI/OpenAdapt) | Launcher and installer (`pip install openadapt`) | -| [`openadapt-flow`](https://github.com/OpenAdaptAI/openadapt-flow) | Records, compiles, verifies, and replays workflows | -| [`openadapt-capture`](https://github.com/OpenAdaptAI/openadapt-capture) | Cross-platform local desktop recording | -| [`openadapt-types`](https://github.com/OpenAdaptAI/openadapt-types) | Canonical action and UI-state schema | -| [`openadapt-grounding`](https://github.com/OpenAdaptAI/openadapt-grounding) | Local OCR text-anchoring plus optional model grounding | -| **`openadapt-privacy`** | PHI/PII detection and redaction (this package) | - -Documentation for the whole stack lives at -[docs.openadapt.ai](https://docs.openadapt.ai). +Finds and removes personal and health data from GUI automation artifacts: the +text a recording captured, the element trees around it, and the screenshots. +Presidio does the detection; this package wires it to the shapes a recording +actually has. -## Installation +A demonstration recording contains everything that was on screen and everything +that was typed. So before you store one, share one, or export one, something +has to strip the identifiers out, and that something needs to fail loudly when +it can't. That's this. -```bash -pip install openadapt-privacy -``` +[Documentation](https://docs.openadapt.ai) · +[openadapt-flow](https://github.com/OpenAdaptAI/openadapt-flow) · +[Repository lifecycle](https://github.com/OpenAdaptAI/.github/blob/main/REPOSITORY_LIFECYCLE.md) -For Presidio-based scrubbing (recommended): +## Install ```bash -pip install openadapt-privacy[presidio] +pip install "openadapt-privacy[presidio]" python -m spacy download en_core_web_sm ``` -### Model security and recall +The Presidio provider takes only the preinstalled, allowlisted `en_core_web_sm` +pipeline and never downloads a model at runtime. A missing model, an +unsupported language, or an inconsistent configuration raises +`PrivacyModelUnavailable` before analysis starts, rather than quietly scrubbing +worse than you expected. -The Presidio provider accepts only the preinstalled, allowlisted -`en_core_web_sm` pipeline. It never downloads a model at runtime. A missing -model, an unsupported language, or inconsistent model configuration raises -`PrivacyModelUnavailable` before analysis instead of continuing with weaker or -simulated scrubbing. +## Read this before you rely on it -`tests/test_phi_recall.py` is a quantitative regression gate covering 24 -synthetic identifiers across names, contact information, financial identifiers, -dates of birth, addresses, network identifiers, medical record numbers, member -IDs, and provider licenses. The current gate requires 24/24 detections and also -checks clean operational UI text for false positives. +Scrubbing is one control inside a reviewed egress process. It is not a +guarantee that an artifact is free of protected data, and the evidence behind +it is synthetic, not clinical. -This synthetic corpus is regression evidence, not clinical validation or a -guarantee that an artifact is PHI-free. Production egress should scrub a copy, -verify every output file, and bind human or policy approval to the verified -artifact rather than treating successful model execution as sufficient. +`tests/test_phi_recall.py` is the regression gate: 24 synthetic identifiers +across names, contact details, financial identifiers, dates of birth, +addresses, network identifiers, medical record numbers, member IDs, and +provider licenses. It requires 24 out of 24, and it also checks that ordinary +operational UI text comes back untouched. -## Quick Start - -### Text Scrubbing +Detection is contextual, so the same value scrubs differently depending on what +surrounds it. Every output below was produced by running 1.0.3 on 2026-08-28, +not written by hand. Some of it will surprise you: ```python -from openadapt_privacy.providers.presidio import PresidioScrubbingProvider +>>> from openadapt_privacy.providers.presidio import PresidioScrubbingProvider +>>> s = PresidioScrubbingProvider() -scrubber = PresidioScrubbingProvider() +>>> s.scrub_text("Patient John Smith, john.smith@example.com, 555-123-4567, SSN 923-45-6789") +'Patient , , , ' -text = "Contact John Smith at john.smith@example.com or 555-123-4567" -scrubbed = scrubber.scrub_text(text) -``` +>>> s.scrub_text("SSN: 923-45-6789") +': ' -**Input:** -``` -Contact John Smith at john.smith@example.com or 555-123-4567 +>>> s.scrub_text("Card 4111111111111111 on file") +' on file' ``` -**Output:** -``` -Contact at or -``` +The last one is the important one. A card number gets redacted, but as +`DATE_TIME`, not as `CREDIT_CARD`, and the label "Card" goes with it. The +redaction holds; the entity type you get back is not the one you would predict. +Do not build a policy that keys off the placeholder name without measuring it +against your own data first. -### Example Inputs & Outputs +For production egress: scrub a copy, verify every output file, and bind the +human or policy approval to the verified artifact. A model that ran without +error is not evidence that the artifact is clean. -| Input | Output | -|-------|--------| -| `My email is john@example.com` | `My email is ` | -| `SSN: 923-45-6789` | `SSN: ` | -| `Card: 4532-1234-5678-9012` | `Card: ` | -| `Call me at 555-123-4567` | `Call me at ` | -| `DOB: 01/15/1985` | `DOB: ` | -| `Contact John Smith` | `Contact ` | +## Scrubbing text -## Dict Scrubbing +```python +from openadapt_privacy.providers.presidio import PresidioScrubbingProvider -Scrub PHI/PII from nested dictionaries (e.g., GUI element trees): +scrubber = PresidioScrubbingProvider() +scrubber.scrub_text("Contact John Smith at john.smith@example.com or 555-123-4567") +# 'Contact at or ' +``` + +## Scrubbing nested dicts + +For element trees and event payloads: ```python from openadapt_privacy import scrub_dict from openadapt_privacy.providers.presidio import PresidioScrubbingProvider scrubber = PresidioScrubbingProvider() -action = { - "text": "Email: john@example.com", - "metadata": { +scrub_dict( + { "title": "User Profile - John Smith", "tooltip": "Click to contact john@example.com", + "value": "Call 555-123-4567", + "coordinates": {"x": 100, "y": 200}, }, - "coordinates": {"x": 100, "y": 200}, -} -scrubbed = scrub_dict(action, scrubber) + scrubber, +) ``` -**Input:** ```json { - "text": "Email: john@example.com", - "metadata": { - "title": "User Profile - John Smith", - "tooltip": "Click to contact john@example.com" - }, + "title": "User Profile - ", + "tooltip": "Click to contact ", + "value": "Call ", "coordinates": {"x": 100, "y": 200} } ``` -**Output:** -```json -{ - "text": "Email: ", - "metadata": { - "title": "User Profile - ", - "tooltip": "Click to contact " - }, - "coordinates": {"x": 100, "y": 200} -} -``` +Only the keys in `PrivacyConfig.SCRUB_KEYS_HTML` are scrubbed, and non-string +values pass through, which is why the coordinates survive. Pass `scrub_all=True` +to scrub every string regardless of key. -## Recording Pipeline +One caveat in 1.0.3: the `text` key is treated as character-separated action +text, joined by `ACTION_TEXT_SEP` (`-`). Scrubbing a plain sentence under that +key returns it hyphenated, one character at a time. Use `value` or `title` for +ordinary prose until that's fixed. -Process complete GUI automation recordings: +## Scrubbing screenshots ```python -from openadapt_privacy import DictRecordingLoader +from PIL import Image from openadapt_privacy.providers.presidio import PresidioScrubbingProvider scrubber = PresidioScrubbingProvider() -loader = DictRecordingLoader() - -recording = loader.load_from_dict({ - "task_description": "Send email to John Smith at john@example.com", - "actions": [ - {"id": 1, "action_type": "click", "text": "Compose", "timestamp": 1000}, - {"id": 2, "action_type": "type", "text": "john@example.com", "timestamp": 2000}, - {"id": 3, "action_type": "click", "text": "Send", "window_title": "Email to john@example.com", "timestamp": 3000}, - ], -}) - -scrubbed = recording.scrub(scrubber) +scrubbed = scrubber.scrub_image(Image.open("screenshot.png")) +scrubbed.save("screenshot_scrubbed.png") ``` -**Input Recording:** -``` -task_description: "Send email to John Smith at john@example.com" - -actions: - [1] click: "Compose" - [2] type: "john@example.com" - [3] click: "Send" (window: "Email to john@example.com") -``` - -**Output Recording:** -``` -task_description: "Send email to at " +![Original screenshot with PII](assets/screenshot_original.png) +![Scrubbed screenshot with PII redacted](assets/screenshot_scrubbed.png) -actions: - [1] click: "Compose" - [2] type: "" - [3] click: "Send" (window: "Email to ") -``` +OCR finds the text regions, the analyzer classifies them, and the detected +regions get filled with a solid colour (`SCRUB_FILL_COLOR`, red by default). +Anything OCR misses is not redacted. -## Image Scrubbing +## Whole recordings -Redact PHI/PII from screenshots using OCR + NER: +`DictRecordingLoader` takes a recording as a dict and scrubs the task +description and every action together: ```python -from PIL import Image +from openadapt_privacy import DictRecordingLoader from openadapt_privacy.providers.presidio import PresidioScrubbingProvider -scrubber = PresidioScrubbingProvider() - -image = Image.open("screenshot.png") -scrubbed_image = scrubber.scrub_image(image) -scrubbed_image.save("screenshot_scrubbed.png") +recording = DictRecordingLoader().load_from_dict({ + "task_description": "Send email to John Smith at john@example.com", + "actions": [ + {"id": 1, "action_type": "click", "text": "Compose", "timestamp": 1000}, + {"id": 2, "action_type": "click", "text": "Send", + "window_title": "Email to john@example.com", "timestamp": 3000}, + ], +}) +scrubbed = recording.scrub(PresidioScrubbingProvider()) ``` -**Input Screenshot:** - -![Original screenshot with PII](assets/screenshot_original.png) - -**Output Screenshot:** - -![Scrubbed screenshot with PII redacted](assets/screenshot_scrubbed.png) - -The image redactor: -1. Runs OCR to detect text regions -2. Analyzes text for PII entities (email, phone, SSN, etc.) -3. Fills detected PII regions with solid color (configurable, default: red) - -## Custom Data Loader - -Implement your own loader for custom storage formats: +Subclass `RecordingLoader` for your own storage. Implement `load` and `save`, +and you get `load_and_scrub` for free: ```python from openadapt_privacy import RecordingLoader, Recording class SQLiteRecordingLoader(RecordingLoader): - def __init__(self, db_path: str): - self.db_path = db_path - - def load(self, recording_id: str) -> Recording: - # Load from SQLite database - ... - - def save(self, recording: Recording, recording_id: str) -> None: - # Save to SQLite database - ... - -# Usage -loader = SQLiteRecordingLoader("recordings.db") -scrubber = PresidioScrubbingProvider() - -# Load, scrub, and save -scrubbed = loader.load_and_scrub("recording_001", scrubber) -loader.save(scrubbed, "recording_001_scrubbed") + def load(self, recording_id: str) -> Recording: ... + def save(self, recording: Recording, recording_id: str) -> None: ... ``` ## Configuration @@ -267,41 +174,34 @@ loader.save(scrubbed, "recording_001_scrubbed") ```python from openadapt_privacy.config import PrivacyConfig -custom_config = PrivacyConfig( - SCRUB_CHAR="X", # Character for scrub_text_all - SCRUB_FILL_COLOR=0xFF0000, # Red for image redaction (BGR) - SCRUB_KEYS_HTML=[ # Keys to scrub in dicts - "text", "value", "title", "tooltip", "custom_field" - ], - SCRUB_PRESIDIO_IGNORE_ENTITIES=[ # Entity types to skip - "DATE_TIME", - ], +PrivacyConfig( + SCRUB_CHAR="X", # for scrub_text_all + SCRUB_FILL_COLOR=0xFF0000, # image redaction, BGR + SCRUB_KEYS_HTML=["text", "value", "title", "tooltip"], + SCRUB_PRESIDIO_IGNORE_ENTITIES=["DATE_TIME"], ) ``` -## Supported Entity Types +The full field list is `SCRUB_CHAR`, `SCRUB_LANGUAGE`, `SCRUB_FILL_COLOR`, +`SCRUB_KEYS_HTML`, `ACTION_TEXT_NAME_PREFIX`, `ACTION_TEXT_NAME_SUFFIX`, +`ACTION_TEXT_SEP`, `SCRUB_CONFIG_TRF`, `SCRUB_PRESIDIO_IGNORE_ENTITIES`, and +`SPACY_MODEL_NAME`. -| Entity | Example Input | Example Output | -|--------|---------------|----------------| -| `PERSON` | `John Smith` | `` | -| `EMAIL_ADDRESS` | `john@example.com` | `` | -| `PHONE_NUMBER` | `555-123-4567` | `` | -| `US_SSN` | `923-45-6789` | `` | -| `CREDIT_CARD` | `4532-1234-5678-9012` | `` | -| `US_BANK_NUMBER` | `635526789012` | `` | -| `US_DRIVER_LICENSE` | `A123-456-789-012` | `` | -| `DATE_TIME` | `01/15/1985` | `` | -| `LOCATION` | `Toronto, ON` | `` | +The analyzer's supported entity set comes from Presidio: `CREDIT_CARD`, +`CRYPTO`, `DATE_TIME`, `EMAIL_ADDRESS`, `IBAN_CODE`, `IP_ADDRESS`, `LOCATION`, +`MAC_ADDRESS`, `MEDICAL_LICENSE`, `NRP`, `PERSON`, `PHONE_NUMBER`, `UK_NHS`, +`URL`, `US_BANK_NUMBER`, `US_DRIVER_LICENSE`, `US_ITIN`, `US_PASSPORT`, +`US_SSN`, plus `ORGANIZATION` from the spaCy pipeline. Which one fires on a +given string depends on the text around it, so measure rather than assume. -## Architecture +## Layout ``` openadapt_privacy/ ├── base.py # ScrubbingProvider, TextScrubbingMixin -├── config.py # PrivacyConfig dataclass +├── config.py # PrivacyConfig ├── loaders.py # Recording, Action, Screenshot, RecordingLoader ├── providers/ -│ ├── __init__.py # ScrubProvider registry │ └── presidio.py # PresidioScrubbingProvider └── pipelines/ └── dicts.py # scrub_dict, scrub_list_dicts @@ -309,4 +209,4 @@ openadapt_privacy/ ## License -MIT +[MIT](LICENSE)