Skip to content

Add the Global AI Workspace - #1004

Open
whyisjake wants to merge 58 commits into
developfrom
feat/ai-workspace
Open

whyisjake wants to merge 58 commits into
developfrom
feat/ai-workspace

Conversation

@whyisjake

@whyisjake whyisjake commented Sep 4, 2026 •

Copy link
Copy Markdown
Member

Implements #282: a full-screen conversational admin screen where a site owner talks to an AI that can read their content — under their own capabilities, never above them.

What this adds

A new ai-workspace experiment (Tools → AI Workspace, manage_options), built on the Abilities API and core's ability-backed tool loop:

  • A turn endpoint and tool loop. Each function call runs through the resolver's single-call execute_ability() rather than the batch path, so every invocation — including denials — produces exactly one log row and a provenance envelope.
  • Two read abilities. ai/search-content returns titles and excerpts only; ai/read-content-bodies returns full bodies for at most five posts named by ID. Both filter row by row at execute time against the requesting user's own capabilities.
  • Propose-then-confirm writes. Draft creation is proposed with resolved values a person approves before anything is written. No write ability is registered.
  • Buffered replies. Every reply arrives whole. Token-by-token streaming is split into the follow-up AI Workspace: stream provider replies once the PHP AI Client can #1083 (see below).
  • A block editor handoff, opening the workspace seeded with the current post's identity — never its body, which would be a second, unenforced way in.
  • A retrieval trace, reporting per invocation what was searched, what came back, and what permission withheld.

Notes for review

Permission filtering is execute-time, not declaration-time. The coarse capability on a tool decides whether it is declared; which rows come back is decided per row, per user, on every call. Both matter, and the tests mutate each independently to prove neither is inert.

Withheld counts are measured, not derived — a deliberate divergence from the plan. The plan's U13 prescribed the opposite, specifying the withheld count as "the difference between the ability's total and its returned rows". Implementing it revealed that verification to be wrong, so it was not followed. total - count(results) conflates pagination with permission filtering: page one of fifty matches would announce forty posts hidden by the person's role on an ordinary search. ai/search-content counts withholding as the permission walk drops rows from the page it builds. ai/read-content-bodies deliberately reports no count at all — it answers for IDs the caller named, so counting unknown and unreadable apart would say whether those exact posts exist.

Tool surface is a hand-maintained allowlist of three abilities, which is deliberate for a first cut and is the subject of follow-up #1003 — every ability the model can call is also reachable by an instruction embedded in content someone else wrote.

Streaming is a follow-up (#1083)

Provider streaming depended on WordPress/php-ai-client#255, which is not merged or released, so this PR used to carry a vendored copy of that PR's streaming classes. It no longer does. Streaming now lives in #1083, stacked on this PR, as a single commit that reverts its removal. It can merge once #255 is released and re-vendored from that release.

What that means for this PR:

  • No unreleased upstream dependency. Everything here runs on released code: the core AI client (WordPress 7.0+) and the Abilities API, including ability filtering and the public meta flag from WordPress 7.1.
  • Replies arrive whole. Every turn uses the buffered path the workspace already used for providers other than Anthropic.
  • No new vendored code. The embeddings files already on develop under includes/Vendor/AiClient/ are unchanged; the streaming overlay entry and its classes are gone.

Testing

  • npm run test:php — 1,990 tests, 6,045 assertions, 42 skipped, 0 failures (on 68425d6).
  • npm run test:e2e — all 38 AI Workspace e2e tests pass, against mocked providers with sequenced tool-calling turns. No network access.
  • npm run lint:php (PHPCS, WordPress-VIP-Go + Slevomat) — clean.
  • npm run lint:php:stan (PHPStan level 8) — no errors.
  • npm run build, npm run typecheck, npm run lint:js — clean.

Test suites executed by name, as the WordPress AI Guidelines require:

PHPUnit (tests/Integration/) — AI_WorkspaceTest, Turn_ControllerTest, Tool_SelectorTest, Proposal_ControllerTest, Propose_DraftsTest, Prompt_Model_ClientTest, Search_ContentTest, Read_Content_BodiesTest, SDK_OverlayTest, UninstallTest.

Playwright (tests/e2e/specs/experiments/) — ai-workspace.spec.js, ai-workspace-tools.spec.ts, ai-workspace-post-results.spec.ts, ai-workspace-markdown.spec.ts, ai-workspace-handoff.spec.ts.

Individual tests named in the mutation evidence below: test_execute_callback_clamps_to_five_without_schema_validation, test_private_body_of_another_author_is_withheld_from_author, test_draft_body_of_another_author_is_withheld_from_lower_roles, test_unexposed_post_types_are_never_read, test_mixed_request_returns_only_readable_posts, test_retrieval_summary_never_counts_paginated_rows_as_withheld, test_pagination_alone_reports_nothing_withheld, test_withheld_counts_the_rows_the_permission_walk_dropped, test_vendored_files_use_the_prefixed_psr_dependencies.

Security-relevant tests were verified by mutation rather than by passing alone. Forcing Read_Content_Bodies::check_read_permission() open and removing its clamp fails five tests, including one asserting an unexposed post type's body sentinel appears nowhere in the encoded result. Replacing the withheld count with subtraction fails test_retrieval_summary_never_counts_paginated_rows_as_withheld and test_pagination_alone_reports_nothing_withheld.

Still open

Opening as a draft — the presentation layer is being worked separately, and three decisions are unresolved: the confirmation modal's default selection state, the full-screen gap (is-fullscreen-mode is applied but @wordpress/interface CSS is never enqueued), and the editor handoff's placement in the Options menu.


AI assistance: Yes. Tool(s): Claude Code (Opus 5). Used for: implementation of all units in the linked plan, test authoring including the mutation testing described above, and this description.

@whyisjake — this line needs your sign-off and I have deliberately not written it for you. The WordPress AI Guidelines require the contributor to understand every line submitted and be able to explain it under review. That attestation is yours to make or reword, and it is not something an agent can satisfy on your behalf. Replace this block before marking the PR ready.

🤖 Generated with Claude Code

https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L

Open WordPress Playground Preview

@github-actions

github-actions Bot commented Sep 4, 2026 •

Copy link
Copy Markdown

✅ WordPress Plugin Check Report

✅ Status: Passed

📊 Report

All checks passed! No errors or warnings found.


🤖 Generated by WordPress Plugin Check Action • Learn more about Plugin Check

@codecov

codecov Bot commented Sep 4, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 86.43596% with 341 lines in your changes missing coverage. Please review.
✅ Project coverage is 83.87%. Comparing base (da700e0) to head (225802b).

Files with missing lines Patch % Lines
...Experiments/AI_Workspace/REST/Stream_Responder.php 4.91% 58 Missing ⚠️
...cludes/Experiments/AI_Workspace/Proposal_Store.php 70.16% 37 Missing ⚠️
...Experiments/Abilities_Explorer/Ability_Handler.php 67.70% 31 Missing ⚠️
includes/Abilities/Content/Search_Content.php 91.56% 29 Missing ⚠️
includes/Abilities/Content/Read_Content_Bodies.php 88.52% 28 Missing ⚠️
includes/Experiments/AI_Workspace/Turn_Runner.php 90.07% 27 Missing ⚠️
...s/Experiments/AI_Workspace/Prompt_Model_Client.php 54.90% 23 Missing ⚠️
...eriments/AI_Workspace/Function_Calling_Support.php 0.00% 22 Missing ⚠️
.../Experiments/AI_Workspace/REST/Turn_Controller.php 88.32% 16 Missing ⚠️
...cludes/Experiments/AI_Workspace/Propose_Drafts.php 90.97% 13 Missing ⚠️
... and 10 more
Additional details and impacted files
@@              Coverage Diff              @@
##             develop    #1004      +/-   ##
=============================================
+ Coverage      81.53%   83.87%   +2.33%     
- Complexity      3071     3737     +666     
=============================================
  Files            129      148      +19     
  Lines          12253    14352    +2099     
=============================================
+ Hits            9991    12038    +2047     
- Misses          2262     2314      +52     
Flag Coverage Δ
unit 83.87% <86.43%> (+2.33%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actions

github-actions Bot commented Sep 5, 2026 •

Copy link
Copy Markdown

The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the props-bot label.

If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message.

Co-authored-by: whyisjake <[email protected]>
Co-authored-by: jeffpaul <[email protected]>
Co-authored-by: thelovekesh <[email protected]>

To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook.

@jeffpaul

Copy link
Copy Markdown
Member

@whyisjake we'll likely want the streaming PR in the PHP AI client into a tagged release there as a precursor to this PR (and perhaps something in the Requests library as well?), so if you could help on those upstream effort(s) and then come back to this PR that'll help keep things unblocked.

@whyisjake

whyisjake commented Sep 23, 2026 •

Copy link
Copy Markdown
Member Author

Thanks @jeffpaul, agreed on the sequencing.

WordPress/php-ai-client#255 is rebased and mergeable and is waiting on @JasonTheAdams's review. The streaming subset vendored here is byte-identical to the current head of that PR, apart from the core-prefixed PSR imports, and the AI Workspace test suites run it under core's prefixed environment. I'll post that as a consumer review there and weigh in on the two open design questions (the opt-in streaming interface, and aborting a stream mid-way).

On Requests: WP's HTTP API buffers the whole body, so this branch opens the stream with fopen() and re-applies wp_http_validate_url, connector approval, and logging itself. Requests already fires request.progress per chunk, exposed in WP as the requests-request.progress action, but that is push-based. The StreamedGenerativeAiResult in WordPress/php-ai-client#255 pulls from a PSR-7 stream. Bridging the two needs Requests or WP_Http to return a lazy body stream. If that's the upstream change you had in mind, I'll open an issue proposing it before doing anything more here.

This stays a draft until WordPress/php-ai-client#255 is in a tagged release, then I'll re-vendor from trunk.

whyisjake and others added 13 commits September 23, 2026 08:44
…reen

Adds a full-screen, capability-gated admin screen behind its own experiment
toggle. The screen renders an app shell only; conversation behaviour, REST
routes, and streaming land in later units.

The capability check is layered rather than single-point: the menu entry, the
asset enqueue, and the render callback each gate on manage_options, so a direct
call to the render callback cannot emit the app shell or its localized data to
a user without the capability.

Advances R1, R2, R3, R4 (U1).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L
Registers ai/search-content, a bounded full-text search over the post types
exposed to the Abilities API. core/read-content is exact-match only and is kept
byte-similar to core's copy, so this lands as a sibling rather than a change to
it.

Results are filtered at execute time against the requesting user's read
permission, mirroring core/read-content's permission walk including the
inherited-parent chain, so every ability consumer inherits the same filtering
rather than relying on the query alone. The permission callback can only gate
coarsely; the row filter is the authoritative check.

Two query behaviours worth noting. per_page is capped at 20 in the schema and
clamped again in the callback, so the cap holds on transports that skip schema
validation. perm => 'readable' is passed only for a single-post-type query:
WP_Query resolves a multi-type query to the placeholder capability
read_private_multiple_post_types, which no role holds, silently hiding private
posts from users who can read them.

Registered from the workspace experiment rather than the Custom Abilities
gated collection, so the workspace always has its search tool and every other
ability consumer can still reach it.

Advances R12, R13 (U2).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L
…overlay

Adds an independently-gated `streaming` feature to SDK_Overlay, alongside
`embeddings`, so environments whose bundled PHP AI Client predates streaming
get the streaming types from a vendored copy.

A spike established what the plan had wrong. WordPress core bundles the SDK
without its vendor directory, so upstream PR #255's Guzzle-based streaming
transport does not apply here; core supplies its own transporter over
wp_safe_remote_request. Only the SDK-level types are useful, and a WordPress
streaming transporter is required regardless of whether that PR merges.

Vendored files are NOT verbatim, unlike the embeddings feature. Core prefixes
its PSR and Nyholm dependencies under WordPress\AiClientDependencies, so every
such import is rewritten; the unprefixed names do not resolve under a WordPress
bootstrap. The README records the rewrite table.

PromptBuilder and AiClient are deliberately excluded. PromptBuilder is 45KB of
trunk-era code and the largest drift risk against the bundled 0.3.1, and the
method it adds only validates, resolves the model, and delegates. Two further
candidates were trimmed as unreferenced by the kept set.

The sentinel is StreamedGenerativeAiResult rather than the streaming model
interface: resolve() probes with class_exists(), which returns false for
interfaces, so an interface sentinel would activate the feature even where the
environment already ships streaming. A new test enforces that for every feature.

Streaming is not yet reachable end to end: the transport and an Anthropic model
implementing the streaming interface are still missing (U12).

Adds U11 (new; the plan has no unit for this work).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L
Makes streaming reachable in PHP. A Streaming_Http_Transporter decorates the
configured transporter: non-streaming calls delegate unchanged, and a streaming
request is opened through fopen() with a stream context so the body is pulled
lazily and wrapped as a PSR-7 stream the SDK's SSE parser reads unmodified.

Because this path bypasses wp_safe_remote_request(), it also bypasses connector
approval and request logging, so both are restored explicitly. Approval mirrors
Http_Guard step for step -- same Connector_Key_Index lookup, same
Caller_Identifier, same Approvals_Store check, same pending-approval record --
and runs before the opener is touched. A test asserts zero calls to the opener,
which is the only egress point, so "before egress" is proven rather than
implied. Logging wraps the whole body in a finally, so a refused request is
logged as an error too. Two honest limits are documented on the class: a
streamed entry records time-to-headers, not time-to-last-byte, and carries no
token counts, because those arrive inside a body that has not been read yet.

The same bypass loses that function's SSRF protection, so wp_http_validate_url()
is applied before connecting and CRLF is stripped from header names and values.

Anthropic's stream is mapped in its own class. Its SSE uses named events rather
than a single delta shape, which is why upstream PR #255 does not cover it, and
tool arguments arrive as input_json_delta fragments that must be concatenated
across deltas before parsing -- parsing a fragment early yields plausible but
wrong arguments. Interleaved fragments from concurrent tool calls are covered.

The transporter is deliberately not installed globally and nothing constructs
the streaming model in production code; ordering against Logging_Http_Transporter
in the registry belongs to the turn endpoint that will drive it.

phpstan.neon.dist excludes the two classes that implement or extend SDK and
provider symbols PHPStan cannot resolve. Verified necessary: without it the run
reports 15 errors that PHPStan itself marks unignorable. This follows the
existing precedent for Logging_Http_Transporter; the mapper, opener and
interface stay fully analysed.

Adds U12 (new; the plan has no unit for this work).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L
The spike in U3 contradicted three assumptions KTD3 rested on, and its outcome
added two units the plan did not contain. Brings the artifact back in line with
the code so U4 is planned against what exists.

- KTD3 records the corrected findings: upstream PR #255 streams via Guzzle,
  which core does not bundle, so a WordPress-side transport is needed either
  way; vendoring cannot be verbatim because core prefixes its PSR dependencies;
  and the guards are restorable by construction through the public transporter
  seam rather than bypassed.
- U3 is marked complete with its outcome; U11 and U12 are added.
- U2's file list pointed at the gated-abilities collection, which accepts only
  gated base classes and runs only when a different experiment is enabled.
- U4's approach corrected on three counts found by review and confirmed in the
  code: the batch resolver call offers no per-call seam for provenance or
  logging; a null-input permission call cannot filter tools because this repo's
  content callbacks are input-dependent; and cancellation cannot rely on abort
  detection on a buffered turn.
- Stop conditions and the streaming open question updated: a host that cannot
  stream now degrades to a buffered request instead of blocking the plan.

Also adds the Anthropic provider to the wp-env configs, since Claude is the
chosen provider and the environment previously loaded only Google and OpenAI.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L
…l loop

Runs a bounded, permission-filtered, logged tool-calling conversation behind
POST ai/v1/workspace/messages, with a companion cancel route. Both gate on
capability independently of nonce validation.

Tool selection uses a coarse, input-free capability predicate. Filtering with a
null-input permission check does not work here: this repository's content
permission callbacks are input-dependent and return false without a post id or
slug, so that filter would deny every tool to every user including
administrators. Object-level authorization stays at execute time, where
WP_Ability::execute() runs the ability's own callback.

The loop iterates the assistant message's function-call parts and invokes the
resolver's single-call execute_ability() per call. The batch form runs every
call internally and exposes no hooks, leaving no seam for the provenance
envelope or for one log row per invocation.

Tool results are wrapped as provenance-tagged data before returning to the
model, and every invocation writes exactly one log row -- allowed, denied, and
failed alike. Denials are distinguishable from failures on the indexed status
column without decoding context, and context.surface discriminates workspace
rows from MCP rows sharing the same ability name.

Cancellation is out of band: a separate route sets a marker the loop re-reads
between rounds. Client-abort detection cannot serve here because PHP observes a
disconnect only after writing output, which a buffered turn never does and the
strict no-output test setting forbids. The marker is written by a different
request, so the read drops its cache entries first -- otherwise the loop would
answer from the value cached at request start and never observe a cancel.

Conversation state lives in user-scoped transients: session-lifetime,
regenerable, and expiring, which matters because it accumulates retrieved
private and draft post bodies. Ownership is enforced twice -- the key is derived
from the user id, and the stored owner is compared again on load -- so a second
user holding the same capability cannot read another's conversation. A miss is
a 404 rather than a silently-new conversation.

Streaming is driven through U12's transport behind a seam, and every reference
to the provider-dependent model is isolated so nothing autoloads where the
Anthropic plugin is absent. A transport that cannot stream on this host falls
back to a buffered request rather than failing.

phpstan.neon.dist excludes one small driver file whose provider parent is
unresolvable; the errors are non-ignorable and the surface was deliberately
isolated so the rest of the loop stays fully analysed.

Advances R6, R7, R8, R9, R10, R13, R18, R20, R21 (U4).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L
…mission

Builds the conversational surface: transcript with tool steps, prompt input
with send/stop/clear, the two context scopes, and the streaming client. Server
side, Stream_Responder consumes the filter seam the turn endpoint left and
emits SSE frames.

Model output is rendered by a restricted-subset markdown renderer rather than a
parser plus sanitizer. The renderer emits a node tree, never an HTML string, and
the components map nodes to elements, so dangerouslySetInnerHTML appears nowhere
in the unit. Dangerous constructs are impossible to emit rather than stripped
after parsing.

Links get the same treatment as images, which the plan understated. Blocking
images alone leaves the exfiltration path half open: a model steered by injected
post content can emit a link whose destination carries retrieved private content
in its query string, behind anchor text that reads as a legitimate next step. A
link is live only when it resolves to http or https on this host; everything
else renders as text with its destination visible. Images are inert including
reference-style syntax, which is not implemented at all.

Streaming asks for SSE and branches on the response content type, so a host that
cannot stream renders the identical turn arriving at once with a quiet line
saying so -- not an error state. Headers are sent lazily on the first delta, so a
turn that never streams falls through to the ordinary JSON body.

Accessibility replaces the plan's loose wording: one polite visually-hidden live
region updated at sentence and paragraph boundaries and on completion, with the
transcript itself no longer aria-live. Announcing every chunk was the failure
mode that wording invited.

Retry resends the original prompt as a new turn rather than replacing the failed
one. A turn that failed mid-stream may already have run tools, and replacing it
in place would hide that it ran twice -- exactly the fact a person needs once
writes exist. Only errors offer retry; round-cap and cancelled turns do not,
because the model ended those, not the transport.

ContextScope in the TypeScript types was 'site' | 'post-type' | 'selection',
which matched neither R6 nor the endpoint's enum. Reconciled to 'site' |
'general'.

Known gap, commented at the call site: on the first turn of a new conversation
the client does not yet know the server-minted conversation id, so Stop closes
the reader locally while that round finishes server-side. Later turns cancel
properly. Closing it needs the turn endpoint to publish the id early.

Advances R2, R6, R8, R9, R11, R19 (U5).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L
…the transcript

Closes a gap the transcript could not work around: the turn route returned only
tool invocation records, never the results, so no post list could reach the
client at all. Each tool_calls entry now carries the ability's own return value
on success, and null on a denial -- a refusal is not a result.

The passthrough is deliberately narrow. Nothing is re-fetched, joined, or
enriched, because re-fetching post data client-side would bypass the
execute-time permission filtering the whole read path rests on. A test asserts
the value the client receives is identical to direct ability execution for the
same user and input, so the response cannot leak more than the tool did. The
client then treats it as unknown and rebuilds each row from the fields the
search ability declares, so a future ability with a wider payload cannot push
unexpected fields into the table.

The table is chrome-free by construction rather than by configuration: composing
DataViews with an explicit Layout child means search, filters, view config and
pagination are never mounted, since the component renders its default UI only
when given no children. Per-field flags and controlled view state back that up.

The plan's instruction to copy the DataViews stylesheet in webpack was
unnecessary -- the copy plugin already emits it once to a fixed path outside the
entry map -- so the config is untouched and the page enqueues the existing file
conditionally, mirroring the request-logs page.

R14 is only half-met by this unit: rows link to the post, not to its editor. The
search ability returns no edit URL and nothing client-side may invent one, so
the edit action is gated on a row-supplied URL and is currently never rendered.
Completing it belongs on the ability that already owns the permission walk.

Advances R14 (U6).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L
Completes R14. The transcript table already had an edit action gated on a
row-supplied URL, but no ability emitted one, so the action never rendered and
rows linked only to the published post.

The URL doubles as the permission proof: it is present only when the current
user can edit the post, so a consumer does not re-derive the capability and
cannot construct an editor URL for a post it may only read. That also avoids
guessing wp-admin/post.php, which breaks on non-standard admin paths.

The explicit capability check is deliberate redundancy, and the docblock says
so: get_edit_post_link() already returns nothing for a user who cannot edit, and
a mutation run with the gate removed left the new tests green. The gate stays
because the emptiness of this field is a permission guarantee consumers rely on,
and that guarantee should not rest on the internals of a core function that is
free to change. The tests lock the contract rather than the mechanism.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L
…ally work

Two bugs found by running the workspace against a real provider rather than
against tests. Both were invisible because each degraded politely instead of
failing.

OptionEnum::FUNCTION_DECLARATIONS is not a constant. That call threw an
undefined-constant Error, a catch-everything swallowed it, and the screen
therefore always reported "no compatible model available" -- even though all
eleven Anthropic models advertise function calling. The enum extends
AbstractEnum, so the member is reached with isFunctionDeclarations().

The catch is now narrower in effect: an Error is re-thrown rather than
swallowed, so a programming mistake in that method surfaces instead of
presenting as a permanent, plausible-looking capability gap. An unreachable
provider still degrades quietly, which is the case the catch is for. Throwable
is still the caught type because the project's coding standard requires it over
Exception.

The streaming driver called GenerativeAiResultChunk::toText(), which does not
exist -- the method is getDeltaText(). The Throwable guard turned that into
"this host cannot stream", so every turn silently fell back to a buffered
request and the UI truthfully reported a condition that was not true. Verified
end to end afterwards: a turn now streams, and the fallback notice is gone.

Only the visible delta text is emitted; getReasoningDeltaText() carries the
model's thinking, which is not the reply.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L
…hinking signatures

Two defects found by running a two-turn conversation against the real provider.

The model was whichever candidate the registry happened to return first, which
on a live site resolved to the most capable and most expensive model available
-- selected by array position, not intent. Selection now consults an ordered
preference, filterable per site, and falls through cleanly when a preferred
model is absent.

That accidental choice also exposed the second defect. Always-on-thinking models
return thinking blocks, and Anthropic rejects a replayed thinking block whose
signature is missing, so the second turn of any conversation failed with
"messages.1.content.0.thinking.signature: Field required". The signature arrives
as its own signature_delta after the block's thinking text, so the mapper now
records which block indices opened as thinking and carries the signature onto
the part it belongs to, matched by index rather than arrival order.

The storage round trip was never at fault: the SDK's MessagePart already carries
a thought signature and Message::fromArray()/toArray() preserve it. The loss was
ours, in the mapper, on the streaming path only -- which is why no test saw it
and why it took a live second turn to surface.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L
The model can propose drafts; only a person's confirmation writes them. The
proposal is persisted server-side and the confirmation renders the stored
resolved values, never the model's summary of them -- that summary is
attacker-influenceable, which is the reason the requirement exists.

There is no registered write ability. The plan called for one kept off the MCP
surface, but that is an intention without a mechanism: a registered ability is
reachable by the MCP surface, the Abilities Explorer, and any third-party
caller, none of which have a confirm gate. Instead the writer is a plain class
reachable only from the proposal controller, so the confirm gate is a structural
property of the write path rather than a validation rule someone can route
around. The model's only reach is a propose ability that writes nothing, is
hidden from REST and MCP, and refuses outside an active turn context.

Ownership is bound, not merely capability-checked. Capability is not identity:
without this, a second user holding the same capability could execute another's
proposal by id, and the values they approved on screen would not be the ones
written. The store derives its key from the owner, compares the stored owner
again on read, and compares the conversation at execution -- three checks
independent of the capability re-check at write time. Verified by mutation:
with the key no longer user-scoped, the stored-owner comparison still refuses a
peer, so both layers hold on their own.

Set approval is bounded. Proposals cap at twenty items, matching the search
tool's row bound, and items start deselected so an item appended by injected
content must be chosen deliberately rather than ridden in on a batch approval.
A status the user cannot publish to fails the proposal instead of being
downgraded silently.

Partial failure is reported per item to both the person and the model, with no
auto-retry, and re-executing a confirmed proposal creates nothing. Every write
attempt is logged in the shape the tool-call rows already use, so writes and
reads join on one conversation.

Advances R15, R16, R17, R20 (U7).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L
An Options-menu action in the post editor opens the workspace with the current
post in scope. It carries identity only -- post ID, status, type, title -- and
never the body: the workspace reads content through the permission-checked tool
path, so there is one enforcement path and nothing trusts a client-supplied
body. A test asserts a marker in the post content never reaches the localized
data.

The seed is re-resolved server-side from the query argument and checked with
read_post, so the URL cannot hand someone a post they may not read; a refusal
renders an explanation rather than content.

A post title is author-controlled text, so it is flattened to single spaces and
clamped before it leaves PHP -- a title cannot smuggle a multi-line instruction
block -- and it reaches the model only inside a message the person has seen and
can edit. The composer is prefilled, never auto-sent. The residual surface is a
person who sends a prefilled prompt without reading it; that is bounded to one
clamped line with a human in the loop, and removing it entirely would mean not
naming the post at all, which the available tools make useless.

Mounted as a post-level Options-menu item rather than the block toolbar the plan
named. This is a navigation action for the whole post: a block toolbar entry
would either repeat on every block or be arbitrarily scoped to one block type
the way content resizing is. The repository already has post-level plugin
precedent.

The webpack entry is part of this change. That build declares every bundle
explicitly, and the asset loader bails silently on a missing asset file, so
without the entry the action would never appear and nothing would report why.

R5 is only half met and should not be read as complete: the handoff carries
identity, but no read-full-body tool is in the workspace allowlist, so the
assistant can find the post and see its excerpt while the body stays out of
reach. Widening the allowlist is a security-relevant scope decision and was
deliberately left alone.

Advances R5, R18 (U8).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BRSJEyQyYp3XTeum8fiq3L
whyisjake and others added 18 commits September 23, 2026 08:46
The workspace stopped hiding the admin menu in ec1f727, but six
descriptions still called it full-screen, including the experiment
description a site owner reads in Settings before enabling it, and a
known-limitation entry documenting a body class the code no longer
applies. Those are corrected, and the limitation entry is deleted rather
than reworded because the problem it describes no longer exists.

Correcting the copy surfaced a second, worse inaccuracy. The old Settings
string said the assistant "creates or updates content" -- an overclaim,
since no registered ability updates anything. The first rewrite replaced
it with "can only create drafts", which is an underclaim in the more
dangerous direction: `Propose_Drafts` accepts a per-item status of draft,
pending, private or publish, and `Draft_Writer` passes it to
`wp_insert_post()` behind a `publish_posts` check that exists precisely
because publishing is reachable. An approved proposal can create a
published post. Both strings now describe the guarantee that is actually
true -- nothing is written until a person approves a proposal showing the
exact values, post status included.

Also records that the draft confirmation deliberately diverges from its
design reference: the reference draws a plain list, the implementation
gives each item a checkbox starting unchecked. That difference is load
bearing rather than cosmetic, and the note cites the server-side contract
that carries it so a later reader does not "fix" the code toward the
drawing.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The ordered-item pattern matched the leading digits but captured only the
text after them, so every list rendered without a start attribute. That is
invisible in a contiguous list, where the browser counts on its own. But a
model writing a numbered lineup puts an excerpt between the items, and the
blank line ends the list — so five titles became five one-item lists, each
restarting at 1 and all displaying "1.".

Capture the opening ordinal and carry it to the ol, omitting it when it is
1 so an ordinary list gains no attribute.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The test proves model-authored markup is stripped even for an approver core
would skip its own kses pass for, and it guards that premise by asserting the
approver holds unfiltered_html. On multisite nobody below a super admin does,
so the guard failed and took all three multisite jobs down with it.

Granting the approver super admin for the whole class fixes this one test and
breaks two others: a super admin holds every capability regardless of role, so
the tests that revoke a capability and assert nothing is written stop revoking
anything. The grant is scoped to the single test that needs it and handed back
in a finally.

Skipping on multisite was the other option and a worse one. The case only
exists because of the user core treats differently, so the platform where that
user is rarer is not the platform to stop checking.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
…ssor

WordPress 7.1 lets wp_get_abilities() filter the registry. WordPress 7.0's
takes no parameters, and PHP discards extra arguments to userland functions
without a word, so asking 7.0 to filter returns everything. A tool policy
that trusted the call would be default-allow on the plugin's own minimum.

Tool_Policy answers two questions and nothing else: whether filtered
discovery is available here, and whether an ability has declared itself fit
for a conversational surface. Both fail closed. The declaration key stays
private and provisional while issue #354 settles its public shape.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
…loor

The workspace chose its tools from a hand-maintained list of three. An
ability can now earn a place by declaring itself fit for a conversational
surface, and the list stops being the only way in.

Default-deny alone would have emptied the surface: nothing declares yet, and
this does not retrofit the abilities that ship here. So the curated three are
a floor the policy adds to, never a fallback it replaces. A test pins that by
name, because it is the property a change to merge order breaks silently.

Two ways in are refused. Core's ability filters are site-wide and fire on
every call including ours, so a plugin hooking one could hand the model an
ability nobody declared; the query is a candidate set and every row is
re-checked here. And the candidates filter, which may skip the declaration,
may not skip the effect class -- otherwise a tool that writes reaches a model
that has been told it cannot write.

Admission requires readonly, not destructive, and not open-world, each
asserted explicitly. Core defaults all three to null, and an absent
open-world hint means the ability may reach outside the site, so silence is
refused rather than assumed. The curated floor is exempt: ai/propose-drafts
writes through the confirm gate and would otherwise strip itself out.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The system instruction promised two things it had no way to know: that the
model cannot write to the site, and that a proposal tool is there to call.
Both were literals. Once the surface can grow or be narrowed, a literal is a
claim the code cannot keep.

Each sentence is now earned. The proposal paragraph appears only when the
proposal tool was declared, so the model is never pointed at a tool it does
not have. The write denial appears only while nothing declared can change the
site on its own -- ai/propose-drafts does not count against it, because it
stores values a person has to approve rather than writing them.

The warning that tool results are untrusted stays unconditional. It does not
describe the surface, it describes what to do with anything that comes back,
and that is true whatever is admitted.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The policy decided the surface and nobody could see it. The Abilities
Explorer now says, per ability, whether the assistant holds it and why not
when it does not -- told apart properly, so an author who declared correctly
but writes is told that, and is not sent back to re-check a declaration that
was already right.

Six reasons, and the report enumerates the whole registry to produce them.
The admission query cannot: it only returns what already matched, so it can
never explain a non-match.

The owner can take any ability off the surface, including a curated one, and
switch the policy off entirely. Both persist as options that uninstall
already cleans by prefix. The mutation checks a nonce and manage_options,
because a control that reshapes what an assistant may call is worth a CSRF.

The removal filter is registered from the experiment bootstrap, not from a
constructor. Hooking on construction would have made the owner's removal
depend on something happening to build a policy first, so a candidate read
that did not would quietly serve a tool the owner took away. Two tests
proving removal works were relying on exactly that.

Descriptions are shown escaped. They are third-party text rendered into
wp-admin, and the point is that the owner reads the same words the model
reads -- not that wp-admin renders someone else's markup.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The experiment doc still said three abilities ship in an allowlist and left
it there. Three still ship, but they are a floor now, and the interesting
half is what an ability has to do to stand beside them: declare itself, and
say plainly that it only reads, adds nothing destructive, and stays inside
the site. Two opt-ins, which is worth stating, because an author who adds
the declaration and stops is not admitted.

Also records what a reader would otherwise have to find out by trying: the
key is private until #354 settles it, so nothing changes yet; WordPress 7.0
has no filtering to do this with and gets the floor; and the Explorer is
where an owner sees the surface and takes something off it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Review found the safety story was not true. The declaration key is a private
constant, and the docs said that meant no third-party ability could opt in
yet -- but private hides a name from callers, not from an author who reads
the source of an open-source plugin. Admission was live on merge.

So admission is now off by default, behind a temporary switch that goes away
when #354 settles the key's public shape. A declared ability shows in the
Explorer as eligible and waits there. The docs say this plainly instead of
claiming a barrier that was never one.

The owner's removals no longer ride on a filter the workspace bootstrap
installs. The Abilities Explorer is a separate experiment that can run while
the workspace does not, and in that configuration nothing registered the
filter -- so a removed ability was shown as held, beside a button offering to
return it. Exclusions apply inside the candidate build now, and a test pins
that without any bootstrap.

Reasons stop guessing. An ability removed by site code, dropped for its
effect class, or waiting on the gate each says so, instead of all three
arriving as "withheld by your capabilities" and sending the owner to look at
roles. A surface the owner emptied no longer reports that nothing is
registered.

The docblock claimed two independent opt-ins. They are two keys in one meta
array written by one hand. Both are self-attestation; the owner's controls
and the permission callback are the boundaries that hold.

One correction to what the last commit said: core discards a meta-mismatched
ability before wp_get_abilities_item_include fires, so that filter cannot
re-widen the query. wp_get_abilities_result can, and the re-verification loop
is what stops it. The test now proves the smuggle before proving the defense.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Tool_Selector said "admission decides what the model is told exists" twice, in
two halves of the same docblock, and named wp_get_abilities_item_include as the
filter a third party could re-widen the query with. Core discards a
meta-mismatched ability before that filter fires; wp_get_abilities_result is
the one that can inject, and the one the re-verification loop is for.

The rest is wording: inverted openers straightened out, a binary contrast
stated directly, and the em dashes thinned.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
A full tool description in every row made two rows fill the screen, so the
column meant to show what the assistant can reach was the reason the table
could not be read.

The state is a badge now, an icon beside a short word. The icon is decorative
and marked so; the word carries the meaning, because a column whose whole
purpose is being read should not depend on being seen.

The description moved into a details element, closed by default. It stays
there rather than moving to a tooltip or a title attribute: it is the exact
string handed to the model, the owner is the only person who can judge whether
it is honest, and judging it means reading and selecting long-form prose.

Two tests asserted on the old wording. They assert on the state class now, so
copy can change without a test pretending something broke, and a new one holds
the description in the cell -- collapsing it is the point, losing it is not.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The column asked "is this on the assistant", which made every other row a
negative and told the owner nothing. An ability can be exposed over REST, over
MCP, to the assistant, or any combination, and until now there was nowhere to
see those three answers together.

So the column is "Exposed in" and carries a badge per surface, read from the
same meta each consumer reads for itself. core/get-site-info turns out to be
REST only, ai/get-post-terms MCP only, and an ability can be eligible for the
assistant while exposed nowhere else.

That spread is the argument in issue #354. Three consumers invented three
flags, and this column is the first place the cost of that shows up as
something an owner looks at rather than something a developer reads about.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
… flag

WordPress 7.1 added meta.public as the general exposure flag, and a channel
resolves as meta[channel] ?? meta.public ?? the channel's own default. Core
applies that to show_in_rest at registration and writes the answer back, so
reading that key gives the resolved value.

Nothing applies it to mcp.public. Reading that key alone reported an ability
as absent from MCP when the general flag had put it there, so the column was
telling the owner an ability was less exposed than it is -- the wrong
direction for a screen whose job is showing reach.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
WordPress 7.1 added meta.public as the general exposure flag, and the
precedence every channel resolves by: its own key, then that flag, then the
channel's default. The workspace was reading a bespoke key of its own, which
is the fragmentation issue #354 is about.

It reads ai-workspace.public ?? public ?? false now. An ability marked public
is eligible here without naming this surface, and naming it overrides in
either direction.

Two things follow, neither of them cosmetic.

The discovery query lost its meta condition. Core matches meta exactly and
cannot express the fallback, so querying the channel key would have silently
skipped every ability eligible only through the general flag. Resolution moved
per item, and a test holds the query empty so nobody optimises it back.

Two exclusion reasons became one. Core writes meta.public onto every ability
at registration and validates it as a boolean, so "the author declared
nothing" is not a state that exists any more -- an ability with no opinion is
eligible-false by core's default. Reporting it as undeclared would send an
author to add a key already there with the value they meant.

Inheriting widens who is eligible, so the effect class is what keeps that from
meaning every public ability can be called by a model that has been talked
into it. Five cases cover that: not readonly, destructive, open-world, the
hint absent, and no annotations at all.

has_declaration() is gone. Since core seeds the key, its fallback branch could
never return false, and nothing outside tests called it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Inheriting meta.public cost a lock nobody was counting. core/get-user-info
ships public, read-only and not destructive, so the only thing keeping a
reader of people's personal data off the assistant was an absent open_world
hint -- and open_world => false would be a correct, tidy thing for core to
add. One good-faith cleanup upstream and a PII tool lands on a surface
reachable by instructions embedded in content someone else wrote.

The effect class cannot catch that. It asks whether an ability writes or
reaches outside the site, not whether handing it to a model is a bad idea. So
a short list answers the second question for the four core abilities where it
is plainly yes: user info, users, settings, environment.

WordPress registers those, not this plugin, so a list here is the only place
the decision can live. It is filterable, because a site with code access
deciding otherwise is different from a default deciding for them, and it is
checked before the curated floor so it holds on every route in.

A test registers a fixture that is declared, annotated impeccably, and on the
list, then asserts it never reaches the model -- so the refusal is provably
the list rather than something the fixture got wrong.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Fifteen rows said "Not public, and has not opted in to the assistant" under a
badge already reading "not the assistant", which buried the three rows where
the reason is the whole point: held back for reading personal data, refused on
effect class, or waiting on the admission gate.

The commonest reason is left unsaid now. The rest still print, because those
are the ones an owner or an author can act on.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
#1002 renamed core/read-users to core/users-query and kept the old name as a
deprecated alias. Both are registered, and the withheld list named only the
old one -- so the ability it exists to hold back was reachable under the name
core now prefers.

Not reachable in practice yet: core/users-query is not public and does not
assert an admissible effect class. But the point of the list is to hold when
those change, and it was not holding.

Both names are listed, because the alias is a real registration rather than a
redirect. A rename upstream is a hole here until someone adds the new name,
which the docblock now says.

PHPStan goes with it. The composer script pinned 1G, and this branch pushed
analysis past it once develop merged in, so the documented command crashed.
2G clears it. CI never used that script -- it runs phpstan with no limit at
all -- so this is the local command catching up to what CI already did.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
#1002 renamed core/read-content to core/content-query. Nine docblock and doc
references here still pointed at the old name, which now resolves to a
deprecated alias -- accurate enough to pass review and wrong enough to send
the next reader somewhere that will be removed.

The withheld list keeps naming both sides of the users rename, and the docs
now say why: the alias is a real registration that copies the replacement's
meta, so exposure changes reach both names at once and holding back only one
holds back nothing.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
@thelovekesh

thelovekesh commented Sep 23, 2026 •

Copy link
Copy Markdown
Member

Bridging the two needs Requests or WP_Http to return a lazy body stream. If that's the upstream change you had in mind, I'll open an issue proposing it before doing anything more here.

Hey @whyisjake, I’ve proposed adding streaming support to the Requests library in WordPress/Requests#1056, and there’s also an open PR for it.

whyisjake and others added 2 commits September 23, 2026 11:33
…xists

The Anthropic streaming model extends a class that ships in the
ai-provider-for-anthropic plugin. Requiring the file where that parent is
absent is a fatal error, and the autoloader requires it for any mention of
the class name. PHPStan 2.2.13 now resolves the docblock union naming this
class in Streaming_Turn_Driver's constructor eagerly, which autoloads the
file in an analysis environment without the provider plugin and aborts the
run with an internal error.

Guard the declaration on the parent's presence so the file declares nothing
on such a host. class_exists() on this class then answers false, which is
the answer every probe wanted.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
WordPress 7.1 introduced the `public` ability meta, seeds it to false on
every registration, and resolves show_in_rest from it. Two tests assert on
that seeding and fail on the 7.0.4 legs of the CI matrix, where the key does
not exist. Skip them there, the way the filtered-discovery tests already
skip on the same versions.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
## What?
Closes #1016. See #203.

Moves the whole Abilities Explorer screen (list, statistics, the
"Exposed in" column and its actions, detail view, test runner) off
`WP_List_Table`, vanilla JS and admin-ajax onto a `@wordpress/dataviews`
screen backed by four new `ai/v1` routes. Stacked on `feat/ai-workspace`
(#1004), because the "Exposed in" column and its actions exist only
there.

## Why?
The Explorer was the plugin's last admin screen on the old stack, and
its hand-written search, sort and filters kept needing repair (#883,
#641, #648, #344, #588). The "Exposed in" column answers the question an
owner most needs answered, which abilities the assistant can reach, but
it could not be filtered without more hand-written filtering. #203 asked
for a custom-column hook, which would have shipped an API against a
table about to be replaced.

## How?
**REST routes**
(`includes/Experiments/Abilities_Explorer/REST/Abilities_Controller.php`),
registered only while the experiment is on and never as abilities:

| Route | Method | Purpose |
|---|---|---|
| `ai/v1/abilities` | GET | every registered ability, including ones
without `show_in_rest`, plus the assistant policy state |
| `ai/v1/abilities/item?name=` | GET | one ability with schemas, raw
data and example input |
| `ai/v1/abilities/invoke` | POST | runs the ability through
`WP_Ability::execute()`, so its own permission check still applies |
| `ai/v1/abilities/surface` | POST | remove, restore, disable_policy or
enable_policy |

All four share one permission check: `manage_options`, cookie
authentication, a valid `wp_rest` nonce, and no application password.
The nonce is checked in the callback itself because core skips its own
check when an earlier authentication filter has already answered.
Ability names travel in the query string or body, never the path, so an
encoded `/` cannot 404 on Apache. Each list item is encode-checked on
its own, so one ability with unencodable data cannot blank the list.

**Screen** (`src/experiments/abilities-explorer/`): a React app mounted
like the AI Request Logs page. It routes between list, detail, runner
and not-found on the existing `action` and `ability` query args, with
browser Back and Forward. The provider filter keeps the #883 rule
(origin for Core, Plugin and Theme, exact label otherwise). "Exposed in"
is now filterable by assistant state. Surface changes wait for the
server, and responses carry a sequence number, tracked per row, so a
late refresh cannot undo a newer change.

**#203:** plugins that depend on `wp-hooks` can add read-only fields
with the `ai.abilitiesExplorer.fields` filter. Built-in field IDs win on
a collision, a filter that throws or returns a non-array leaves the
built-ins, and a saved view keeps a third-party column's ID while its
plugin is inactive. Documented in
`docs/experiments/abilities-explorer.md`.

**Behavior changes worth knowing:**
- Invoke now runs under REST, where `is_admin()` is false. The route
loads `wp-admin/includes/admin.php` first.
- The list is built from a REST request too, so an ability a plugin
registers only when `is_admin()` is true no longer appears.
- Errors show the ability's code, message and data instead of the
always-null `trace`.
- The admin-ajax action `ai_ability_explorer_invoke` and the
`Ability_Table` class are removed.
- Old list bookmarks with `s`, `orderby` or filter args land on the
unfiltered list.
- Back to List keeps the view's layout, columns, sort and page size.
Search and filters reset.
- Row actions (View, Test, and Remove from or Return to assistant) sit
under each ability's name and show on hover or keyboard focus, as in
`WP_List_Table`. They are always shown below 782px.

### Use of AI Tools
AI assistance: Yes
Tool(s): Claude Code
Model(s): Claude Opus 5.5
Used for: Planning, implementation of the REST controller, React screen,
tests and docs, and a multi-reviewer code review whose confirmed
findings were fixed in this branch.

<!-- JAKE: add your review attestation here in your own words before
opening. -->

## Testing Instructions
1. Enable Experiments and the Abilities Explorer experiment, then open
**Tools > Abilities Explorer**.
2. Filter by Provider, Category and "Exposed in", search, sort and page.
The statistics should not change with search or filters.
3. With the AI Workspace experiment on, remove an ability from the
assistant, return it, and turn the admission policy off and on. Each
change should show a notice.
4. Open an ability's detail view, copy a schema, then open **Test
Ability**. For `ai/content-classification`, `max_suggestions` 11 should
report "must be at most 10", and 5 should validate.
5. Use browser Back and reload on each view.

### Tests run locally
- `npm run build`, `npm run typecheck`, `npm run lint:js`, `npm run
lint:php`, `npm run lint:php:stan`: all pass.
- `npm run test:php`: 1988 tests, 40 skipped, 0 failures.
- Explorer e2e, 24 tests across four files, all passing:
  - `tests/e2e/specs/experiments/abilities-explorer-list.spec.js`:
    - shows the statistics and the table under a single heading
    - narrows the rows with the "Category" filter
    - matches "Plugin" by origin and "Acme" by its exact label
    - keeps the statistics unaffected by search
    - filters "Exposed in" down to what the assistant can reach
- restores layout, fields and sort from a saved view, but not search or
filters
    - loads a saved view with an unknown field ID and keeps the ID
  - `tests/e2e/specs/experiments/abilities-explorer-surface.spec.js`:
    - removes an ability and returns it, with a notice for each
- returning a withheld ability shows the withheld reason, not the
assistant badge
    - turns the admission policy off and on, updating every row
    - disables a pending change, so a second click sends nothing
    - does not let a refresh answered after a remove undo it
    - moves focus to the list when a change filters its row out
  - `tests/e2e/specs/experiments/abilities-explorer-runner.spec.js`:
    - shows each detail section and copies the JSON
    - validates deep-linked input against the schema
    - shows the success and error panels, and Clear hides them
    - never shows one runner the result of another ability in flight
  - `tests/e2e/specs/experiments/abilities-explorer-navigation.spec.js`:
    - keeps one heading per view, and Back and reload keep the view
- fetches an ability again when it is opened after returning to the list
- shows an extension column on first load and keeps it through
deactivation
    - keeps the built-in columns when the filter returns a non-array
    - does not let an extension replace a built-in field
    - is unreachable when the experiment is off
    - is unreachable when AI is globally off
- Full `npm run test:e2e`: running at the time this PR was opened; the
result will be added below.

<details><summary>PHPUnit tests added or moved in this branch
(67)</summary>

`Abilities_ExplorerTest` (2):
- `test_rest_routes_are_registered_when_enabled`
- `test_rest_routes_are_absent_when_disabled`

`Ability_HandlerTest` (5):
- `test_validate_input_accepts_a_type_list`
- `test_generate_example_input_returns_empty_for_empty_schema`
- `test_generate_example_input_uses_default_values`
- `test_generate_example_input_uses_example_values`
- `test_generate_example_input_generates_type_defaults`

`Admin_PageTest` (12):
- `test_load_hook_registers_help_tabs_and_assets`
- `test_assets_are_hooked_only_when_the_screen_loads`
- `test_enqueue_assets_enqueues_bundle_and_dataviews_fallback`
- `test_enqueue_assets_skips_dataviews_fallback_when_core_registers_it`
- `test_enqueue_assets_does_nothing_for_an_editor`
- `test_localized_settings_route_map_matches_controller_constants`
- `test_localized_settings_carry_reason_and_provider_labels`
- `test_render_page_works_for_admin`
- `test_render_page_ignores_action_for_server_render`
- `test_render_page_outputs_nothing_for_editor`
- `test_render_page_outputs_nothing_when_logged_out`
- `test_help_tabs_register_on_screen`

`Abilities_ControllerTest` (48):
- `test_list_includes_an_ability_not_shown_in_rest`
-
`test_list_keeps_origin_and_provider_apart_and_sends_the_category_label_unescaped`
- `test_list_decodes_a_category_label_escaped_by_a_filter`
- `test_list_survives_an_ability_with_unencodable_meta`
- `test_item_returns_schemas_and_example_input`
- `test_item_returns_404_for_an_unregistered_name`
- `test_item_returns_400_for_a_missing_name`
- `test_invoke_with_valid_input_returns_the_data`
- `test_invoke_with_input_failing_explorer_validation_returns_400`
-
`test_invoke_with_input_failing_core_validation_returns_the_ability_error`
- `test_invoke_accepts_a_property_typed_with_a_type_list`
- `test_invoke_without_an_input_schema_accepts_empty_and_null`
- `test_invoke_with_input_omitted_invokes_with_no_input`
- `test_invoke_decodes_scalar_input_on_the_server`
- `test_invoke_with_malformed_json_returns_400`
- `test_invoke_refuses_non_string_input`
- `test_invoke_returns_404_for_an_unregistered_ability`
- `test_invoke_reports_the_ability_permission_denial_as_its_outcome`
- `test_invoke_writes_no_request_log_row`
- `test_surface_remove_stores_the_exclusion_and_a_repeat_is_no_change`
- `test_surface_restore_clears_the_exclusion`
-
`test_surface_restore_of_a_withheld_ability_keeps_it_off_the_assistant`
- `test_surface_policy_switch_disables_and_enables_and_returns_the_list`
- `test_surface_refuses_an_unknown_change`
- `test_surface_remove_refuses_an_unregistered_or_missing_name`
- `test_surface_get_changes_nothing`
- `test_list_reports_only_known_origins`
- `test_list_carries_a_custom_provider_beside_the_known_origins`
- `test_list_lets_a_custom_provider_row_match_its_origin_and_its_label`
- `test_list_marks_a_declared_admitted_ability_as_on_the_assistant`
- `test_list_carries_the_reason_an_ability_is_off_the_assistant`
- `test_list_and_function_declaration_share_one_description_source`
-
`test_list_does_not_report_a_removed_ability_as_held_without_the_workspace_bootstrap`
- `test_list_reports_the_general_public_flag_on_both_channels`
- `test_list_reports_rest_and_mcp_exposure_independently`
- `test_surface_remove_takes_the_ability_off_the_model_declarations`
- `test_surface_restore_returns_the_ability_to_the_model_declarations`
- `test_surface_policy_switch_withdraws_and_returns_admitted_abilities`
- `test_editor_is_refused_on_every_route`
- `test_application_password_is_refused_on_every_route`
- `test_application_password_is_refused_even_alongside_the_cookie_flag`
-
`test_determine_current_user_without_a_cookie_is_refused_on_every_route`
- `test_cookie_administrator_without_a_nonce_is_refused_on_every_route`
-
`test_cookie_administrator_with_an_invalid_nonce_is_refused_on_every_route`
-
`test_cookie_administrator_with_a_nonce_for_another_action_is_refused_on_every_route`
- `test_cookie_administrator_with_the_nonce_as_a_parameter_is_accepted`
- `test_cookie_administrator_reaches_every_route`
- `test_no_route_is_registered_as_an_ability`
</details>

## Unapplied review findings
From the branch's code review, not blocking:
- [ ] P3 — `src/experiments/abilities-explorer/api.ts` — The response
sequencer and `validate.ts` have no JS unit tests, because the repo has
no JS unit test setup. They are covered only through e2e.
- [ ] P3 —
`includes/Experiments/Abilities_Explorer/REST/Abilities_Controller.php`
— The invoke success payload is not passed through the strict encode
check the list and item use, so an ability returning NaN or INF produces
a generic client error.
- [ ] P3 — `src/experiments/abilities-explorer/fields.tsx` —
`BUILT_IN_FIELD_IDS` is a hand-kept list. Deriving it from
`getBuiltInFields()` would stop a new built-in field from being
overridable.
- [ ] P3 — Response ordering uses one server's `microtime()`. On
multi-node hosting with clock skew, a stale list could briefly show a
row's old state until the next refresh.

## Changelog Entry
> Changed - Rebuilt the Abilities Explorer on DataViews with REST
routes, a filterable "Exposed in" column, and a JavaScript filter for
adding read-only table fields.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- wp-playground-preview:start -->
<a
href="https://playground.wordpress.net?blueprint-url=data:application/json,%7B%22preferredVersions%22%3A%7B%22wp%22%3A%22latest%22%7D%2C%22login%22%3Atrue%2C%22landingPage%22%3A%22%2Fwp-admin%2Fadmin.php%3Fpage%3Dai-wp-admin%22%2C%22steps%22%3A%5B%7B%22step%22%3A%22installPlugin%22%2C%22pluginData%22%3A%7B%22resource%22%3A%22url%22%2C%22url%22%3A%22https%3A%2F%2Fgithub.com%2FWordPress%2Fai%2Freleases%2Fdownload%2Fci-artifacts%2Fpr-1074-88aa03bda577ce567570cbd90eb3be378851aad0.zip%22%7D%7D%5D%7D"
target="_blank" rel="noopener noreferrer">
<img
src="https://raw.githubusercontent.com/adamziel/playground-preview/refs/heads/trunk/assets/playground-preview-button.svg"
alt="Open WordPress Playground Preview" width="220" height="57" />
</a>
<!-- wp-playground-preview:end -->

---------

Co-authored-by: Claude Opus 5.5 <[email protected]>
whyisjake added a commit that referenced this pull request Sep 30, 2026
Every change on this branch already landed on feat/ai-workspace when #1004
was rebased: 52 commits as identical patches and 3 re-applied with
conflict resolutions. Take feat/ai-workspace's tree as-is so the PR has
no conflicts and merging it changes no code.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
whyisjake and others added 2 commits September 30, 2026 14:29
Resolve the one conflict: develop edited
tests/e2e/specs/experiments/abilities-explorer.spec.js for #985, while
this branch had already split that spec into four files, so it stays
deleted.

#985 also removed the global "Enable AI" toggle and its e2e helpers.
Carry that into this branch's own tests the way #985 did for the other
experiments:

- Drop AI_WorkspaceTest::test_global_toggle_off_prevents_registration and
  the "is unreachable when AI is globally off" e2e test, since there is
  no global switch left to turn off.
- Remove the enableExperiments() and disableExperiments() calls from the
  AI Workspace e2e specs, keeping the per-experiment enable and disable.
- Stop setting the no longer registered wpai_features_enabled option in
  AI_WorkspaceTest, the Abilities Explorer controller test and the
  Explorer's e2e setup.
- Name the enableExperiment() helper in the docs instead of the removed
  enableExperiments().

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…lient streams

Provider streaming depended on WordPress/php-ai-client#255, which is not
merged or released, so this branch vendored its streaming classes into the
SDK overlay. Take that dependency out of #1004 so the workspace can land
on released code: every round now goes through the buffered path the
workspace already used for providers other than Anthropic, and replies
arrive whole.

Removed here and restored by the follow-up PR, which reverts this commit:
- the Anthropic streaming driver, model, stream mapper and fopen() transport
  under includes/Experiments/AI_Workspace/Streaming/, with their tests;
- the vendored streaming classes and HTTP DTOs, and the SDK overlay's
  streaming feature;
- the default streaming driver in Prompt_Model_Client, the "cannot stream"
  notice in the transcript, and the streaming docs, hooks and changelog
  lines.

The SSE responder and the optional Stream_Driver_Interface injection stay,
so restoring streaming only adds a driver back.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@whyisjake

Copy link
Copy Markdown
Member Author

Proposal: merge the AI Workspace (#1004)

Ask: can someone with write access review #1004? I'd like to land it in the AI plugin during the 7.2 cycle.

What it is: a full-screen chat at Tools > AI Workspace (admins only) that reads site content through the Abilities API, under the current user's capabilities. It never writes on its own: it proposes drafts, and a person approves them. It ships as an opt-in experiment.

Dependencies — all released
✅ Core AI client (wp_ai_client_prompt), WordPress 7.0+
✅ Ability filtering and the public ability flag, WordPress 7.1. On 7.0 it falls back to a curated set of three abilities.

Streaming is now a follow-up
I split token-by-token streaming out into #1083 (draft), so #1004 no longer depends on the unreleased php-ai-client#255. Replies arrive whole for now. #1083 restores streaming once #255 ships. Streaming in the PHP AI Client is already on the 7.2 roadmap.

Housekeeping: #1004 is up to date with develop and has no conflicts. The Abilities Explorer DataViews work (#1074) is merged into it, and #1014 is closed because its changes are already in #1004.

Open questions I'd like input on:

  1. What should the confirmation modal select by default?
  2. How should we close the full-screen CSS gap?
  3. Where should the editor handoff live in the Options menu?

A follow-up turn with Claude failed with "messages.1.content.0.thinking.
signature: Field required". The workspace replays the whole conversation
each round, including the model's earlier thinking. ai-provider-for-
anthropic 1.0.4 drops a thinking block's signature when it reads a reply
and sends it back unsigned, which Anthropic rejects. Until #1004 took
streaming out, the streaming model restored or omitted that block; the
buffered path did not.

Drop thought parts that carry no signature before each round. Anthropic
accepts an omitted thinking block, including between tool calls, and the
Google and OpenAI connectors sign their thought parts, so theirs pass
through unchanged.

Verified against the live API on a dev site: a two-turn conversation and
a tool-calling turn both complete, and the second turn returns the 400
again with the filter disabled.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Chat experiment: Integration outside the editor and outside single-task AI use

4 participants