Skip to content

Add classification capability - #300

Draft
the-hercules wants to merge 4 commits into
WordPress:trunkfrom
the-hercules:add/296-classification-capability
Draft

the-hercules wants to merge 4 commits into
WordPress:trunkfrom
the-hercules:add/296-classification-capability

Conversation

@the-hercules

Copy link
Copy Markdown

What?

Adds a non-generative classification capability, for models that answer typed questions about a given state instead of generating content. Each answer is a probability, a chosen option, or a position on a scale, optionally with a confidence and a per-option or per-level probability distribution.

It follows the embeddings precedent: a capability enum value, a model interface, a builder, a result DTO, and lifecycle events.

Closes #296.

Why?

Today the only way to use a classification or judgement model is TEXT_GENERATION with a JSON response schema. That asks a generative model to print a number, and it throws away the calibrated confidence a classification model returns. It also means provider packages that support these models, such as Fueled/ai-provider-for-ollama#103, have to ship their own DTOs and entry points outside the client.

docs/REQUIREMENTS.md already lists classification as a non-generative feature the client must support.

How?

  • CapabilityEnum::CLASSIFICATION.
  • ClassificationModelInterface::classifyResult(array $state, array $questions): ClassificationResult.
  • ClassificationQuestion, typed by ClassificationQuestionTypeEnum:
    • binary: a yes/no statement, with no criteria.
    • choice: criteria map each option key to a description, with at least two options. Keys must be non-integer strings, because PHP turns integer-like keys into integers, which can then encode as a JSON list.
    • score: criteria are an ordered list of level descriptions, lowest to highest, with at least two levels.
  • ClassificationAnswer:
    • Records its question type.
    • Typed getters getProbability(), getChoice() and getScore() throw if called on the wrong type. getConfidence() and getProbabilities() are always available.
    • A score is a number on the scale, from 0 to levels − 1, and may fall between levels.
    • NaN and out-of-range values are rejected.
  • ClassificationResult implements ResultInterface.
  • ClassificationBuilder, created via AiClient::classify($state):
    • Picks a model through ModelResolutionTrait, as PromptBuilder does.
    • Validates state and question keys.
    • Checks that every question was answered and that each answer fits its question (type, option, scale, probability keys). Otherwise it throws a RuntimeException.
  • BeforeClassifyEvent and AfterClassifyEvent go through the client's PSR-14 event dispatcher. The after-event fires only for valid results.
  • README and docs/ARCHITECTURE.md updates.

Intentionally not included:

  • No provider implementation; that belongs in provider packages.
  • Details specific to Jev or Ollama stay in providers: the noul name, Ollama's limits (26 options, 64 questions, 64 KB body), the score legend, and optional true/false descriptions on binary questions. Each could be added to core later without a breaking change.
  • No static one-call API yet, only the fluent builder.

Naming is open for maintainer feedback: CLASSIFICATION vs DECISION, AiClient::classify(), and the event names. See the shape proposal in #296.

Changelog Entry

Added - Classification capability for models that answer typed binary, choice, and score questions with probabilities and confidence, via AiClient::classify().

Use of AI Tools

AI assistance: Yes
Tool(s): Claude Code, [tool used for the initial draft]
Model(s): Claude Opus 5.5, [model used for the initial draft]
Used for: An initial implementation was drafted with AI. I then reviewed it with Claude Code against #296 and Fueled/ai-provider-for-ollama#103, which led to reworking the question and answer shape, adding answer validation and lifecycle events, and expanding the tests. All code was reviewed and tested by me.

🤖 Generated with Claude Code

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the props-bot label.

If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message.

Co-authored-by: the-hercules <[email protected]>
Co-authored-by: juanlentino <[email protected]>

To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook.

@codecov

codecov Bot commented Oct 2, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 87.78%. Comparing base (77a5995) to head (7ba63a8).

Additional details and impacted files
@@             Coverage Diff              @@
##              trunk     #300      +/-   ##
============================================
+ Coverage     86.58%   87.78%   +1.20%     
- Complexity     1383     1503     +120     
============================================
  Files            69       75       +6     
  Lines          4449     4872     +423     
============================================
+ Hits           3852     4277     +425     
+ Misses          597      595       -2     
Flag Coverage Δ
unit 87.78% <100.00%> (+1.20%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@the-hercules
the-hercules marked this pull request as draft October 2, 2026 10:07
@jeffpaul

jeffpaul commented Oct 3, 2026

Copy link
Copy Markdown
Member

Is "classification" the right/best naming option here (versus "decisioning" or something else)?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Non-generative capability for typed judgement models: classification and scoring with probabilities and confidence

2 participants