Add classification capability - #300
the-hercules wants to merge 4 commits into
Conversation
|
The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message. To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## trunk #300 +/- ##
============================================
+ Coverage 86.58% 87.78% +1.20%
- Complexity 1383 1503 +120
============================================
Files 69 75 +6
Lines 4449 4872 +423
============================================
+ Hits 3852 4277 +425
+ Misses 597 595 -2
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
Is "classification" the right/best naming option here (versus "decisioning" or something else)? |
What?
Adds a non-generative classification capability, for models that answer typed questions about a given state instead of generating content. Each answer is a probability, a chosen option, or a position on a scale, optionally with a confidence and a per-option or per-level probability distribution.
It follows the embeddings precedent: a capability enum value, a model interface, a builder, a result DTO, and lifecycle events.
Closes #296.
Why?
Today the only way to use a classification or judgement model is
TEXT_GENERATIONwith a JSON response schema. That asks a generative model to print a number, and it throws away the calibrated confidence a classification model returns. It also means provider packages that support these models, such as Fueled/ai-provider-for-ollama#103, have to ship their own DTOs and entry points outside the client.docs/REQUIREMENTS.mdalready lists classification as a non-generative feature the client must support.How?
CapabilityEnum::CLASSIFICATION.ClassificationModelInterface::classifyResult(array $state, array $questions): ClassificationResult.ClassificationQuestion, typed byClassificationQuestionTypeEnum:binary: a yes/no statement, with no criteria.choice: criteria map each option key to a description, with at least two options. Keys must be non-integer strings, because PHP turns integer-like keys into integers, which can then encode as a JSON list.score: criteria are an ordered list of level descriptions, lowest to highest, with at least two levels.ClassificationAnswer:getProbability(),getChoice()andgetScore()throw if called on the wrong type.getConfidence()andgetProbabilities()are always available.ClassificationResultimplementsResultInterface.ClassificationBuilder, created viaAiClient::classify($state):ModelResolutionTrait, asPromptBuilderdoes.RuntimeException.BeforeClassifyEventandAfterClassifyEventgo through the client's PSR-14 event dispatcher. The after-event fires only for valid results.docs/ARCHITECTURE.mdupdates.Intentionally not included:
noulname, Ollama's limits (26 options, 64 questions, 64 KB body), the scorelegend, and optional true/false descriptions on binary questions. Each could be added to core later without a breaking change.Naming is open for maintainer feedback:
CLASSIFICATIONvsDECISION,AiClient::classify(), and the event names. See the shape proposal in #296.Changelog Entry
Use of AI Tools
AI assistance: Yes
Tool(s): Claude Code, [tool used for the initial draft]
Model(s): Claude Opus 5.5, [model used for the initial draft]
Used for: An initial implementation was drafted with AI. I then reviewed it with Claude Code against #296 and Fueled/ai-provider-for-ollama#103, which led to reworking the question and answer shape, adding answer validation and lifecycle events, and expanding the tests. All code was reviewed and tested by me.
🤖 Generated with Claude Code