knowledge(performance): grouped query (Count + ColumnFilter = HAVING) for distinct values and duplicates - #215
Open
Michael Dieringer (MichaelDieringer) wants to merge 1 commit into
Conversation
… for distinct values and duplicates Adds use-grouped-query-for-distinct-values-and-duplicates with good/bad samples, a worklist cue in al-performance-review, and registration in the performance review-fixtures override. Co-Authored-By: Claude Opus 5.5 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
performance/use-grouped-query-for-distinct-values-and-duplicates. AL'sRecordhas noSELECT DISTINCTand noGROUP BY ... HAVING. So table-wide duplicate detection is often written as a loop that filters a second record variable on each row's value and callsCount(). That costs one extra SQL statement per looped row, and an unfiltered loop runs as many statements as the table has rows. A query object does it in one statement. An aggregateMethodgroups the dataset by the other columns, and aColumnFilteron aMethod = Countcolumn is applied asHAVING.ColumnFilter = NameCount = filter(> 1)returns only the duplicate groups.aggregate-before-persisting-intermediate-results, which covers grouped totals.query/setfilter-overwrites-query-columnfilter, because a runtime filter on the count column replaces itsColumnFilter.Verified:
HAVING, any other filter toWHERE. A filter row is not included in the dataset.Counttakes only a name, and the "distinct values" section.Finance/FinancialReports/AccSchedLineDescCount.Query.al), whichCheckDuplicateAccScheduleLineDescriptioncalls (AccSchedChartManagement.Codeunit.allines 390-398).ColmLaytColmHeaderCount.Query.alandInventory/Analysis/AnalysisLineDescCount.Query.alfollow the same shape. BCApps links are pinned to 837ef80.Exclusions, to avoid false positives:
OnValidate, beforeInsert).ServiceContractHeader.Table.allines 2692-2705 is cited as a legitimate per-row lookup that must not be flagged.Wiring: added to
al-performance-reviewwith a worklist cue for the nested filter-and-count loop, and registered in the performancereview-fixtures.jsonoverride.Test plan
validate_frontmatter.py: 0 errors (2 warnings, both in files this PR doesn't touch)Test-KnowledgeIndex.ps1,Test-SkillIndex.ps1,Test-ReviewContract.ps1,Test-KnowledgeRetrieval.ps1Test-ReviewFixtures.ps1: 230 cases, including the-PrepareDirectorydeterministic ranking check🤖 Generated with Claude Code