Skip to content

knowledge(performance): grouped query (Count + ColumnFilter = HAVING) for distinct values and duplicates - #215

Open
Michael Dieringer (MichaelDieringer) wants to merge 1 commit into
microsoft:mainfrom
Curabis:community-contribution/r4-grouped-query-having
Open

Michael Dieringer (MichaelDieringer) wants to merge 1 commit into
microsoft:mainfrom
Curabis:community-contribution/r4-grouped-query-having

Conversation

@MichaelDieringer

Copy link
Copy Markdown
Contributor

Summary

  • performance/use-grouped-query-for-distinct-values-and-duplicates. AL's Record has no SELECT DISTINCT and no GROUP BY ... HAVING. So table-wide duplicate detection is often written as a loop that filters a second record variable on each row's value and calls Count(). That costs one extra SQL statement per looped row, and an unfiltered loop runs as many statements as the table has rows. A query object does it in one statement. An aggregate Method groups the dataset by the other columns, and a ColumnFilter on a Method = Count column is applied as HAVING. ColumnFilter = NameCount = filter(> 1) returns only the duplicate groups.

    • Complements aggregate-before-persisting-intermediate-results, which covers grouped totals.
    • Links query/setfilter-overwrites-query-columnfilter, because a runtime filter on the count column replaces its ColumnFilter.
  • Verified:

    • Filtering in query objects: a filter on a column with a totals method corresponds to HAVING, any other filter to WHERE. A filter row is not included in the dataset.
    • Aggregating data in query objects: implicit grouping, Count takes only a name, and the "distinct values" section.
    • Base Application uses this exact shape in query 762 "Acc. Sched. Line Desc. Count" (Finance/FinancialReports/AccSchedLineDescCount.Query.al), which CheckDuplicateAccScheduleLineDescription calls (AccSchedChartManagement.Codeunit.al lines 390-398). ColmLaytColmHeaderCount.Query.al and Inventory/Analysis/AnalysisLineDescCount.Query.al follow the same shape. BCApps links are pinned to 837ef80.
  • Exclusions, to avoid false positives:

    • A single uniqueness check (OnValidate, before Insert).
    • An outer loop bounded to one parent document.
    • A lookup against a different subset than the looped rows.
    • A loop that acts on each hit.
    • A filtered or counted record that is temporary.

    ServiceContractHeader.Table.al lines 2692-2705 is cited as a legitimate per-row lookup that must not be flagged.

  • Wiring: added to al-performance-review with a worklist cue for the nested filter-and-count loop, and registered in the performance review-fixtures.json override.

Test plan

  • Samples checked with AL compiler 30.0 against Base Application 28.4 symbols
  • validate_frontmatter.py: 0 errors (2 warnings, both in files this PR doesn't touch)
  • Test-KnowledgeIndex.ps1, Test-SkillIndex.ps1, Test-ReviewContract.ps1, Test-KnowledgeRetrieval.ps1
  • Test-ReviewFixtures.ps1: 230 cases, including the -PrepareDirectory deterministic ranking check

🤖 Generated with Claude Code

… for distinct values and duplicates

Adds use-grouped-query-for-distinct-values-and-duplicates with good/bad
samples, a worklist cue in al-performance-review, and registration in the
performance review-fixtures override.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant