Skip to content

Issue #67 - Add an event parsing mode to toCSV - #68

Merged
NassimBtk merged 7 commits into
mainfrom
feature/issue-67-add-a-streaming-mode-to-tocsv
Oct 8, 2026
Merged

NassimBtk merged 7 commits into
mainfrom
feature/issue-67-add-a-streaming-mode-to-tocsv

Conversation

@NassimBtk

@NassimBtk NassimBtk commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Closes #67

What changed

This PR adds an opt-in event parsing mode with the same API and the same output as the default mode.

  • new JFlat(String source, boolean eventParsing) and new JFlat(Reader reader, boolean eventParsing). The existing constructors keep the default mode.
  • parse() only checks the document with the event parser of javax.json.stream, instead of loading it as a tree. When the document is not an object or an array, or is invalid, it runs the default parse, so the exceptions are the same.
  • toCSV() reads the document again with the event parser. It flattens the value of the first element of the entry key (items in /items/status/conditions) one array element at a time, or as a single unit when the value is not a non-empty array.
    • Each unit is flattened under its real path (/items[3]), with the same markers and number formatting as navigateTree().
    • The rows are built with the existing entry expansion and row code, now extracted to expandEntries() and appendRows(). Wildcards, ../, ., {object}/{array}/NULL, empty and null lists, and case-insensitive keys therefore behave as in the default mode.
    • Rows keep the array index order.
  • Pruning. Nodes that neither the entry key nor any property can reach are skipped. There is no pruning when the entry key contains a wildcard or a property contains non-ASCII characters. A subtree whose key contains /, [, ] or non-ASCII characters is always kept.
  • Fallback to the default mode when the event parsing cannot give the exact result:
    • the document is an array;
    • the entry key is /, or its first element is * or contains an index;
    • the first element of the entry key is present more than once in the root object, whatever the case;
    • a root key contains / or [;
    • a property refers to a node above the element (../kind from /items);
    • an object of an element has a duplicate key.
  • getFlatTree() builds the map of the whole document, as before.
  • What this mode does not do. The input is not streamed: the document is still kept in memory as a whole, and a Reader is read entirely by parse(). The output is not streamed either: the CSV is returned as a whole. The mode is named after the way the document is read, so that streaming the input or the output can be added later under its own name.
  • Version: 1.2.00-SNAPSHOT (new feature; main is at 1.1.01-SNAPSHOT).
  • Docs: an "Event Parsing Mode" section in src/site/markdown/index.md.

Tests

  • Every existing test runs in both modes (@ParameterizedTest on eventParsing).
  • csvList covers a Kubernetes-like list: row order, entry key below the elements, markers and numbers, a property above the element, case-insensitive keys.
  • eventParsingEdgeCases checks, in both modes and with and without removeNodes, the documents that must give the default result:
    • empty keys, root keys with / or [, the list present twice, duplicate keys;
    • ../ within and above the element, wildcards below the value and below an element;
    • empty, scalar and null values, the escaped wildcard;
    • case-colliding and non-ASCII keys;
    • number formats, trailing content, a number out of range.
  • eventParsingMatchesDefault is a differential test with a fixed seed. It covers 4,000 generated documents × 6 calls, with random entry keys, properties, separators and removeNodes. It compares the CSV of two consecutive calls and the flat tree, or the exception class and message, and asserts that most calls are actually read by events (not fallbacks).
  • edgeCases also checks that a Reader that fails gives an IOException and is closed, in both modes.
  • Mutation check. I disabled one at a time each fallback (duplicate keys, empty path segments, list present twice, ../ above the element, root key with /), the wildcard rule of the pruning, the number check and the parent array path. Each of these 8 mutations makes the tests fail. Disabling the pruning entirely, which does not change the output, does not.

mvn verify: 16 tests, prettier check OK.

Parity on real data

There are 196 cases, each run with and without removeNodes, so 392 runs. All 392 give identical output with jflat 1.1.00, this branch in the default mode, and this branch in event parsing mode. 332 of the runs are actually read by events; the others fall back by design (entry key /, ../kind, invalid documents).

The cases are:

  • the 26 json2Csv definitions of the MetricsHub OpenShift connector, on captures of a real cluster and on the integration-test fixtures;
  • 17 Kubernetes lists of the same cluster, each with 5 entry key/property sets;
  • edge cases.

Benchmark

Kubernetes DeploymentList (compact JSON, managedFields included), entry key /items, properties /metadata/name /metadata/namespace /metadata/uid /spec/replicas /status/availableReplicas. Temurin 21.0.2, G1, one JVM at a time.

How to read the table:

  • Minimum heap is the smallest -Xmx that succeeds, found by bisection with child JVMs.
  • Load only is the minimum heap needed to hold the document String alone.
  • Times are medians of 9 runs (2.8 MiB) or 3 runs.
  • Both modes give the same output, checked by row count and hash.
List Payload Default: min heap Event parsing: min heap Load only Default: time Event parsing: time Default: allocated Event parsing: allocated
239 Deployments (real cluster) 2.8 MiB 57 MiB 9 MiB 9 MiB 136 ms 16 ms 58.8 MiB 4.9 MiB
5,000 Deployments 59.5 MiB 1,171 MiB 63 MiB 61 MiB 2.66 s 0.31 s 1,254 MiB 91.5 MiB
10,000 Deployments 119 MiB 2,343 MiB 123 MiB 121 MiB 5.85 s 0.61 s 2,485 MiB 183 MiB

With the nested entry key /items/status/conditions (properties ../../metadata/name /type /status /reason /lastTransitionTime):

List Default: time Event parsing: time Event parsing: min heap
2.8 MiB 149 ms 18 ms 9 MiB
59.5 MiB 6.06 s 0.31 s 65 MiB
119 MiB 29.7 s 0.66 s 127 MiB

In event parsing mode the memory is the document String plus a few MiB, so the String itself is now the main cost. In MetricsHub, reading the HTTP response briefly needs 3 to 4 times the payload on top of that. That is handled separately in simple-http-java (MetricsHub/simple-http-java#60).

Notes

  • Behavior kept as is. A * that follows an element of an array passes the entry through instead of expanding the element's children (case 2 of the wildcard handling looks for the last [ anywhere in the path). The event parsing mode reproduces this exactly.
  • Engine side. Since the output is identical, the plan is to use new JFlat(body, true) in ClientsExecutor.executeJson2Csv.
  • Branch name. The branch still says "streaming"; renaming it would mean a new pull request.

🤖 Generated with Claude Code

new JFlat(source, true) only checks the document in parse(), and toCSV()
reads it again and flattens the value of the first element of the entry
key one array element at a time, skipping the nodes that neither the
entry key nor the properties can reach. The memory used is the document
plus one element, and the result is the same as in the default mode.
The documents that cannot be streamed exactly are processed as in the
default mode.

Every existing test runs in both modes, and a differential test
compares both modes on generated documents. Version 1.2.00-SNAPSHOT.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-08T15:08:48.789566Z 16fab24 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 233b5cffe0

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/main/java/org/metricshub/jflat/JFlat.java Outdated
@NassimBtk

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Breezy!

Reviewed commit: 03719a377a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread pom.xml Outdated
NassimBtk and others added 3 commits October 8, 2026 15:16
"Streaming" could be read as streaming the input or the CSV output, which
this mode does not do: the document is still kept in memory and the CSV
is returned as a whole. The mode reads the document with the event
parser of javax.json.stream instead of loading it as a tree, so the flag
is now eventParsing (JFlat(String, boolean pEventParsing)), and the
documentation says what is still kept in memory.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@NassimBtk NassimBtk changed the title Issue #67 - Add a streaming mode to toCSV Issue #67 - Add an event parsing mode to toCSV Oct 8, 2026
@NassimBtk

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Delightful!

Reviewed commit: 16fab24060

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@NassimBtk
NassimBtk merged commit 921b336 into main Oct 8, 2026
4 checks passed
@NassimBtk
NassimBtk deleted the feature/issue-67-add-a-streaming-mode-to-tocsv branch October 8, 2026 17:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add an event parsing mode to toCSV to bound memory to one array element

1 participant