Repository navigation
Issue #67 - Add an event parsing mode to toCSV - #68
Conversation
new JFlat(source, true) only checks the document in parse(), and toCSV() reads it again and flattens the value of the first element of the entry key one array element at a time, skipping the nodes that neither the entry key nor the properties can reach. The memory used is the document plus one element, and the result is the same as in the default mode. The documents that cannot be streamed exactly are processed as in the default mode. Every existing test runs in both modes, and a differential test compares both modes on generated documents. Version 1.2.00-SNAPSHOT. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 233b5cffe0
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Co-Authored-By: Claude Opus 5.5 <[email protected]>
|
@codex review |
|
Codex Review: Didn't find any major issues. Breezy! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
"Streaming" could be read as streaming the input or the CSV output, which this mode does not do: the document is still kept in memory and the CSV is returned as a whole. The mode reads the document with the event parser of javax.json.stream instead of loading it as a tree, so the flag is now eventParsing (JFlat(String, boolean pEventParsing)), and the documentation says what is still kept in memory. Co-Authored-By: Claude Opus 5.5 <[email protected]>
|
@codex review |
|
Codex Review: Didn't find any major issues. Delightful! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
Closes #67
What changed
This PR adds an opt-in event parsing mode with the same API and the same output as the default mode.
new JFlat(String source, boolean eventParsing)andnew JFlat(Reader reader, boolean eventParsing). The existing constructors keep the default mode.parse()only checks the document with the event parser ofjavax.json.stream, instead of loading it as a tree. When the document is not an object or an array, or is invalid, it runs the default parse, so the exceptions are the same.toCSV()reads the document again with the event parser. It flattens the value of the first element of the entry key (itemsin/items/status/conditions) one array element at a time, or as a single unit when the value is not a non-empty array./items[3]), with the same markers and number formatting asnavigateTree().expandEntries()andappendRows(). Wildcards,../,.,{object}/{array}/NULL, empty and null lists, and case-insensitive keys therefore behave as in the default mode./,[,]or non-ASCII characters is always kept./, or its first element is*or contains an index;/or[;../kindfrom/items);getFlatTree()builds the map of the whole document, as before.Readeris read entirely byparse(). The output is not streamed either: the CSV is returned as a whole. The mode is named after the way the document is read, so that streaming the input or the output can be added later under its own name.1.2.00-SNAPSHOT(new feature;mainis at1.1.01-SNAPSHOT).src/site/markdown/index.md.Tests
@ParameterizedTestoneventParsing).csvListcovers a Kubernetes-like list: row order, entry key below the elements, markers and numbers, a property above the element, case-insensitive keys.eventParsingEdgeCaseschecks, in both modes and with and withoutremoveNodes, the documents that must give the default result:/or[, the list present twice, duplicate keys;../within and above the element, wildcards below the value and below an element;eventParsingMatchesDefaultis a differential test with a fixed seed. It covers 4,000 generated documents × 6 calls, with random entry keys, properties, separators andremoveNodes. It compares the CSV of two consecutive calls and the flat tree, or the exception class and message, and asserts that most calls are actually read by events (not fallbacks).edgeCasesalso checks that aReaderthat fails gives anIOExceptionand is closed, in both modes.../above the element, root key with/), the wildcard rule of the pruning, the number check and the parent array path. Each of these 8 mutations makes the tests fail. Disabling the pruning entirely, which does not change the output, does not.mvn verify: 16 tests, prettier check OK.Parity on real data
There are 196 cases, each run with and without
removeNodes, so 392 runs. All 392 give identical output with jflat 1.1.00, this branch in the default mode, and this branch in event parsing mode. 332 of the runs are actually read by events; the others fall back by design (entry key/,../kind, invalid documents).The cases are:
json2Csvdefinitions of the MetricsHub OpenShift connector, on captures of a real cluster and on the integration-test fixtures;Benchmark
Kubernetes
DeploymentList(compact JSON,managedFieldsincluded), entry key/items, properties/metadata/name /metadata/namespace /metadata/uid /spec/replicas /status/availableReplicas. Temurin 21.0.2, G1, one JVM at a time.How to read the table:
-Xmxthat succeeds, found by bisection with child JVMs.With the nested entry key
/items/status/conditions(properties../../metadata/name /type /status /reason /lastTransitionTime):In event parsing mode the memory is the document String plus a few MiB, so the String itself is now the main cost. In MetricsHub, reading the HTTP response briefly needs 3 to 4 times the payload on top of that. That is handled separately in
simple-http-java(MetricsHub/simple-http-java#60).Notes
*that follows an element of an array passes the entry through instead of expanding the element's children (case 2 of the wildcard handling looks for the last[anywhere in the path). The event parsing mode reproduces this exactly.new JFlat(body, true)inClientsExecutor.executeJson2Csv.🤖 Generated with Claude Code