Skip to content

Use release-selected eager FlatBuffers for production AST parsing - #267

Open
lm-sousa wants to merge 78 commits into
clava-optimizationsfrom
ast-flatbuffers
Open

lm-sousa wants to merge 78 commits into
clava-optimizationsfrom
ast-flatbuffers

Conversation

@lm-sousa

@lm-sousa lm-sousa commented Oct 3, 2026

Copy link
Copy Markdown
Member

The AST reader must match the exact dumper release before eager FlatBuffers can replace Text. This switches the production parser to eager FlatBuffers, builds strict consumer bindings from the schema selected by clang-dumper-release.tag, verifies the pinned compiler/runtime and schema hashes, and rejects malformed or incompatible streams without Text fallback.

Testing uses v18.1.8_5-rc1. The lazy prototype and historical measurements remain on lazy-flatbuffers-experiment. The branch also fixes release resource lookup, bounds metadata caches, preserves assembly source syntax and documents field/node generation and compatible releases. Generated-output drift fails CI.

Stacked on #258 (clava-optimizations), with dumper PR #56 and dependency PR #37. No changes to lara-framework were needed for this task.

Validation is being completed against a separate RC checkout containing only Clava, lara-framework and specs-java-libs, with no sibling dumper. Native RC corpus validation passes all 1,062 baseline-clean LLVM inputs. Focused malformed-input, schema, cache, assembly and reference roundtrip tests pass. Final full parser tests, consumer corpus proofs and measured memory/runtime comparisons will be attached before handoff.

The broader corpus also exposes existing Text generator limitations in template partial specializations and lambda/temporary expressions. They are recorded separately from wire regressions; the isolated alias/reference/auto fixture validates the changed type representations.

This PR is for review and RC testing. Stable publication and master merging require separate approval.

Add a mapped hybrid reader, eager reference linking, field memoization and real Clava construction/codegen benchmarks. Preserve six-fixture fidelity checks, cache counters, memory probes, measurements and the recommendation to prefer eager integration with selective laziness.
…elds

Generate typed Java bindings from the companion schema, map bounded blocks, and resolve all graph references before dropping construction indexes. Keep memoized scalar fields in existing Clava stores and retain immutable backing files. Add full-suite cache experiments and cross-format fidelity, mapping and lifetime checks.
Archive 27 baseline-equivalent runs, fidelity and memory checks, and excluded race/quota trials. Isolate experiment scratch files and stop matrices on unexpected failures. Document that the generated schema works but does not improve suite time or retained heap with blanket laziness.
Syntax validation now sends one native-only syntax-check option before compiler arguments. Add focused coverage for its argument placement.
Generate structural descriptors from official schema reflection and validate bounded offsets, vtables, vectors, strings, unions, scalar domains and required fields. Bound recursion and verification work, and reject references outside a single block. Validation: 15 isolated JUnit tests pass; templates, NAS LU and fidelity native streams verify across 191 blocks. Consumer generation and framing integrate this verifier in the following cutover change.
Exercise schema mismatch, header ordering, duplicate headers, missing and incorrect End records, empty blocks, trailing records and truncated framing. Verify mapped files can be deleted after both successful and failed reads. Validation: 25 structural and envelope tests pass in isolated JUnit runs on JDK 17 and JDK 26 against the current eager reader sources.
Reflect exact supplied Clava classes and their inherited public DataKeys into deterministic JSON with declaring constants and generic value types. Classify references, enums, optional values and lists without schema annotations or a naming alias map. Validation: Java 17 compilation succeeds; all 112 payload classes expose 2,186 keys and repeated inventory output is byte-identical.
Cover negative 64-bit IDs that would narrow to a valid 32-bit null sentinel, invalid zero and negative references, and scope isolation for translation-unit node IDs. Validation: 27 structural, envelope and ID tests pass against the current eager reader sources on JDK 17.
Verify schema-specific cache directories, reject invalid hashes and false CCACHE_DISABLE values, and check compressed cache invocation settings.
Cover exact release tag resolution, missing local manifests, canonical schema hashing, malformed archives and incompatible compiler versions.
Preserve experimental implementations on lazy-flatbuffers-experiment. Add fixed suite runners, structural corpus checks, batch consumer comparisons, cross-TU and round-trip probes, and source-identified repeated memory comparisons.
Keep nullable-node metadata beside the Java DataKey and reflect it into the consumer inventory, preserving VariableArrayType size expressions without producer annotations or binding aliases.
Resolve the exact selected schema and pinned compiler/runtime without sibling checkouts. Document schema evolution and add generated-source drift checks, including ARM compiler bootstrap and archive verification tests.
Find revision manifests stored directly in a frozen distribution and hash local native tools alongside the parser JAR.
Apply the same collection attempts and waiting period to every runtime, including historical controls that retain an AST. This avoids extra control-side GCs changing the peak measurement.
Run eager FlatBuffers, historical Text, and Protobuf runtimes sequentially against the same filtered Clava-JS suite. Record source and runtime identities, test counts, cache state, and the shared Vitest wall-time boundary. Remove ccache and its aliases from the runner PATH while setting CCACHE_DISABLE=true.
Use the shared filtered PATH for all transports and inspect live parser file descriptors alongside mappings, AST collection and temporary folders. Require every eager cycle to release those resources.
…anup

Add a strict selected-corpus reparse mode with preserved compiler options, source hashes and stable generated bytes. Check mappings and open parser files across all previous parse iterations, and capture frozen runtime provenance from its own directory.
Run all memory controls with explicit G1 and honored explicit GC, scrub ambient JVM option injection from measured processes, and record the effective Java policy. Allow historical Text/Protobuf local dumper builds without release manifests when their exact executable hashes and source checkout snapshots are captured; keep eager release validation strict.
Pin the Text control's schema generator to its own source checkout and FlatBuffers SDK, and report their hashes. Historical Text/Protobuf preflights use the release tag and local tool hash rather than requiring the modern selected-release.json; retain strict resolved RC metadata for eager.
Seed the parser's exact default empty local_options.xml during preflight, before classpath fingerprints, and require the same file hash for Text, Protobuf, and eager. This captures the runtime's first-parse initialization without excluding mutable classpath content.
Set java.io.tmpdir in the forked Gradle test worker after resolving the run root. This lets the parser consume the preverified release cache staged under that observation and lets the harness verify its manifest and native tool directly before and after the test.
Prestage the selected v18.1.8_5-rc1 resource assets under each isolated XDG cache and record matching manifest, canonical schema, host executable, and extracted-includes hashes before and after observations. Query ccache's exact schema/tool namespace so cold and warm counters describe the production cache used by Clava.
Reuse one preverified release cache for every parse in a memory observation, and verify its manifest, native tool, schema and extracted includes before and after the JVM run. Keep per-parse AST cleanup and collection checks unchanged.
Pass the source-identified published cache and schema assets into the eager bypass observation so resource setup stays outside Vitest timing and the runner verifies the selected RC inputs.
Compare each Vitest source path relative to its Clava-JS workspace so current and staged historical overlays share stable test identities.
Match the shipped Windows headers, which alias high_resolution_clock to steady_clock. Keep Linux expectations unchanged and reuse the existing libc++ golden.
Resolve release provenance for the strict cleanup runtime instead of recognizing one hardcoded RC tag. Keep manifest, native, schema and header verification active for subsequent candidates.
def prepare_empty_local_options(checkout: Path, classpath: dict[str, Any]) -> dict[str, Any]:
"""Seed the runtime's default writable-JAR options file before hashing the classpath."""
main_classes = (checkout / "ClangAstParser" / "build" / "classes" / "java" / "main").resolve()
classpath_paths = {Path(path).resolve() for path in classpath["test_classpath"]}
Document the required release asset and resource cache inputs so the published examples pass the provenance preflight.
Require the testing RC2 release tag and exact Linux x64 CI artifact hash before accepting eager parser timing observations.
Retain partial-specialization parameters, init-capture names, direct-list syntax, and explicit cast types when printing eager ASTs. Preserve nested template parameter spelling and stable constructor member initializer parentheses. Add parser round-trip coverage and refresh the generated wire inventory for the current schema.
Use the testing release with partial-specialization parameter and lambda init-capture metadata. Preserve the developer local release override in the worktree.
Emit the sizeof operand envelope once and remove the inline separator when a comment has no declaration code. Cover explicit nested grouping and comments attached to empty declarations through a real parser round trip.
Print array qualifiers on their element types while preserving pointer qualifiers and existing element address spaces. Verify generated C++ compiles and remains stable after reparsing.
Read local tool paths by line prefix and tokenize ccache statistics before validating numbers, avoiding overlapping regex backtracking. The validation harness suite passes all 16 tests.
Give the first CUDA test thirty seconds to download and assemble its built-in headers. RC2 Linux CI passed the subsequent CUDA tests but timed out this initial setup at the default five-second limit.
Print capture initialization styles and pack placement from producer metadata, omit implicit captures from source lists, and reject misaligned capture vectors during eager reading. Validate source-stable roundtrips and rejection/file cleanup for each malformed vector.
Verify source/output hashes and exact compiler flags before checking each generated translation unit with LLVM 18. Record paired acceptance, baseline generator failures, commands, and diagnostics without waiving fidelity failures. The fixed RC2 corpus audit completed with no valid-Text-to-invalid-eager case.
Keep out-of-line template headers and specialization qualification, resolve lexical parameter names, and print bare constructor/destructor names. Preserve dependent braces and structured member-pointer precedence and suffixes. Validate grouped parameter metadata and stable roundtrips for partial templates, renamed parameters, member arrays, noexcept references, and dependent construction.
Expect bare constructor/destructor names, typed member-function pointer formatting, and one sizeof delimiter pair. Preserve source inputs and golden indentation; all five affected parser tests pass.
Consume the testing release with lambda, member pointer and template metadata. Preserve the user local dumper override while selecting RC3 in the committed release tag.
Keep CUDA and OpenMP outside the historical comparison workload while requiring the full integration suite to be checked separately.
"compiler": command[0],
"compiler_version": compiler_versions[command[0]],
"command": command,
"command_shell_quoted": shlex.join(command),
Comment on lines +176 to +177
completed = subprocess.run(command, capture_output=True, text=True,
check=False, timeout=timeout_seconds)
Comment on lines +151 to +152
completed = subprocess.run([compiler, "--version"], capture_output=True,
text=True, check=False, timeout=10)

def read_json(path: Path) -> dict[str, Any]:
try:
value = json.loads(path.read_text())
raise SystemExit(f"runtime {label} has no rows")

command = [str(part) for part in raw_command]
work_root = Path(command[-1]).resolve()
Render class arguments from each enclosing template parameter group, including nested primary and partial-specialization scopes. Keep enclosing headers before the member template header. Verify renamed, packed, non-type and mixed nested parameters through stable parse-generate-reparse tests.
Pin full corpus and audit hashes, preserve exact input options and record every excluded case before strict round-trip validation. Select all eager syntax-pass cases without manual exclusions or waivers.
Unwrap parenthesized pointees for cv/ref/noexcept suffix discovery, preserve method ref qualifiers, and globally qualify nonlocal member-pointer classes. Cover namespace shadowing, local classes, and template specializations through stable source round trips.
@sonarqubecloud

sonarqubecloud Bot commented Oct 4, 2026

Copy link
Copy Markdown

Quality Gate Failed Quality Gate failed

Failed conditions
E Security Rating on New Code (required ≥ A)

See analysis details on SonarQube Cloud

Catch issues before they fail your Quality Gate with our IDE extension SonarQube for IDE

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants