feat(data-handling): standardized table export with a schema sidecar - #594
Open
breimanntools wants to merge 1 commit into
Open
breimanntools wants to merge 1 commit into
breimanntools wants to merge 1 commit into
Conversation
The package had no table-export surface at all: every downstream consumer re-implemented `to_csv` and hand-wrote the metadata needed to read a `df_feat` back months later, so exported results were easy to mis-read. `to_table(df, file_path, name="df_feat", sep=None, random_state=None)` writes an output table to CSV/TSV plus a JSON sidecar named after it (`<stem>.meta.json`), and returns that record. The sidecar is assembled from what already ships rather than inventing a metadata format: column meanings, dtypes, allowed values and ranges come from the data schemas (with the simple feature contract covering the optional and post-hoc `df_feat` columns), and the package/environment record comes from `get_provenance`. Every exported column carries a non-empty description; columns outside the named contract are flagged in `columns_undocumented` rather than silently described as if they were part of it. The row index is not written, so the CSV has no unnamed first column and round-trips losslessly through plain `pd.read_csv`. Delimited text only: no new dependency and no parquet path. Co-Authored-By: Claude Fable 5.1 <[email protected]>
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #594 +/- ##
========================================
Coverage 95.39% 95.39%
========================================
Files 222 224 +2
Lines 23387 23493 +106
Branches 4073 4088 +15
========================================
+ Hits 22309 22412 +103
- Misses 631 633 +2
- Partials 447 448 +1
... and 1 file with indirect coverage changes
🚀 New features to boost your workflow:
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The package had zero
to_csv/to_parquet/exportsurface (verified before building).aa.to_tablewrites a table plus a JSON sidecar that says what the columns mean.Two artifacts: the table (
.csv/.tsv, no row index) andsignature.meta.json, which is also returned.Assembly, not authorship — the sidecar is built from what already ships: column meanings and dtypes from
DICT_DF_SCHEMAS(plusDICT_DF_FEATfor the optional post-hocdf_featcolumns), and the environment/version/seed block fromget_provenance. Nothing new is invented, which was the point of the issue.nameselects which shipped schema documents the frame — any key ofDICT_DF_SCHEMAS(df_feat,df_seq,df_parts,X,df_pred, …) — and is validated: a frame missing that schema's required columns raises aValueErrornaming them.sep=Nonederives the separator from the extension; a mismatched pair (.tsvwith,) is rejected.Verification
78 test cases (27 + 19 across the house classes), including a round-trip losslessness assertion on the real
DOM_GSECfeature output —read_back.equals(df_feat)isTruewith dtypes preserved. 819 tests pass. pyright 0 errors (it first flagged 6 from=Nonedefaults on local helpers — the trap from earlier today — fixed with honest signatures). Docstring checker 0 defects, drift 0. The generated example page has no section titles, so the CRITICAL baseline cannot rise.Three things it deliberately did NOT build
read_tablecompanion. Codex's plan claimed plainread_csvcould not round-trip losslessly; that claim was tested and is false for these frames. A reader would have been a second public symbol, a second test file and a second notebook to buy nothing. The sidecar carriesdtype_pandasso an exact load-back is a documented one-liner.pyproject.tomlchange. Requirement 3 of the issue is not implemented. The recommendation is to close it as won't-do: the sidecar already solves the dtype/meaning problem parquet's embedded schema would solve, CSV is what dashboards and external pipelines ingest, anddf.to_parquet()is one call away. Your call.Two smaller decisions for you
namedefaults to"df_feat". Convenient for the primary output, but a mislabelled export is then caught only by the required-column check. Could be made required.df_featonly compares equal afterreset_index(drop=True)(documented and tested). Anindex=parameter would cover that if it matters.Refs #33.
🤖 Generated with Claude Code