Skip to content

feat(data-handling): standardized table export with a schema sidecar - #594

Open
breimanntools wants to merge 1 commit into
masterfrom
feat/33-export-formats
Open

breimanntools wants to merge 1 commit into
masterfrom
feat/33-export-formats

Conversation

@breimanntools

Copy link
Copy Markdown
Owner

The package had zero to_csv / to_parquet / export surface (verified before building). aa.to_table writes a table plus a JSON sidecar that says what the columns mean.

dict_metadata = aa.to_table(df_feat, "signature.csv", name="df_feat", random_state=42)

Two artifacts: the table (.csv / .tsv, no row index) and signature.meta.json, which is also returned.

Assembly, not authorship — the sidecar is built from what already ships: column meanings and dtypes from DICT_DF_SCHEMAS (plus DICT_DF_FEAT for the optional post-hoc df_feat columns), and the environment/version/seed block from get_provenance. Nothing new is invented, which was the point of the issue.

name selects which shipped schema documents the frame — any key of DICT_DF_SCHEMAS (df_feat, df_seq, df_parts, X, df_pred, …) — and is validated: a frame missing that schema's required columns raises a ValueError naming them. sep=None derives the separator from the extension; a mismatched pair (.tsv with ,) is rejected.

Verification

78 test cases (27 + 19 across the house classes), including a round-trip losslessness assertion on the real DOM_GSEC feature output — read_back.equals(df_feat) is True with dtypes preserved. 819 tests pass. pyright 0 errors (it first flagged 6 from =None defaults on local helpers — the trap from earlier today — fixed with honest signatures). Docstring checker 0 defects, drift 0. The generated example page has no section titles, so the CRITICAL baseline cannot rise.

Three things it deliberately did NOT build

  1. No read_table companion. Codex's plan claimed plain read_csv could not round-trip losslessly; that claim was tested and is false for these frames. A reader would have been a second public symbol, a second test file and a second notebook to buy nothing. The sidecar carries dtype_pandas so an exact load-back is a documented one-liner.
  2. No parquet, and no pyproject.toml change. Requirement 3 of the issue is not implemented. The recommendation is to close it as won't-do: the sidecar already solves the dtype/meaning problem parquet's embedded schema would solve, CSV is what dashboards and external pipelines ingest, and df.to_parquet() is one call away. Your call.
  3. No workflow edit to add the notebook to the nbmake subset — CONFIRM-FIRST.

Two smaller decisions for you

  • name defaults to "df_feat". Convenient for the primary output, but a mislabelled export is then caught only by the required-column check. Could be made required.
  • The row index is not written, so a filtered df_feat only compares equal after reset_index(drop=True) (documented and tested). An index= parameter would cover that if it matters.

Refs #33.

🤖 Generated with Claude Code

The package had no table-export surface at all: every downstream consumer
re-implemented `to_csv` and hand-wrote the metadata needed to read a `df_feat`
back months later, so exported results were easy to mis-read.

`to_table(df, file_path, name="df_feat", sep=None, random_state=None)` writes an
output table to CSV/TSV plus a JSON sidecar named after it (`<stem>.meta.json`),
and returns that record. The sidecar is assembled from what already ships rather
than inventing a metadata format: column meanings, dtypes, allowed values and
ranges come from the data schemas (with the simple feature contract covering the
optional and post-hoc `df_feat` columns), and the package/environment record
comes from `get_provenance`. Every exported column carries a non-empty
description; columns outside the named contract are flagged in
`columns_undocumented` rather than silently described as if they were part of it.

The row index is not written, so the CSV has no unnamed first column and
round-trips losslessly through plain `pd.read_csv`. Delimited text only: no new
dependency and no parquet path.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
@codecov

codecov Bot commented Sep 18, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.16981% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 95.39%. Comparing base (3c11c90) to head (db6eaaf).
⚠️ Report is 10 commits behind head on master.

Files with missing lines Patch % Lines
aaanalysis/data_handling/_backend/export_table.py 95.08% 2 Missing and 1 partial ⚠️
Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff            @@
##           master     #594    +/-   ##
========================================
  Coverage   95.39%   95.39%            
========================================
  Files         222      224     +2     
  Lines       23387    23493   +106     
  Branches     4073     4088    +15     
========================================
+ Hits        22309    22412   +103     
- Misses        631      633     +2     
- Partials      447      448     +1     
Files with missing lines Coverage Δ
aaanalysis/__init__.py 92.85% <ø> (ø)
aaanalysis/_constants.py 100.00% <100.00%> (ø)
aaanalysis/data_handling/__init__.py 100.00% <100.00%> (ø)
aaanalysis/data_handling/_to_table.py 100.00% <100.00%> (ø)
aaanalysis/data_handling/_backend/export_table.py 95.08% <95.08%> (ø)

... and 1 file with indirect coverage changes

Components Coverage Δ
cpp_core 95.97% <ø> (ø)
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant