Skip to content

Source-level LilyPond API and datasets for curated corpora; release 0.4.0 - #12

Merged
matteospanio merged 9 commits into
masterfrom
dev
Sep 28, 2026
Merged

matteospanio merged 9 commits into
masterfrom
dev

Conversation

@matteospanio

Copy link
Copy Markdown
Member

LilyPond source you can inspect and edit, and datasets for curated corpora: Epics K and L, released together as 0.4.0. See docs/changelog.md for the list of changes, docs/devlog.md for the notes and the new docs/python-api.md for the API.

What changes

  • \version:
    • lytk.LilyPondVersion compares numerically and accepts what LilyPond 2.24 accepts;
    • lilypond_version, set_lilypond_version and strip_lilypond_version work on the parse tree;
    • Score.lilypond_version, and version= on the LilyPond writers;
    • a new invalid-version error.
  • Tokens: lytk.tokenize and strip_comments. Embedded Scheme is tokenized as Scheme, and LilyPond inside it as LilyPond. Nothing but whitespace is lost, on the whole corpus board.
  • Statistics: lytk.info(score) (the library home of lytk info --json, computed in Rust) and lytk.source_stats(text).
  • Includes:
    • found on the tree, anywhere in a line and never in comments or strings;
    • flatten_string flattens text;
    • every LilyPond reader and check_lilypond follow includes with include_paths=, with diagnostics mapped back into the caller's text;
    • flatten no longer refuses several \header blocks, no longer drops earlier \language lines, and keeps LilyPond's own includes.
  • Movements: from_lilypond_music_movements, from_lilypond_movements_string, from_abc_tunes_string.
  • Datasets:
    • RecordsDataset.from_jsonl / from_records (ids, records, the records' own splits);
    • on_error="skip"|"warn" with dataset.errors;
    • movements="all";
    • reading options passed to the readers;
    • a content-keyed, atomic cache that Subset uses;
    • ids, return_ids on the torch and TensorFlow adapters;
    • group-aware ratio splits.
  • Python API reference: docs/python-api.md, generated from the stubs; tests fail when it is stale or a public name has no stub.

Checks

  • 1,158 Rust + 262 Python tests locally (280 in CI's Python 3.12 job with torch and TensorFlow).
  • LilyPond corpus board: 1 valid file with an error, 0 refused, no text lost by the tokens.
  • All CI jobs green on each commit.

🤖 Generated with Claude Code

matteospanio and others added 9 commits September 28, 2026 11:46
lytk.LilyPondVersion compares numerically and accepts what LilyPond 2.24
accepts (x.y.z with a free fourth part, x.y for an even y).
lilypond_version(text) reads the first \version statement from the tree;
set_lilypond_version and strip_lilypond_version edit the statements.
Score/MusicDocument.lilypond_version carry the source's version, and
to_lilypond/to_lilypond_music take version=. An invalid \version is the
new invalid-version error, as LilyPond's lexer rejects it.

The corpus boards leave out LilyPond's tests of its own errors
(expect-error = ##t) and report how many of them lytk catches (2 of 13).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
lytk.tokenize(text) gives the tokens of LilyPond text from the parse tree
(kind, text, character span, line, column): comment, string, scheme,
command, symbol, number, fraction, punctuation, and error for text the
grammar cannot tokenize. Nothing but whitespace is lost, checked on the
fixtures, broken input and the whole corpus board. lytk.strip_comments
removes LilyPond comments as LilyPond skips them.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
lytk.info(score) is the library home of `lytk info --json` (which now
calls it), computed in Rust: parts and notes as before, plus voices, bars,
length in quarter notes, lyric syllables, chord symbols and grace notes.
lytk.source_stats(text) counts bytes, lines, tokens, comments, Scheme
expressions and error tokens. lytk.from_lilypond_music_movements gives
every movement of a file as a MusicDocument.

The devlog's test counts for K1 and K3 were two too high; corrected.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
… request

flatten finds includes on the parse tree (anywhere in a line, never in a
comment or string; a one-line %{ %} no longer hides the rest of the file),
keeps every \language (keeping only the last changed the music before it),
accepts several \header blocks (LilyPond merges them; MultipleHeaders is
gone), and keeps includes of LilyPond's own files (list generated from its
ly/ directory). lytk.flatten_string flattens text.

Every LilyPond reader and check_lilypond take include_paths=: includes are
followed, and each diagnostic is mapped back into the caller's text (one
in an included file at its \include, naming the file). Without it,
includes are still not followed; SECURITY.md says what following them
reads.

Epic K is complete.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
lytk.tokenize descends into embedded Scheme, which the grammar already
parses: `scheme` is now the # or $ that opens an expression, followed by
its symbols, numbers, strings, comments, brackets and quotes, plus the new
boolean, character and keyword kinds; Token.scheme tells Scheme tokens from
LilyPond ones, and LilyPond inside #{ #} is LilyPond again.
strip_comments strips Scheme's comments too. Nothing but whitespace is
lost, on the fixtures and the whole corpus board.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The string twin of from_lilypond_movements, with the same language,
strict and include_paths keywords.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ds, API reference)

- RecordsDataset.from_jsonl / from_records: records holding music as
  text, keeping their ids and records; with split_field, split() returns
  the records' own splits instead of re-shuffling them.
- on_error="raise"|"skip"|"warn" (iteration skips and records in
  dataset.errors; indexing raises), movements="all", and language,
  include_paths, strict and MIDI quantize passed to the readers.
- The cache is keyed by content, reading options, lytk version and
  normalized converter arguments, written atomically; Subset uses its
  parent's cache.
- ids on every dataset, return_ids on the torch and TensorFlow datasets
  and data loaders, split(groups=...) for group-aware ratio splits.
- docs/python-api.md generated from the stubs by scripts/python_api.py;
  tests fail when it is stale or a public name has no stub.
- lytk.from_abc_tunes_string for ABC records.

Epic L ships in 0.4.0 with Epic K, as the owner chose.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
set_lilypond_version and the writers' version= check a string version
and write it as given ("2.24" stays "2.24"); a LilyPondVersion is written
as str() gives it. Rust: set_lilypond_version takes the version text.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Epics K and L (source-level LilyPond API and datasets for curated
corpora), with Scheme tokens and from_lilypond_movements_string: version
0.4.0 in Cargo.toml (and Cargo.lock); the changelog's [Unreleased] becomes
[0.4.0] - 2026-09-28; roadmap and development log updated.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@matteospanio
matteospanio merged commit e6d4369 into master Sep 28, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant