Skip to content

Make the library and the command line one loop: collect.resolve, join on read, jeval status - #7

Merged
rlaope merged 4 commits into
mainfrom
feat/runtime-labels-and-status
Sep 26, 2026
Merged

rlaope merged 4 commits into
mainfrom
feat/runtime-labels-and-status

Conversation

@rlaope

@rlaope rlaope commented Sep 26, 2026

Copy link
Copy Markdown
Owner

Summary

  • collect.resolve(source_key=..., question=..., answer=...) records the answer a human settled on, into labels.jsonl next to the records. Like track, it never raises.
  • Every command that reads records joins those answers in memory by (source_key, question). An existing label is kept. An answer the question cannot produce is refused. A later answer wins, but silver never replaces a human answer. Nothing is written back to disk.
  • jeval status shows what has been collected per model and question, and how many more gold labels the verdict needs.
  • docs/library.md is the library reference. docs/instrumenting-a-service.md is rebuilt around track/resolve/status, with every output captured from the new examples/service-quickstart.py. Its old threshold output predated the currency format and could not be reproduced.

Test plan

  • 25 new tests in tests/test_runtime_labels.py, including the independent review's findings, each reproduced before the fix:
    • resolve never raises on hostile inputs
    • silver does not beat gold
    • yes/no normalisation matches harvest
    • label --apply does not persist joined answers
    • status handles mixed timezones and score-only logs
  • Every output quoted in the two docs was re-run from a fresh directory and matches line for line; the only exception is the run-time timestamp, which the docs mark as such.
  • docs/library.md is now covered by the documented-features guard.
  • uv run pytest, ruff, mypy clean.

collect.resolve(source_key=..., question=..., answer=...) appends the answer a human settled on to
labels.jsonl beside the records, and never raises. Every command that reads records joins those
answers in memory by (source_key, question): an existing label is kept, an answer the question could
not have produced is refused, a later answer to the same case wins, and the outcome of every answer is
printed on stderr. jeval status shows what has been collected per model and question and how many more
gold labels the verdict needs. Together they make the library and the command line one loop with no
ingest step between them.

Signed-off-by: rlaope <[email protected]>
…reproducible run

docs/library.md is the reference for track, record and resolve: arguments, environment variables,
counters, promises and limits. The walkthrough is rewritten around track, resolve and status, with
every output captured from examples/service-quickstart.py; its earlier threshold output predated the
currency format and could not be reproduced. The README gains a two-ways-to-use-it section and a
jeval status row.

Signed-off-by: rlaope <[email protected]>
…els off disk

Review findings, each reproduced and pinned by a test:
- resolve raised on an object whose __str__ raises, a non-datetime ts or a non-path path; its whole
  body is now guarded, a directory is refused as a records path, and fields over 512 characters are
  refused. record() guards its path resolution too.
- a later silver answer in labels.jsonl replaced an earlier human one; a silver answer now never
  wins over human_review or human_override, and superseded lines are counted.
- yes/no answers such as True or false were refused; they are validated through the schema first,
  exactly as harvest_labels does.
- jeval label --apply wrote the joined answers into records.jsonl, freezing them against later
  corrections; it now reads records unjoined.
- jeval status crashed on mixed naive and aware timestamps and printed local times with a Z; it
  reads every stamp in UTC, refuses an empty records file, and does not ask score-only logs for
  labels the verdict cannot use.
- docs/library.md joins the documented-features guard.

Signed-off-by: rlaope <[email protected]>
A silver answer never replaces a human one; later means appended later; label --apply writes no
joined answer; JEVAL_COLLECT=<file> puts answers beside that file, but the command line joins only
<root>/.jeval/labels.jsonl.

Signed-off-by: rlaope <[email protected]>
@rlaope
rlaope merged commit 3a35331 into main Sep 26, 2026
5 checks passed
@rlaope
rlaope deleted the feat/runtime-labels-and-status branch September 26, 2026 01:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant