Make the library and the command line one loop: collect.resolve, join on read, jeval status - #7
Merged
Merged
Conversation
collect.resolve(source_key=..., question=..., answer=...) appends the answer a human settled on to labels.jsonl beside the records, and never raises. Every command that reads records joins those answers in memory by (source_key, question): an existing label is kept, an answer the question could not have produced is refused, a later answer to the same case wins, and the outcome of every answer is printed on stderr. jeval status shows what has been collected per model and question and how many more gold labels the verdict needs. Together they make the library and the command line one loop with no ingest step between them. Signed-off-by: rlaope <[email protected]>
…reproducible run docs/library.md is the reference for track, record and resolve: arguments, environment variables, counters, promises and limits. The walkthrough is rewritten around track, resolve and status, with every output captured from examples/service-quickstart.py; its earlier threshold output predated the currency format and could not be reproduced. The README gains a two-ways-to-use-it section and a jeval status row. Signed-off-by: rlaope <[email protected]>
…els off disk Review findings, each reproduced and pinned by a test: - resolve raised on an object whose __str__ raises, a non-datetime ts or a non-path path; its whole body is now guarded, a directory is refused as a records path, and fields over 512 characters are refused. record() guards its path resolution too. - a later silver answer in labels.jsonl replaced an earlier human one; a silver answer now never wins over human_review or human_override, and superseded lines are counted. - yes/no answers such as True or false were refused; they are validated through the schema first, exactly as harvest_labels does. - jeval label --apply wrote the joined answers into records.jsonl, freezing them against later corrections; it now reads records unjoined. - jeval status crashed on mixed naive and aware timestamps and printed local times with a Z; it reads every stamp in UTC, refuses an empty records file, and does not ask score-only logs for labels the verdict cannot use. - docs/library.md joins the documented-features guard. Signed-off-by: rlaope <[email protected]>
A silver answer never replaces a human one; later means appended later; label --apply writes no joined answer; JEVAL_COLLECT=<file> puts answers beside that file, but the command line joins only <root>/.jeval/labels.jsonl. Signed-off-by: rlaope <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
collect.resolve(source_key=..., question=..., answer=...)records the answer a human settled on, intolabels.jsonlnext to the records. Liketrack, it never raises.(source_key, question). An existing label is kept. An answer the question cannot produce is refused. A later answer wins, but silver never replaces a human answer. Nothing is written back to disk.jeval statusshows what has been collected per model and question, and how many more gold labels the verdict needs.docs/library.mdis the library reference.docs/instrumenting-a-service.mdis rebuilt aroundtrack/resolve/status, with every output captured from the newexamples/service-quickstart.py. Its oldthresholdoutput predated the currency format and could not be reproduced.Test plan
tests/test_runtime_labels.py, including the independent review's findings, each reproduced before the fix:resolvenever raises on hostile inputslabel --applydoes not persist joined answersdocs/library.mdis now covered by the documented-features guard.uv run pytest,ruff,mypyclean.