Brush's logfile option now writes structured, valid CSV logs that cover the whole run, without any third-party library. This PR also adds a guide notebook showing how to load and plot them. - #72
Open
gAldeia wants to merge 8 commits into
Conversation
…, and flush every gen
Open
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Better run logs
Every call to
fit/partial_fitappends to these files. Each row carries arun_id, so several runs can share onelogfile.<logfile>run_id,random_state,best_model(the best expression),best_size,best_complexity,stall_count,archive_size,n_evaluations(cumulative)<logfile>_islands.csv<logfile>_simplifications.csvconstants/inexact), return type,original→replacement, anddistance(MSE between the program's predictions before and after; inexact only)<logfile>_simplification_tablerun_idcolumn<logfile>_runs.jsonlParametersFixes and improvements
Pow(x0,x1)were written without quotes, sopd.read_csvfailed on them. All fields are now quoted when needed.Parameters, and itscurrent_genstayed at the default (1). Ids repeated across generations as a result. It is now synced every generation. Ids are only used for parent tracking and serialization, so evolution results are unchanged.Logger::verbositywas uninitialized; it now starts at 0.using namespace stdis removed fromlogger.h.Implementation
src/util/csv.h: a small header-onlyCsvWriter. It quotes fields when needed, writes the header once, flushes every row and writes floats with enough digits to read back exactly. It holds its stream in ashared_ptr, soEnginestays copyable.Engine: the statistics loop fromcalculate_stats()moved into a newsummarize(), which is reused for per-island stats. Newopen_logs()/log_stats()/close_logs().run_idis a timestamp plus a random suffix fromstd::random_device, so logging never touches Brush's random generator.SimplificationRecords*. Replacements are only recorded whenlogfileis set.Evaluation::n_evaluationscounts calls toassign_fit.Docs
docs/guide/logging.ipynb(added to the guide's table of contents). It generates logs for 3 seeds and shows run metadata, convergence, stall count, archive size, evaluations, how the best expression changed, per-island curves, the final Pareto fronts, and replacements per generation with examples.Tests
tests/cpp/test_logging.cpp(new): quoting, float formatting, header written once, row width checked, error on different columns.tests/cpp/test_brush.cpp: the engine test starts from clean log files and checks the header, one row per generation, and that the other files were created.tests/python/test_sklearn_interface.py: fits twice into the samelogfile, parses all five files and checks columns,run_ids, per-island rows and thatn_evaluationsonly increases. A new test checks that two runs with the samerandom_statelog identicalbest_modelcolumns.logfile.