From 145ccafd2ca36dc61f976cdbeaaf147af21e7c2e Mon Sep 17 00:00:00 2001 From: Stefan Jansen Date: Thu, 24 Sep 2026 13:14:25 -0400 Subject: [PATCH] docs: make core workflows verifiable and repair book links --- docs/api/index.md | 12 + docs/audit/book-integration-audit.md | 190 -------------- docs/book-guide/index.md | 244 +++++++----------- docs/getting-started/installation.md | 13 +- docs/getting-started/quickstart.md | 145 ++++------- docs/index.md | 49 ++-- .../assets/stylesheets/ml4t-docs-theme.css | 14 +- docs/user-guide/bars.md | 18 +- docs/user-guide/dataset-builder.md | 11 +- docs/user-guide/discovery.md | 9 +- docs/user-guide/features.md | 39 ++- docs/user-guide/fractional-differencing.md | 16 +- docs/user-guide/labeling.md | 33 +-- docs/user-guide/ml-readiness.md | 2 +- docs/user-guide/preprocessing.md | 6 +- 15 files changed, 254 insertions(+), 547 deletions(-) delete mode 100644 docs/audit/book-integration-audit.md diff --git a/docs/api/index.md b/docs/api/index.md index f0de241..804b566 100644 --- a/docs/api/index.md +++ b/docs/api/index.md @@ -87,6 +87,18 @@ start with the [User Guide](../index.md) or the [Book Guide](../book-guide/index options: show_root_heading: true +## Fractional Differencing + +::: ml4t.engineer.features.fdiff + options: + show_root_heading: true + signature_crossrefs: false + members: + - get_ffd_weights + - ffdiff + - find_optimal_d + - fdiff_diagnostics + ## Next Steps - Read [Features](../user-guide/features.md) for the main computation workflow. diff --git a/docs/audit/book-integration-audit.md b/docs/audit/book-integration-audit.md deleted file mode 100644 index 19c3871..0000000 --- a/docs/audit/book-integration-audit.md +++ /dev/null @@ -1,190 +0,0 @@ -# Book Integration Audit: ml4t-engineer - -*Audit date: 2026-03-03 | Library version: v0.1.0a11 | Book: Machine Learning for Trading, 3rd Edition* - -## Executive Summary - -ml4t-engineer's core value proposition is validated by heavy book usage: 120 features, 7 labeling methods, and 11 bar samplers are all exercised in chapters 3, 7-9 and all 9 case studies. Several high-quality modules (MLDatasetBuilder, FeatureCatalog.search()) previously had zero book exposure but are now showcased in Ch7 NB10. The pipeline engine and DuckDB store are honestly low-value. - ---- - -## Usage Matrix: Chapter x Module - -### Feature Computation (`compute_features`) - -| Chapter / Case Study | Notebook | Features Used | Notes | -|---------------------|----------|---------------|-------| -| Ch7 | `10_ml4t_library_ecosystem.py` | rsi, sma, ema, atr, macd, bollinger_bands | Registry tour, 3 input formats, MLDatasetBuilder | -| Ch8 | `01_price_volume_features.py` | momentum (31), trend (10), volatility (15) | Core feature teaching | -| Ch8 | `02_microstructure_features.py` | microstructure (12) | Kyle Lambda, VPIN, Amihud | -| Ch8 | `03_structural_cross_instrument_features.py` | cross-asset (10) | beta_to_market, correlations | -| Ch8 | `04_fundamentals_macro_calendar.py` | ML features, calendar | Lag, encodings, macro features | -| Ch9 | `02_structural_breaks.py` | statistics | Structural break detection | -| Ch9 | `03_fractional_differencing.py` | ffdiff (4) | Fractional differencing + ADF | -| Ch9 | `05_spectral_features.py` | ML features | Spectral, FFT | -| Ch9 | `08_garch_volatility.py` | volatility (15) | GARCH, EWMA, realized vol | -| Ch9 | `09_har_rough_volatility.py` | volatility | HAR model features | -| Ch9 | `11_hmm_regimes.py` | regime (6) | Hurst, HMM state probabilities | -| Ch9 | `13_regime_as_feature.py` | regime (6) | Regime encoding as features | -| Ch9 | `14_panel_features.py` | cross-asset (10) | Cross-sectional panel features | -| ETFs | `03_features.py`, `04_temporal.py` | momentum, volatility, volume, ffdiff | Full pipeline | -| US Equities Panel | `03_features.py`, `04_temporal.py` | momentum, volatility, ffdiff | Full pipeline | -| CME Futures | `02_labels.py`, `03_features.py` | momentum, volatility, atr | Futures-specific | - -### Labeling Methods - -| Chapter / Case Study | Notebook | Method | Config Style | -|---------------------|----------|--------|--------------| -| Ch7 | `03_label_methods.py` | triple_barrier_labels | LabelingConfig.triple_barrier() | -| Ch7 | `03_label_methods.py` | rolling_percentile_binary_labels | Direct call | -| Ch7 | `03_label_methods.py` | trend_scanning_labels | Direct call | -| Ch7 | `03_label_methods.py` | meta_labels + compute_bet_size | Meta-labeling workflow | -| Ch7 | `03_label_methods.py` | sequential_bootstrap | Sample weighting | -| CME Futures | `02_labels.py` | atr_triple_barrier_labels | LabelingConfig.atr_barrier() | -| ETFs | `02_labels.py` | rolling_percentile_binary_labels | Direct call | -| US Equities Panel | `02_labels.py` | triple_barrier_labels | LabelingConfig | -| All case studies | `02_labels.py` | fixed_time_horizon_labels | Direct call | - -### Alternative Bar Sampling - -| Chapter | Notebook | Sampler | Notes | -|---------|----------|---------|-------| -| Ch3 | `08_itch_bar_sampling.py` | TickBarSampler, VolumeBarSampler, DollarBarSampler | ITCH tick data | -| Ch3 | `10_itch_information_bars.py` | TickImbalanceBarSampler, FixedTickImbalanceBarSampler | Information-driven bars | -| Ch3 | `13_databento_bar_sampling.py` | Bar sampling on Databento data | Alternative data source | - -### Feature Discovery & Registry - -| Chapter | Notebook | API Used | -|---------|----------|----------| -| Ch7 | `10_ml4t_library_ecosystem.py` | get_registry(), list_all(), get(), list_by_category() | -| Ch7 | `10_ml4t_library_ecosystem.py` | feature_catalog.search(), feature_catalog.list(), describe() | -| Ch7 | `10_ml4t_library_ecosystem.py` | compute_features (3 formats: list, dict, YAML) | - -### MLDatasetBuilder & Preprocessing - -| Chapter | Notebook | API Used | -|---------|----------|----------| -| Ch7 | `10_ml4t_library_ecosystem.py` | create_dataset_builder, train_test_split, scaler="robust" | -| Ch7 | `10_ml4t_library_ecosystem.py` | LabelingConfig.to_yaml(), from_yaml() | -| Ch7 | `02_preprocessing_pipeline.py` | StandardScaler, split-aware preprocessing | - ---- - -## Book Chapter Structure (Actual) - -| Chapter | Directory | Notebooks | Primary ml4t-engineer Usage | -|---------|-----------|-----------|----------------------------| -| Ch3 | `03_market_microstructure/` | 17 | bars module | -| Ch7 | `07_defining_learning_task/` | 10 | labeling, registry, dataset builder | -| Ch8 | `08_feature_engineering/` | 8 + meta | features (all categories) | -| Ch9 | `09_time_series_analysis/` | 14 + meta | volatility, regime, ffdiff, cross-asset | - -### Case Study Structure (Standard Pattern) - -All 9 case studies follow the same 18-file pattern: - -| Step | File | ml4t-engineer Usage | -|------|------|---------------------| -| Setup | `01_setup.py` | — | -| Labels | `02_labels.py` | `atr_triple_barrier_labels`, `rolling_percentile_binary_labels`, `fixed_time_horizon_labels` | -| Features | `03_features.py` | `compute_features`, individual feature functions | -| Temporal | `04_temporal.py` | `ffdiff`, walk-forward CV | -| Evaluation | `05_evaluation.py` | — (ml4t-diagnostic) | -| Models | `06-13_*.py` | — | -| Backtest | `14_backtest.py` | — (ml4t-backtest) | - ---- - -## Feature Triage - -### Heavily Used (Core Value) - -| Module | Lines | Book Coverage | Confidence | Action | -|--------|-------|---------------|------------|--------| -| 120 features (10 categories) | ~8,000 | Ch8 (8 notebooks), Ch9 (14 notebooks), 9 case studies | 59 TA-Lib validated | Keep, document well | -| 7 labeling methods | ~2,000 | Ch7 NB03, all 9 case study `02_labels.py` | AFML validated | Keep, document well | -| 11 bar samplers | ~2,000 | Ch3 (3 notebooks) | Production-ready | Keep, document well | -| ffdiff module | 383 | Ch9 NB03, ETFs/Equities `04_temporal.py` | Unique value | Keep, dedicated guide | -| LabelingConfig | 467 | Ch7 NB03, all case studies | API surface | Keep, document well | -| Registry/Catalog | ~650 | Ch7 NB10 | Discovery | Keep, dedicated guide | -| MLDatasetBuilder | 638 | Ch7 NB10 (newly added) | Leakage-safe prep | Keep, dedicated guide | - -### Honestly Low-Value - -| Module | Lines | Assessment | Recommended Action | -|--------|-------|------------|-------------------| -| Pipeline engine | ~300 | `compute_features` already handles dependency ordering. Thin DAG wrapper adds little. | Label "Advanced" | -| Store (DuckDB) | ~500 | No adoption path, no book usage, no clear user need. | Label "Experimental" | -| FeatureSelector | stub | Correctly moved to ml4t-diagnostic. Stub remains as migration aid. | Keep stub, document redirect | - ---- - -## Case Studies NOT Using ml4t-engineer - -These case studies implement features manually. This is **correct** in most cases: - -| Case Study | Reason for Manual Implementation | Library Overlap | -|-----------|----------------------------------|-----------------| -| Crypto Perps Funding | Domain-specific funding rate features | None — inline appropriate | -| S&P 500 Options / Option Analytics | Greeks, IV surfaces — specialized derivatives analytics | None — out of scope | -| US Firm Characteristics | Accounting ratios from financial statements | None — out of scope | -| NASDAQ-100 Microstructure | Kyle's Lambda, Amihud, VPIN implemented manually for pedagogy | **High** — all in library (callout added) | -| FX Pairs | Garman-Klass volatility, momentum features | **Partial** — some in library (callout added) | - ---- - -## Cross-Reference: Book Notebooks Using ml4t.engineer - -### Direct imports (`from ml4t.engineer`) - -| File | Imports | Status | -|------|---------|--------| -| `07_defining_learning_task/code/10_ml4t_library_ecosystem.py` | compute_features, get_registry, feature_catalog, create_dataset_builder, LabelingConfig | Working | -| `07_defining_learning_task/code/03_label_methods.py` | LabelingConfig, 7 labeling functions | Working (migrated from BarrierConfig, un-skipped) | -| `07_defining_learning_task/code/04_minimum_favorable_adverse_excursion.py` | LabelingConfig | Working (migrated from BarrierConfig, un-skipped) | -| `08_feature_engineering/code/01_price_volume_features.py` | ml4t.engineer.features.volatility, momentum, trend | Working | -| `09_time_series_analysis/code/08_garch_volatility.py` | ml4t.engineer.features.volatility (6 functions) | Working | -| All case study `02_labels.py` | ml4t.engineer.labeling (atr_triple_barrier_labels etc.) | Working | -| All case study `03_features.py` | ml4t.engineer.features (momentum, volatility, regime, trend) | Working | - -### Indirect usage (via `utils/label_functions.py`) - -Some case study `02_labels.py` files use standalone label utility wrappers that mirror the ml4t.engineer API. These are isolated from API changes but are less idiomatic. - ---- - -## Documentation Coverage - -| User Guide Page | Lines | Book Reference | Status | -|----------------|-------|----------------|--------| -| `labeling.md` | 522 | Ch7 `03_label_methods.py`, CME `02_labels.py`, ETFs `02_labels.py` | Complete | -| `features.md` | 388 | Ch8 NB01-04, Ch9 NB08-14, ETFs/Equities/CME `03_features.py` | Complete | -| `bars.md` | 405 | Ch3 `08_itch_bar_sampling.py`, `10_itch_information_bars.py`, `13_databento_bar_sampling.py` | Complete | -| `ml-readiness.md` | 178 | Ch8 `01_price_volume_features.py` | Complete | -| `discovery.md` | 162 | Ch7 `10_ml4t_library_ecosystem.py` | Complete | -| `fractional-differencing.md` | 188 | Ch9 `03_fractional_differencing.py`, ETFs/Equities `04_temporal.py` | Complete | -| `preprocessing.md` | 171 | Ch7 `02_preprocessing_pipeline.py` | Complete | -| `dataset-builder.md` | 201 | Ch7 `10_ml4t_library_ecosystem.py` | Complete | - ---- - -## Value Assessment - -### What ml4t-engineer does well - -1. **Feature computation is the clear winner**: 120 features, validated, fast, config-driven. Used in 30+ notebooks across 8 chapters and 9 case studies. -2. **Labeling methods are comprehensive**: All 7 AFML methods implemented, validated, calendar-aware. Used in every case study. -3. **Bar sampling is uniquely valuable**: No other Python library provides production-quality imbalance bars with threshold spiral warnings. -4. **Registry/discovery is elegant**: Metadata-driven feature selection with TA-Lib compatibility flags and normalization status. -5. **MLDatasetBuilder fills a real gap**: Leakage-safe dataset prep with CV integration — now demonstrated in Ch7 NB10. - -### What should be scoped honestly - -1. **Pipeline engine**: `compute_features` already does dependency ordering. The Pipeline class adds a thin DAG wrapper that few users need. Document as "Advanced". -2. **DuckDB Store**: No user demand, no book usage. Keep but label experimental. -3. **Cross-asset features (8 of 10 unused in book)**: Strong implementations but limited coverage. Only `beta_to_market` and `rolling_correlation` are commonly needed. - ---- - -*This audit was used to drive the user guide expansion and book notebook updates for v0.1.0a11.* diff --git a/docs/book-guide/index.md b/docs/book-guide/index.md index a8c996d..75955a5 100644 --- a/docs/book-guide/index.md +++ b/docs/book-guide/index.md @@ -1,151 +1,97 @@ # Book Guide -Use this guide to move between *Machine Learning for Trading, Third Edition* and -`ml4t-engineer` without guessing which notebook maps to which production API. - -`ml4t-engineer` appears in two distinct ways throughout the book: - -- pedagogical notebooks that build ideas step by step -- reusable library workflows that collapse those ideas into stable APIs - -The book teaches the method. The library is where you should go when you want to -reuse that method across assets, case studies, and production pipelines. - -## How to Use This Guide - -Start from the book if you want intuition, derivations, and plots. Start from the -library guides if you want reusable implementations, validated APIs, and pipeline -integration. - -Use this page when you need to answer one of these questions: - -- "Which chapter teaches the concept behind this API?" -- "Which notebook should I read before using this workflow in production?" -- "Which library guide replaces the manual code from the book?" - -## Recommended Reader Journey - -1. Read the relevant chapter notebook to understand the modeling idea. -2. Jump to the matching user-guide page for the production API. -3. Use the case-study pipeline files as the bridge from teaching code to reusable - workflows. -4. Use the [API Reference](../api/index.md) when you need exact signatures. - -## Chapter Map - -### Chapter 3: Market Microstructure - -| Book path | What the book teaches | Library entry point | Docs page | -|-----------|-----------------------|---------------------|-----------| -| `03_market_microstructure/code/08_itch_bar_sampling.py` | Why time bars are statistically weak and how tick, volume, and dollar bars improve sampling | `TickBarSampler`, `VolumeBarSampler`, `DollarBarSampler` | [Alternative Bars](../user-guide/bars.md) | -| `03_market_microstructure/code/10_itch_information_bars.py` | Imbalance bars and threshold dynamics | `TickImbalanceBarSampler`, `FixedTickImbalanceBarSampler`, `FixedVolumeImbalanceBarSampler` | [Alternative Bars](../user-guide/bars.md) | -| `03_market_microstructure/code/13_databento_bar_sampling.py` | Applying bar samplers to a modern vendor feed | Same sampler family with production input contracts | [Alternative Bars](../user-guide/bars.md) | - -What changes when you move to the library: - -- the book focuses on the statistical motivation and diagnostics -- the library gives you stable samplers, warnings, and reusable OHLCV outputs -- fixed-threshold imbalance bars are the recommended production path - -### Chapter 7: Defining the Learning Task - -| Book path | What the book teaches | Library entry point | Docs page | -|-----------|-----------------------|---------------------|-----------| -| `07_defining_learning_task/code/02_preprocessing_pipeline.py` | Leakage-safe preprocessing and split-aware scaling | `StandardScaler`, `MinMaxScaler`, `RobustScaler`, `PreprocessingPipeline` | [Preprocessing](../user-guide/preprocessing.md) | -| `07_defining_learning_task/code/03_label_methods.py` | Triple-barrier, percentile, trend-scanning, meta-labeling, and sample weighting | `LabelingConfig`, `triple_barrier_labels`, `rolling_percentile_binary_labels`, `trend_scanning_labels`, `meta_labels` | [Labeling](../user-guide/labeling.md) | -| `07_defining_learning_task/code/04_minimum_favorable_adverse_excursion.py` | Barrier behavior and excursion analysis | `LabelingConfig.triple_barrier()` | [Labeling](../user-guide/labeling.md) | -| `07_defining_learning_task/code/10_ml4t_library_ecosystem.py` | The library-oriented view of feature computation, discovery, and dataset building | `compute_features`, `feature_catalog`, `create_dataset_builder` | [Features](../user-guide/features.md), [Feature Discovery](../user-guide/discovery.md), [Dataset Builder](../user-guide/dataset-builder.md) | - -What changes when you move to the library: - -- manual notebook experiments become serialized `LabelingConfig` workflows -- feature discovery moves from ad hoc inspection to metadata-driven search -- dataset preparation becomes train-only scaling and splitter-aware folds by default - -### Chapter 8: Feature Engineering - -| Book path | What the book teaches | Library entry point | Docs page | -|-----------|-----------------------|---------------------|-----------| -| `08_feature_engineering/code/01_price_volume_features.py` | Momentum, trend, volatility, and volume feature intuition | `compute_features` with registry-backed indicators | [Features](../user-guide/features.md) | -| `08_feature_engineering/code/02_microstructure_features.py` | Microstructure feature construction and interpretation | Microstructure feature functions and registry metadata | [Features](../user-guide/features.md) | -| `08_feature_engineering/code/03_structural_cross_instrument_features.py` | Cross-asset and panel relationships | `ml4t.engineer.features.cross_asset` | [Features](../user-guide/features.md) | -| `08_feature_engineering/code/04_fundamentals_macro_calendar.py` | Lag features, calendar encodings, and ML-oriented transforms | ML feature utilities and preprocessing bridge | [Features](../user-guide/features.md), [ML Readiness](../user-guide/ml-readiness.md) | - -What changes when you move to the library: - -- the book derives and visualizes features individually -- the library lets you request validated feature sets through one computation API -- registry and catalog metadata help you choose features systematically - -### Chapter 9: Time-Series Analysis - -| Book path | What the book teaches | Library entry point | Docs page | -|-----------|-----------------------|---------------------|-----------| -| `09_time_series_analysis/code/03_fractional_differencing.py` | The memory-stationarity tradeoff and ADF-based search for `d` | `ffdiff`, `find_optimal_d`, `fdiff_diagnostics` | [Fractional Differencing](../user-guide/fractional-differencing.md) | -| `09_time_series_analysis/code/08_garch_volatility.py` | Volatility estimators and conditional volatility modeling | Volatility feature family | [Features](../user-guide/features.md) | -| `09_time_series_analysis/code/09_har_rough_volatility.py` | Multi-horizon volatility structure | Volatility features used downstream in pipelines | [Features](../user-guide/features.md) | -| `09_time_series_analysis/code/11_hmm_regimes.py` | Regime detection workflows | Regime feature family | [Features](../user-guide/features.md) | -| `09_time_series_analysis/code/13_regime_as_feature.py` | Turning regimes into model inputs | Regime features in reusable pipelines | [Features](../user-guide/features.md) | -| `09_time_series_analysis/code/14_panel_features.py` | Cross-sectional panel features | Cross-asset feature functions for multi-asset inputs | [Features](../user-guide/features.md) | - -What changes when you move to the library: - -- research notebooks stay focused on method validation -- library functions package the same transforms into repeatable feature pipelines -- the case studies show how to combine these transforms with labels and CV - -## Case-Study Pipeline Map - -Most case studies follow the same structure. `ml4t-engineer` is the handoff point -between raw market data and model-ready datasets. - -| Case-study step | Typical file | Library workflow | Docs page | -|-----------------|--------------|------------------|-----------| -| Labels | `case_studies//code/02_labels.py` | `LabelingConfig`, barrier labels, percentile labels, fixed-horizon labels | [Labeling](../user-guide/labeling.md) | -| Features | `case_studies//code/03_features.py` | `compute_features` plus feature-specific functions | [Features](../user-guide/features.md) | -| Temporal prep | `case_studies//code/04_temporal.py` | fractional differencing and leakage-safe preparation | [Fractional Differencing](../user-guide/fractional-differencing.md), [Dataset Builder](../user-guide/dataset-builder.md) | - -Examples called out in the current integration audit: - -- ETFs: percentile labels, production `compute_features`, fractional differencing -- US Equities Panel: triple-barrier labels, panel features, fractional differencing -- CME Futures: ATR-based barriers and futures-aware labeling workflows -- NASDAQ-100 Microstructure: manual pedagogical implementations with strong overlap to - production-ready microstructure features in this library - -## From Notebook Code to Library API - -Use this translation when moving from the book to reusable code: - -| Book pattern | Library equivalent | -|--------------|--------------------| -| Manually computing several indicators in sequence | `compute_features(data, feature_spec)` | -| Notebook-only feature browsing | `feature_catalog.list()`, `feature_catalog.search()`, `registry.get()` | -| Inline barrier parameters spread across a notebook | `LabelingConfig.triple_barrier(...)` or `LabelingConfig.atr_barrier(...)` | -| One-off train/test scaling | `create_dataset_builder(..., scaler=...)` | -| Manual stationarity experiments | `find_optimal_d()` then `ffdiff()` | -| Bar-construction experiments | sampler classes in `ml4t.engineer.bars` | - -## Maturity and Scope - -These are the workflows readers should prioritize: - -- production-ready: feature computation, labeling methods, alternative bars, - fractional differencing, feature discovery, dataset builder -- advanced: cross-asset feature workflows and pipeline orchestration -- experimental or low-priority: DuckDB store and any workflows not yet used in the - book or case studies - -This matches the current audit: the strongest value in `ml4t-engineer` is the -production API for feature engineering, labeling, and dataset preparation. - -## Where to Go Next - -- Start with [Quickstart](../getting-started/quickstart.md) if you want a working - example first. -- Read [Features](../user-guide/features.md) for the core computation API. -- Read [Labeling](../user-guide/labeling.md) if you are building supervised targets. -- Read [Alternative Bars](../user-guide/bars.md) for microstructure workflows. -- Read [Dataset Builder](../user-guide/dataset-builder.md) for leakage-safe model - inputs. -- Use [API Reference](../api/index.md) for exact signatures. +This guide maps `ml4t-engineer` tasks to public notebooks from *Machine Learning for +Trading, Third Edition*. Every link uses companion commit +[`d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb`](https://github.com/stefan-jansen/machine-learning-for-trading/tree/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb). +All paths and the paired Python sources were checked at that commit. + +The relationship column distinguishes three cases: + +- **Calls Engineer**: the notebook imports and runs the named `ml4t.engineer` API. +- **Teaches manually**: the notebook implements the method for instruction and does not + use Engineer for that task. +- **Related workflow**: the notebook shows where the task fits, but its broader workflow + is not an Engineer API example. + +## Feature computation and discovery + +| Book notebook | Relationship | Engineer API | Task guide | +|---|---|---|---| +| [The ml4t Library Ecosystem](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/10_ml4t_library_ecosystem.ipynb) | Calls Engineer to inspect registry metadata and run `compute_features()` with names and parameter dictionaries | `compute_features`, `get_registry` | [Features](../user-guide/features.md), [Feature Discovery](../user-guide/discovery.md) | +| [Price and Volume Feature Families](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/08_financial_features/01_price_volume_features.ipynb) | Calls Engineer for registry features, volatility, regime, risk, and fractional-differencing functions; also derives selected features manually | `compute_features` and feature modules | [Features](../user-guide/features.md), [ML Readiness](../user-guide/ml-readiness.md) | +| [Microstructure Features](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/08_financial_features/02_microstructure_features.ipynb) | Calls Engineer for the tick rule and liquidity estimators, then builds a wider teaching workflow | `ml4t.engineer.features.microstructure` | [Features](../user-guide/features.md) | +| [Structural and Cross-Instrument Features](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/08_financial_features/03_structural_cross_instrument_features.ipynb) | Calls Engineer for market beta; teaches carry and options features manually | `beta_to_market` | [Features](../user-guide/features.md) | +| [Slow Features and Context](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/08_financial_features/04_fundamentals_macro_calendar.ipynb) | Calls Engineer for calendar encoding; teaches point-in-time joins and slow features manually | `cyclical_encode` | [Features](../user-guide/features.md) | +| [Panel Features](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/09_model_based_features/14_panel_features.ipynb) | Calls Engineer for cross-asset features and compares them with manual statistical work | `ml4t.engineer.features.cross_asset` | [Features](../user-guide/features.md) | +| [ETFs: Feature Engineering](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/case_studies/etfs/03_financial_features.ipynb) | Calls individual Engineer feature functions inside a full case-study pipeline | momentum, trend, volatility, volume, and regime feature modules | [Features](../user-guide/features.md) | + +The chapter notebooks use book datasets and plotting dependencies. The +[Quickstart](../getting-started/quickstart.md) provides an offline synthetic path for +the same released computation API. + +## Labeling + +| Book notebook | Relationship | Engineer API | Task guide | +|---|---|---|---| +| [Label Engineering Methods](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/03_label_methods.ipynb) | Calls Engineer for fixed-horizon, percentile, triple-barrier, ATR-barrier, trend-scanning, meta-labeling, and sample-weighting workflows | `ml4t.engineer.labeling`, `LabelingConfig` | [Labeling](../user-guide/labeling.md) | +| [ETFs: Label Engineering](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/case_studies/etfs/02_labels.ipynb) | Related workflow that constructs and audits case-study labels without calling Engineer | no direct Engineer call | [Labeling](../user-guide/labeling.md) | + +The second notebook is useful for the artifact and timing workflow. It is not evidence +that the case study uses Engineer's labeling functions. + +## Alternative bars + +| Book notebook | Relationship | Engineer API | Task guide | +|---|---|---|---| +| [ITCH Bar Sampling](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/03_market_microstructure/14_itch_bar_sampling.ipynb) | Calls Engineer for tick, volume, dollar, imbalance, and run bars on ITCH trades | bar sampler classes | [Alternative Bars](../user-guide/bars.md) | +| [Information-Bar Formulas and Parameters](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/03_market_microstructure/16_itch_information_bars.ipynb) | Calls Engineer and compares manual formulas with adaptive, fixed, and window samplers | imbalance-bar sampler classes | [Alternative Bars](../user-guide/bars.md) | +| [Databento Bar Calibration](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/03_market_microstructure/17_databento_bar_sampling.ipynb) | Calls Engineer in a multi-day calibration workflow that requires Databento data | bar sampler classes | [Alternative Bars](../user-guide/bars.md) | + +The first two notebooks require book data. The Databento notebook also requires the +vendor dataset. The task guide and `examples/bars_example.py` provide an offline, +synthetic verification path. + +## Preprocessing and fractional differencing + +| Book notebook | Relationship | Engineer API | Task guide | +|---|---|---|---| +| [Preprocessing Pipeline](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/02_preprocessing_pipeline.ipynb) | Calls Engineer's `StandardScaler` for train-only fitting; teaches the broader cleaning pipeline manually | `StandardScaler` | [Preprocessing](../user-guide/preprocessing.md) | +| [Fractional Differencing](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/09_model_based_features/03_fractional_differencing.ipynb) | Calls Engineer's fractional-differencing helpers while teaching the statistical method | `ffdiff`, `find_optimal_d`, `fdiff_diagnostics` | [Fractional Differencing](../user-guide/fractional-differencing.md) | +| [ETFs: Model-Based Features](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/case_studies/etfs/04_model_based_features.ipynb) | Calls `ffdiff` inside a walk-forward case-study workflow; its HMM and GARCH work is outside Engineer's fractional-differencing API | `ffdiff` | [Fractional Differencing](../user-guide/fractional-differencing.md) | + +No checked book notebook at this revision calls `create_dataset_builder`. Use the +[Dataset Builder guide](../user-guide/dataset-builder.md) and the repository's +`examples/complete_workflow_example.py` for that workflow. + +## Alphalens migration scope + +Engineer overlaps with Alphalens only before factor analysis: it can compute factor +values with `compute_features()` and fit preprocessing state on training data. Engineer +does not replace Alphalens tearsheets, information-coefficient analysis, quantile-return +analysis, turnover analysis, or event studies. Use +[ML4T Diagnostic](https://www.ml4trading.io/docs/diagnostic/) for those evaluation +tasks. The book's feature notebooks above show factor construction; topical similarity +does not make them Alphalens replacements. + +## Supported and experimental boundaries + +- The principal documented workflows are feature computation and discovery, labeling, + alternative bars, preprocessing, dataset building, and fractional differencing. +- Cross-asset functions are supported advanced APIs. Validate asset ordering and + point-in-time alignment before use. +- Adaptive imbalance bars require calibration. Fixed-threshold samplers provide the + bounded production path described in the [Alternative Bars guide](../user-guide/bars.md). +- `fdiff_diagnostics()` and `find_optimal_d()` require the `stats` extra. `ffdiff()` is + available from the core installation. +- The optional DuckDB store is experimental and has no verified book adoption path. +- `transfer_entropy()` is not implemented for production use. It is not part of the + principal documented workflow. + +## Run a workflow first + +- [Quickstart](../getting-started/quickstart.md) computes features from synthetic data. +- [Features](../user-guide/features.md) covers configuration and input contracts. +- [Labeling](../user-guide/labeling.md) covers target construction and timing. +- [Alternative Bars](../user-guide/bars.md) covers sampler choice and calibration. +- [Dataset Builder](../user-guide/dataset-builder.md) covers leakage-safe splits. +- [API Reference](../api/index.md) provides exact signatures. diff --git a/docs/getting-started/installation.md b/docs/getting-started/installation.md index 5fdbda7..99aaf01 100644 --- a/docs/getting-started/installation.md +++ b/docs/getting-started/installation.md @@ -44,6 +44,17 @@ pip install TA-Lib pip install TA-Lib ``` +### Statistical diagnostics + +`find_optimal_d()` and `fdiff_diagnostics()` use Statsmodels for the augmented +Dickey-Fuller test: + +```bash +pip install "ml4t-engineer[stats]" +``` + +The core `ffdiff()` transform does not require this extra. + ## Verify Installation ```python @@ -63,7 +74,7 @@ Categories: ['math', 'microstructure', 'ml', 'momentum', 'price_transform', 'reg ## Next Steps -- Read [Quickstart](quickstart.md) for a first working feature and labeling example. +- Read [Quickstart](quickstart.md) for the first feature-computation workflow. - Read the [Book Guide](../book-guide/index.md) if you are coming from the book or case studies. - Use the [API Reference](../api/index.md) once you need exact object locations. diff --git a/docs/getting-started/quickstart.md b/docs/getting-started/quickstart.md index aae1b42..31ebee8 100644 --- a/docs/getting-started/quickstart.md +++ b/docs/getting-started/quickstart.md @@ -1,117 +1,68 @@ # Quickstart -Get up and running with ML4T Engineer in 5 minutes. +Use synthetic OHLCV data to compute three features with the released package. This +workflow needs no credentials, external service, optional dependency, or special +hardware. -If you are coming from *Machine Learning for Trading, Third Edition*, pair this page -with the [Book Guide](../book-guide/index.md) to jump from notebook examples to the -matching reusable library workflows. +## Install -## Basic Feature Computation +Install `ml4t-engineer` in a Python 3.12, 3.13, or 3.14 environment: + +```bash +pip install ml4t-engineer +``` + +## Compute a feature matrix ```python -from datetime import datetime, timedelta +from datetime import date, timedelta import polars as pl from ml4t.engineer import compute_features -# Create sample OHLCV data -df = pl.DataFrame({ - "timestamp": [datetime(2024, 1, 1) + timedelta(days=i) for i in range(100)], - "open": [100.0, 101.0, 102.0, 103.0, 104.0] * 20, - "high": [102.0, 103.0, 104.0, 105.0, 106.0] * 20, - "low": [99.0, 100.0, 101.0, 102.0, 103.0] * 20, - "close": [101.0, 102.0, 103.0, 104.0, 105.0] * 20, - "volume": [1000, 1100, 1200, 1300, 1400] * 20, -}) - -# Compute features -result = compute_features(df, ["rsi", "macd", "atr"]) -print(result.columns) -``` - -## Custom Parameters - -```python -# Use dict format for custom parameters -result = compute_features(df, [ - {"name": "rsi", "params": {"period": 20}}, - {"name": "sma", "params": {"period": 50}}, +close = [100.0 + i * 0.1 + (i % 7) * 0.2 for i in range(100)] +ohlcv = pl.DataFrame( { - "name": "bollinger_bands", - "params": {"period": 20, "nbdevup": 2.5, "nbdevdn": 2.5}, - }, -]) -``` - -## YAML Configuration - -Create a `features.yaml` file: - -```yaml -features: - - name: rsi - params: - period: 14 - - name: macd - params: - fast_period: 12 - slow_period: 26 - - name: atr - params: - period: 14 -``` - -Then use it: - -```python -result = compute_features(df, "features.yaml") -``` - -## Triple-Barrier Labeling - -```python -from ml4t.engineer.config import LabelingConfig -from ml4t.engineer.labeling import triple_barrier_labels - -config = LabelingConfig.triple_barrier( - upper_barrier=0.02, # 2% profit target - lower_barrier=0.01, # 1% stop loss - max_holding_period=20, # 20 bar maximum holding + "timestamp": [date(2024, 1, 1) + timedelta(days=i) for i in range(100)], + "open": close, + "high": [price + 1.0 for price in close], + "low": [price - 1.0 for price in close], + "close": close, + "volume": [100_000 + i * 100 for i in range(100)], + } ) -labels = triple_barrier_labels( - df, - config=config, -) -``` - -## Explore Available Features +features = compute_features(ohlcv, ["rsi", "macd", "atr"]) -```python -from ml4t.engineer import feature_catalog +added = [name for name in features.columns if name not in ohlcv.columns] +print(f"rows={features.height}") +print(f"added={added}") -# List all categories -print(feature_catalog.categories()) -# ['momentum', 'trend', 'volatility', ...] +assert features.height == 100 +assert added == ["rsi", "macd", "atr"] +assert features.select(added).drop_nulls().height > 0 +``` -# List features in a category -print(feature_catalog.list(category="momentum")) -# ['rsi', 'macd', 'stoch', 'cci', ...] +Expected result: -# Get feature details -info = feature_catalog.describe("rsi") -print(info) -# {'name': 'rsi', 'category': 'momentum', 'normalized': True, ...} +```text +rows=100 +added=['rsi', 'macd', 'atr'] ``` -## Next Steps - -- [Features Guide](../user-guide/features.md) - 120 features across 11 categories -- [Labeling Guide](../user-guide/labeling.md) - 7 labeling methods for supervised learning -- [Alternative Bars](../user-guide/bars.md) - Information-driven bar sampling -- [Feature Discovery](../user-guide/discovery.md) - Registry, catalog, and search API -- [Fractional Differencing](../user-guide/fractional-differencing.md) - Memory-preserving stationarity -- [Dataset Builder](../user-guide/dataset-builder.md) - Leakage-safe train/test preparation -- [Book Guide](../book-guide/index.md) - Chapter and case-study map for the book -- [API Reference](../api/index.md) - Complete API documentation +`compute_features()` preserves the input rows and columns, then appends the requested +features. Rolling features contain null values during their warmup windows. The final +assertion verifies that the three features produce values after warmup. + +## Continue with your task + +- [Compute and configure features](../user-guide/features.md) +- [Inspect available features](../user-guide/discovery.md) +- [Create supervised labels](../user-guide/labeling.md) +- [Sample alternative bars](../user-guide/bars.md) +- [Build leakage-safe train/test data](../user-guide/dataset-builder.md) +- [Apply train-only preprocessing](../user-guide/preprocessing.md) +- [Apply fractional differencing](../user-guide/fractional-differencing.md) +- [Look up exact signatures](../api/index.md) +- [Open matching book notebooks](../book-guide/index.md) diff --git a/docs/index.md b/docs/index.md index fd54e0a..0a46ce8 100644 --- a/docs/index.md +++ b/docs/index.md @@ -4,48 +4,39 @@ Feature engineering, labeling, alternative bars, and leakage-safe datasets for financial ML. `ml4t-engineer` is the feature-engineering layer in the ML4T stack. It sits between -`ml4t-data`, which prepares canonical datasets, and `ml4t-diagnostic`, which -evaluates signals and models. Start here if you want a working workflow quickly, use -the [Book Guide](book-guide/index.md) to map notebooks to production APIs, and use the -[API Reference](api/index.md) when you need exact interfaces. - -Chapters 7-10 of *Machine Learning for Trading, Third Edition* develop many of these -methods manually in notebooks. This library packages those computations as tested, -reusable functions. See the [Book Guide](book-guide/index.md) to map notebook code to -library calls. +`ml4t-data`, which prepares canonical datasets, and `ml4t-diagnostic`, which evaluates +signals and models.
-- :material-chart-line:{ .lg .middle } __120 Features, One Call__ +- :material-chart-line:{ .lg .middle } __First successful workflow__ --- - Momentum, volatility, microstructure, trend, and other feature families through - `compute_features(df, indicators)`. - [:octicons-arrow-right-24: Features](user-guide/features.md) + Install the released package, compute three features from synthetic data, and + verify the result. + [:octicons-arrow-right-24: Quickstart](getting-started/quickstart.md) -- :material-check-decagram:{ .lg .middle } __60 TA-Lib Validated__ +- :material-check-decagram:{ .lg .middle } __Task guides__ --- - Indicators tested against TA-Lib to `1e-6` tolerance so notebook and pipeline - outputs stay aligned. - [:octicons-arrow-right-24: Quickstart](getting-started/quickstart.md) + Compute features, create labels, sample bars, and build leakage-safe datasets. + [:octicons-arrow-right-24: Features](user-guide/features.md) -- :material-label:{ .lg .middle } __Labels, Bars, and Leakage Control__ +- :material-label:{ .lg .middle } __Exact API reference__ --- - Triple-barrier labels, alternative bars, preprocessing, and dataset splitting in - the same workflow. - [:octicons-arrow-right-24: Labeling](user-guide/labeling.md) + Look up released functions, classes, signatures, and supported options. + [:octicons-arrow-right-24: API Reference](api/index.md) -- :material-book-open-variant:{ .lg .middle } __Book to Production__ +- :material-book-open-variant:{ .lg .middle } __Checked book notebooks__ --- - The book teaches the methods step by step. This library turns them into reusable - calls for research and scheduled pipelines. + Open commit-pinned notebooks and see whether each one calls Engineer, teaches the + method manually, or illustrates a related workflow. [:octicons-arrow-right-24: Book Guide](book-guide/index.md)
@@ -73,10 +64,11 @@ df = pl.DataFrame({ features = compute_features(df, ["rsi", "macd", "atr"]) assert {"rsi", "macd", "atr"} <= set(features.columns) +assert features.select(["rsi", "macd", "atr"]).drop_nulls().height > 0 ``` -That single call appends validated indicator columns to the same DataFrame you will -pass downstream into labeling, preprocessing, and model training. +The call appends three indicator columns to the input DataFrame. The assertions check +that the columns exist and contain values after their rolling warmup windows. ## Core Workflows @@ -141,12 +133,11 @@ result = find_optimal_d(df["close"]) ffd_close = ffdiff(df["close"], d=result["optimal_d"]) ``` -This is the standard bridge from Chapter 9’s fractional differencing workflow to a -reusable production transform. +The parameter search requires the `stats` extra. The [Fractional Differencing guide](user-guide/fractional-differencing.md) explains the statistical check and the core-only `ffdiff()` transform. ## Documentation Entry Points -- [Quickstart](getting-started/quickstart.md) for a working feature and labeling run +- [Quickstart](getting-started/quickstart.md) for the first feature-computation workflow - [Features](user-guide/features.md) for the core computation API - [Labeling](user-guide/labeling.md) for supervised targets and sample weighting - [Book Guide](book-guide/index.md) for chapter, notebook, and case-study mapping diff --git a/docs/overrides/assets/stylesheets/ml4t-docs-theme.css b/docs/overrides/assets/stylesheets/ml4t-docs-theme.css index 01dca7d..b89b6dd 100644 --- a/docs/overrides/assets/stylesheets/ml4t-docs-theme.css +++ b/docs/overrides/assets/stylesheets/ml4t-docs-theme.css @@ -32,7 +32,7 @@ --md-code-bg-color: #f4f5f3; --md-code-fg-color: var(--ml4t-navy); - /* Footer colors — match website footer */ + /* Footer colors - match website footer */ --md-footer-bg-color: var(--ml4t-navy); --md-footer-bg-color--dark: var(--ml4t-navy-dark); --md-footer-fg-color: var(--ml4t-silver); @@ -148,7 +148,7 @@ body { box-shadow: none; } -/* Collapse the MkDocs header bar to zero — only tabs remain visible */ +/* Collapse the MkDocs header bar to zero - only tabs remain visible */ .md-header__inner { display: none; height: 0; @@ -156,7 +156,7 @@ body { overflow: hidden; } -/* Tabs styled as integrated sub-navigation — same visual +/* Tabs styled as integrated sub-navigation - same visual language as the chapter page tabs on the main website. Uses px units because MkDocs sets html font-size: 125%. */ .md-tabs { @@ -212,7 +212,7 @@ body { font-size: 0.9rem; } -/* FIX #4: TOC sidebar — match website sidebar styling */ +/* FIX #4: TOC sidebar - match website sidebar styling */ .md-sidebar--secondary .md-nav__title { font-size: 0.78rem; font-weight: 700; @@ -268,7 +268,7 @@ body { box-shadow: inset 0 1px 0 rgb(255 255 255 / 0.4); } -/* Code copy button — visible on light background */ +/* Code copy button - visible on light background */ .md-clipboard { color: var(--ml4t-text-light); } @@ -398,6 +398,10 @@ body { } @media screen and (max-width: 44.9375em) { + body:has(h1#book-guide) .md-typeset table:not([class]) { + min-width: 48rem; + } + .md-typeset h1 { font-size: 1.7rem; } diff --git a/docs/user-guide/bars.md b/docs/user-guide/bars.md index 6f922fa..79ea73f 100644 --- a/docs/user-guide/bars.md +++ b/docs/user-guide/bars.md @@ -5,10 +5,7 @@ Transform tick data into information-driven bars instead of time-based bars. Use this page when you want to replace time bars with sampling schemes that better match market activity and microstructure dynamics. -> **Book**: *ML for Trading, 3rd ed.* — Ch3 `08_itch_bar_sampling.py` constructs tick, volume, and dollar bars from ITCH trade data. `10_itch_information_bars.py` builds imbalance bars with threshold analysis. `13_databento_bar_sampling.py` demonstrates bar sampling on Databento data. - -Use the [Book Guide](../book-guide/index.md) for the chapter-level map from the -microstructure notebooks to the production sampler classes. +The book notebooks [ITCH Bar Sampling](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/03_market_microstructure/14_itch_bar_sampling.ipynb), [Information-Bar Formulas and Parameters](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/03_market_microstructure/16_itch_information_bars.ipynb), and [Databento Bar Calibration](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/03_market_microstructure/17_databento_bar_sampling.ipynb) call these sampler classes. The last notebook requires vendor data. The [Book Guide](../book-guide/index.md) records the checked revision and verification limits. ## Why Alternative Bars? @@ -18,8 +15,9 @@ Time bars (1min, 1h, daily) have problems: - Autocorrelation in returns - Poor statistical properties (non-normal, heteroskedastic) -Alternative bars sample based on market activity, producing bars with more uniform -information content and better statistical properties for ML models. +Alternative bars sample on market activity rather than elapsed time. Compare their +return distribution, duration, and autocorrelation on your own trade data before choosing +a sampler and threshold. ## Quick Start @@ -425,10 +423,10 @@ inner loops for optimal performance. For very large datasets, consider: ## See It In The Book -- Ch3 `08_itch_bar_sampling.py` for tick, volume, and dollar bars -- Ch3 `10_itch_information_bars.py` for imbalance-bar intuition and diagnostics -- Ch3 `13_databento_bar_sampling.py` for a modern market-data workflow -- [Book Guide](../book-guide/index.md) for the full chapter-to-API map +- [ITCH Bar Sampling](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/03_market_microstructure/14_itch_bar_sampling.ipynb) calls the standard and run-bar samplers. +- [Information-Bar Formulas and Parameters](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/03_market_microstructure/16_itch_information_bars.ipynb) compares manual formulas with Engineer's imbalance samplers. +- [Databento Bar Calibration](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/03_market_microstructure/17_databento_bar_sampling.ipynb) calls the samplers in a vendor-data workflow. +- [Book Guide](../book-guide/index.md) records the pinned revision and requirements. ## Next Steps diff --git a/docs/user-guide/dataset-builder.md b/docs/user-guide/dataset-builder.md index 7fb1743..816a881 100644 --- a/docs/user-guide/dataset-builder.md +++ b/docs/user-guide/dataset-builder.md @@ -5,10 +5,7 @@ Use this page when you already have features and labels and want a reusable bridge from engineered data to train/test or cross-validation folds. -> **Book**: *ML for Trading, 3rd ed.* — Ch7 `10_ml4t_library_ecosystem.py` demonstrates `MLDatasetBuilder` with triple-barrier labels: features + labels in, scaled train/test split out. Ch7 `02_preprocessing_pipeline.py` covers the underlying preprocessing concepts. - -Use the [Book Guide](../book-guide/index.md) if you want the full bridge from the -Chapter 7 teaching notebooks to reusable dataset workflows in the library. +No checked book notebook at the pinned revision calls `MLDatasetBuilder` or `create_dataset_builder`. [Preprocessing Pipeline](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/02_preprocessing_pipeline.ipynb) calls Engineer's `StandardScaler` and teaches the split-aware preprocessing problem that the builder solves. Use the [Book Guide](../book-guide/index.md) for the exact relationship. ## Basic Usage @@ -228,9 +225,9 @@ X_train, X_test, y_train, y_test = builder.train_test_split(train_size=0.8) ## See It In The Book -- Ch7 `10_ml4t_library_ecosystem.py` for the end-to-end dataset-builder workflow -- Ch7 `02_preprocessing_pipeline.py` for the preprocessing logic that underpins it -- [Book Guide](../book-guide/index.md) for the surrounding chapter and case-study map +- [Preprocessing Pipeline](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/02_preprocessing_pipeline.ipynb) calls `StandardScaler`, not the dataset builder. +- [Book Guide](../book-guide/index.md) records that no checked notebook calls `create_dataset_builder`. +- Run `examples/complete_workflow_example.py` for an end-to-end dataset-builder workflow. ## Next Steps diff --git a/docs/user-guide/discovery.md b/docs/user-guide/discovery.md index cc5409e..e2146f3 100644 --- a/docs/user-guide/discovery.md +++ b/docs/user-guide/discovery.md @@ -2,9 +2,7 @@ ML4T Engineer provides two complementary discovery APIs: the **Feature Registry** for programmatic metadata access, and the **Feature Catalog** for interactive exploration with filtering and search. -If you are arriving from Ch7 `10_ml4t_library_ecosystem.py`, the -[Book Guide](../book-guide/index.md) shows where discovery fits relative to feature -computation, labeling, and dataset preparation. +[The ml4t Library Ecosystem](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/10_ml4t_library_ecosystem.ipynb) calls Engineer's registry to inspect feature metadata. The [Book Guide](../book-guide/index.md) distinguishes that direct use from related workflows. Use this page when you are choosing features, validating metadata, or building registry-driven workflows instead of hardcoding indicator names. @@ -182,9 +180,8 @@ feature_catalog.lookback("sma", period=50) ## See It In The Book -- Ch7 `10_ml4t_library_ecosystem.py` for registry inspection and catalog search -- [Book Guide](../book-guide/index.md) for how discovery connects to feature - computation, labeling, and dataset preparation +- [The ml4t Library Ecosystem](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/10_ml4t_library_ecosystem.ipynb) calls the registry and computes selected features. +- [Book Guide](../book-guide/index.md) records the pinned revision and task mapping. ## Next Steps diff --git a/docs/user-guide/features.md b/docs/user-guide/features.md index 5405e63..acfe48d 100644 --- a/docs/user-guide/features.md +++ b/docs/user-guide/features.md @@ -24,10 +24,7 @@ Standalone cross-asset utilities such as beta, rolling correlation, and cointegration are documented below, but they are not part of the 120-feature registry count. -> **Book**: *ML for Trading, 3rd ed.* — Ch8 notebooks (`01_price_volume_features.py` through `04_fundamentals_macro_calendar.py`) build features manually to explain the economics. Case studies (ETFs, US Equities Panel, CME Futures) then use `compute_features()` in production pipelines. - -Use the [Book Guide](../book-guide/index.md) for the full notebook-to-API map -across Chapters 7-9 and the case studies. +The pinned book notebooks combine direct Engineer calls with manual teaching implementations. Start with [The ml4t Library Ecosystem](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/10_ml4t_library_ecosystem.ipynb) for `compute_features()`, then use the [Book Guide](../book-guide/index.md) to check each notebook's exact relationship to the API. ## Computation API @@ -85,7 +82,7 @@ result = compute_features(df, [ Unknown parameters, repeated output names, and output names that replace input columns raise `ValueError` before feature execution. -> **Book**: Ch7 `10_ml4t_library_ecosystem.py` demonstrates all three input formats on SPY data, including a comparison between library and manual RSI implementations. +> **Book**: [The ml4t Library Ecosystem](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/10_ml4t_library_ecosystem.ipynb) calls `compute_features()` with names and parameter dictionaries, and compares the library RSI with a manual implementation. ## Category Reference @@ -114,7 +111,7 @@ Price momentum and oscillator indicators. Most produce bounded (normalized) outp | `trix` | Triple Exponential Average | Yes | No | 30 | | `cmo` | Chande Momentum Oscillator | Yes | -100 to 100 | 14 | | `ultosc` | Ultimate Oscillator | Yes | 0-100 | 7/14/28 | -| `bop` | Balance of Power | Yes | -1 to 1 | — | +| `bop` | Balance of Power | Yes | -1 to 1 | - | | `imi` | Intraday Momentum Index | No | 0-100 | 14 | | `aroon` | Aroon (up/down) | Yes | 0-100 | 14 | | `aroonosc` | Aroon Oscillator | Yes | -100 to 100 | 14 | @@ -122,7 +119,7 @@ Price momentum and oscillator indicators. Most produce bounded (normalized) outp | `ppo` | Percentage Price Oscillator | Yes | No | 12/26 | | `sar` | Parabolic SAR | Yes | No | 0.02/0.2 | -> **Book**: Ch8 `01_price_volume_features.py` constructs momentum indicators on ETF data, explaining the economic rationale for each. ETFs and US Equities Panel case studies use these in `03_features.py`. +> **Book**: [Price and Volume Feature Families](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/08_financial_features/01_price_volume_features.ipynb) calls Engineer for selected indicators and derives others manually. [ETFs: Feature Engineering](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/case_studies/etfs/03_financial_features.ipynb) calls individual Engineer feature functions in a case-study pipeline. ### Trend (10 indicators) @@ -165,7 +162,7 @@ Volatility estimators ranging from simple (ATR) to advanced (GARCH). Includes ra **Efficiency ranking**: Yang-Zhang > Garman-Klass ~ Rogers-Satchell > Parkinson > Close-to-Close. See Molnar (2012) for theoretical efficiency ratios. -> **Book**: Ch9 `08_garch_volatility.py` and `09_har_rough_volatility.py` compare volatility estimators on real data. Ch8 `01_price_volume_features.py` covers range-based estimators with efficiency analysis. +> **Book**: [Price and Volume Feature Families](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/08_financial_features/01_price_volume_features.ipynb) calls Engineer's volatility functions and compares selected estimators on ETF data. ### Microstructure (15 indicators) @@ -189,7 +186,7 @@ Market microstructure features from De Prado (2018) and empirical market microst | `volume_synchronicity` | Volume synchronicity | | `weighted_mid_price` | Weighted mid price | -> **Book**: Ch8 `02_microstructure_features.py` builds microstructure features from tick and minute data. The NASDAQ-100 Microstructure case study (`03_features.py`) implements Kyle's Lambda, Amihud, and VPIN manually for pedagogical purposes — the ml4t-engineer implementations are production-ready equivalents. +> **Book**: [Microstructure Features](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/08_financial_features/02_microstructure_features.ipynb) calls Engineer's tick-rule and liquidity functions while teaching the wider workflow. [NASDAQ-100 Microstructure: Feature Engineering](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/case_studies/nasdaq100_microstructure/03_financial_features.ipynb) calls `amihud_illiquidity` and implements other case-study features separately. ### ML Features (14 indicators) @@ -212,7 +209,7 @@ Features designed specifically for machine learning pipelines. | `time_decay_weights` | Exponential time decay | No | | `ffdiff` | Fractional differencing | No | -> **Book**: Ch8 `04_fundamentals_macro_calendar.py` covers feature construction patterns including lag features and calendar encodings. +> **Book**: [Slow Features and Context](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/08_financial_features/04_fundamentals_macro_calendar.ipynb) calls Engineer's `cyclical_encode` and teaches the point-in-time data joins manually. ### Risk (6 indicators) @@ -251,7 +248,7 @@ Multi-asset relationship features. These are standalone functions in `ml4t.engin These are called directly (not via `compute_features`) since they require multi-asset DataFrames. -> **Book**: Ch8 `03_structural_cross_instrument_features.py` constructs cross-asset features. Ch9 `14_panel_features.py` applies cross-sectional features to equity panels. +> **Book**: [Structural and Cross-Instrument Features](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/08_financial_features/03_structural_cross_instrument_features.ipynb) calls `beta_to_market`. [Panel Features](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/09_model_based_features/14_panel_features.ipynb) calls several cross-asset functions and combines them with manual analysis. ### Regime (4 indicators) @@ -264,7 +261,6 @@ Market regime detection features. All produce bounded outputs suitable for direc | `fractal_efficiency` | Price path efficiency | 0-1 | | `trend_intensity_index` | Trend strength | 0-100 | -> **Book**: Ch9 `11_hmm_regimes.py` and `13_regime_as_feature.py` apply regime detection to equity indices. ### Statistics (14 indicators) @@ -305,7 +301,7 @@ Statistical features including TA-Lib standard and rolling distribution metrics. | `ad` | Accumulation/Distribution | Yes | | `adosc` | A/D Oscillator | Yes | -> **Book**: ETFs case study `03_features.py` uses volume features in a multi-asset pipeline alongside momentum and volatility. +> **Book**: [ETFs: Feature Engineering](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/case_studies/etfs/03_financial_features.ipynb) calls Engineer's volume, momentum, trend, volatility, and regime feature functions. ### Math (3 indicators) @@ -352,7 +348,7 @@ print(feature_catalog.tags()) See the dedicated [Feature Discovery guide](discovery.md) for complete examples. -> **Book**: Ch7 `10_ml4t_library_ecosystem.py` explores the registry metadata for RSI, ATR, and Garman-Klass, then demonstrates `feature_catalog.search()` and filtered listing. +> **Book**: [The ml4t Library Ecosystem](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/10_ml4t_library_ecosystem.ipynb) inspects registry metadata for RSI, ATR, and Garman-Klass before computing features. ## YAML Configuration @@ -428,17 +424,16 @@ Invalid parameters raise `ValueError` with the valid parameter names. - **Polars-native**: All computations use Polars expressions for automatic parallelism - **Numba JIT**: Numerical kernels (volatility estimators, microstructure) are Numba-accelerated -- **Throughput**: ~480K indicator calculations/second, 11M rows/second streaming -- **TA-Lib parity**: RSI computed at ~1x TA-Lib speed via Polars native implementation -- **Dependency ordering**: `compute_features` resolves feature dependencies via topological sort +- `compute_features` resolves feature dependencies before execution. +- LazyFrame input remains lazy until you collect the result. +- Measure throughput with your feature set, row count, grouping columns, and hardware. ## See It In The Book -- Ch8 `01_price_volume_features.py` through `04_fundamentals_macro_calendar.py` for - the main feature-engineering concepts -- Ch7 `10_ml4t_library_ecosystem.py` for the config-driven `compute_features` API -- Case-study `03_features.py` workflows for production usage -- [Book Guide](../book-guide/index.md) for the full chapter and case-study map +- [The ml4t Library Ecosystem](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/10_ml4t_library_ecosystem.ipynb) calls the registry and `compute_features()`. +- [Price and Volume Feature Families](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/08_financial_features/01_price_volume_features.ipynb) combines Engineer calls with manual teaching implementations. +- [ETFs: Feature Engineering](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/case_studies/etfs/03_financial_features.ipynb) calls individual Engineer functions in a case-study pipeline. +- [Book Guide](../book-guide/index.md) records the pinned revision and relationship for every selected notebook. ## Next Steps diff --git a/docs/user-guide/fractional-differencing.md b/docs/user-guide/fractional-differencing.md index e206937..0d79b5a 100644 --- a/docs/user-guide/fractional-differencing.md +++ b/docs/user-guide/fractional-differencing.md @@ -1,6 +1,6 @@ # Fractional Differencing -Fractional differencing (FFD) produces stationary time series while preserving long-range memory — the key insight from De Prado (2018, Chapter 5). Standard first-differencing (d=1) achieves stationarity but destroys predictive signal; fractional differencing finds the minimum d that passes stationarity tests. +Fractional differencing (FFD) seeks a stationary series while retaining more long-range dependence than first differencing. `find_optimal_d()` searches for the minimum tested order that passes the configured stationarity threshold. Use the [Book Guide](../book-guide/index.md) for the surrounding Chapter 9 workflow and the case-study files that use FFD in production pipelines. @@ -21,7 +21,7 @@ The goal: find the smallest d where the ADF test rejects the null hypothesis of ## Core Functions -### `ffdiff` — Apply Fractional Differencing +### `ffdiff` - Apply Fractional Differencing ```python @@ -55,7 +55,7 @@ assert len(result) == len(ffd_series) == len(df) **How it works**: FFD applies a weighted sum of lagged values where weights are derived from the fractional binomial expansion. Weights decay geometrically, and the `threshold` parameter truncates negligibly small weights for efficiency. Weights are cached via `@lru_cache` and the inner loop is Numba-accelerated. -### `find_optimal_d` — Find Minimum Stationary d +### `find_optimal_d` - Find Minimum Stationary d ```python from ml4t.engineer.features.fdiff import find_optimal_d @@ -81,7 +81,7 @@ print(result) A high correlation (>0.90) means most of the predictive information is preserved. -### `fdiff_diagnostics` — Full Diagnostic Report +### `fdiff_diagnostics` - Full Diagnostic Report ```python from ml4t.engineer.features.fdiff import fdiff_diagnostics @@ -169,7 +169,7 @@ result = df.with_columns( ## Asset-Class Guidelines -Typical optimal d values (these are starting points — always validate on your data): +Typical optimal d values (these are starting points - always validate on your data): | Asset Class | Typical d Range | Notes | |-------------|----------------|-------| @@ -199,9 +199,9 @@ for symbol in ["SPY", "QQQ", "IWM"]: ## See It In The Book -- Ch9 `03_fractional_differencing.py` for the memory-stationarity tradeoff -- ETFs and US Equities Panel `04_temporal.py` workflows for production usage -- [Book Guide](../book-guide/index.md) for the full chapter and case-study map +- [Fractional Differencing](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/09_model_based_features/03_fractional_differencing.ipynb) calls all three documented helpers while teaching the statistical method. +- [ETFs: Model-Based Features](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/case_studies/etfs/04_model_based_features.ipynb) calls `ffdiff` inside a walk-forward workflow. +- [Book Guide](../book-guide/index.md) records the pinned revision and task mapping. ## Next Steps diff --git a/docs/user-guide/labeling.md b/docs/user-guide/labeling.md index 2034666..7f1c26c 100644 --- a/docs/user-guide/labeling.md +++ b/docs/user-guide/labeling.md @@ -14,12 +14,9 @@ ML4T Engineer provides 7 labeling methods for supervised learning in finance, im | Meta-labeling | `meta_labels()` + `compute_bet_size()` | Bet sizing for primary model | | Calendar-aware | `calendar_aware_labels()` | Session-break handling for futures | -All methods return a Polars DataFrame with standardized output columns. Performance is ~50,000 labels/second via Numba-accelerated kernels. +All methods return a Polars DataFrame with standardized output columns. -> **Book**: *ML for Trading, 3rd ed.* — Ch7 `03_label_methods.py` walks through all 7 methods on real ETF data with visualizations. All case study `02_labels.py` notebooks apply these methods in production pipelines. - -Use the [Book Guide](../book-guide/index.md) for the broader mapping from Chapter 7 -and case-study `02_labels.py` files to the production labeling APIs. +[Label Engineering Methods](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/03_label_methods.ipynb) calls Engineer's labeling APIs on ETF data. The ETF case-study label notebook builds labels without Engineer, so the [Book Guide](../book-guide/index.md) classifies it as a related workflow rather than a library example. ## Choosing a Method @@ -196,7 +193,7 @@ config = LabelingConfig.triple_barrier( With `trailing_stop=True`, the lower barrier moves up as the trade moves in favor. This reduces the time spent in losing positions. -> **Book**: Ch7 `03_label_methods.py` applies triple-barrier labeling to SPY with visualization of barrier touches. +> **Book**: [Label Engineering Methods](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/03_label_methods.ipynb) calls `triple_barrier_labels` and visualizes barrier touches. ## ATR-Based Dynamic Barriers @@ -242,7 +239,7 @@ null labels. High and low prices determine intrabar touches. | Regime changes (calm → volatile) | ATR barriers avoid premature stops | | Futures with varying contract sizes | ATR normalizes across contracts | -> **Book**: CME Futures case study `02_labels.py` applies ATR barriers on ES, NQ, and CL futures with session-aware horizons. +> **Book**: [Label Engineering Methods](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/03_label_methods.ipynb) calls Engineer's ATR-barrier labeling API. ## Rolling Percentile Labels @@ -294,7 +291,7 @@ result = rolling_percentile_multi_labels( # Produces: label_long_p90_h5, label_long_p95_h5, label_long_p90_h10, ... ``` -> **Book**: Ch7 `03_label_methods.py` compares rolling percentile labels against triple-barrier on SPY. ETFs case study `02_labels.py` uses percentile labels in its production pipeline. +> **Book**: [Label Engineering Methods](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/03_label_methods.ipynb) calls Engineer's percentile and triple-barrier functions. [ETFs: Label Engineering](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/case_studies/etfs/02_labels.ipynb) constructs its case-study labels without calling Engineer. ## Fixed Time Horizon @@ -350,7 +347,7 @@ remain available. A supplied trend-scanning config controls all four numerical s Constant windows have null outputs. Exact nonconstant linear fits use the largest finite float as the signed `t_value`. -> **Book**: Ch7 `03_label_methods.py` demonstrates trend scanning alongside triple-barrier and percentile methods, showing how the optimal horizon varies with market conditions. +> **Book**: [Label Engineering Methods](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/03_label_methods.ipynb) calls `trend_scanning_labels` alongside the barrier and percentile methods. ## Meta-Labeling & Bet Sizing @@ -411,7 +408,7 @@ result = apply_meta_model( | `"sigmoid"` | `2 / (1 + exp(-scale * (p - 0.5))) - 1` | Smooth, differentiable | | `"discrete"` | `1 if p >= threshold else 0` | Binary position sizing | -> **Book**: Ch7 `03_label_methods.py` implements the complete meta-labeling workflow: primary model signals → meta-labels → bet sizing on SPY. +> **Book**: [Label Engineering Methods](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/03_label_methods.ipynb) calls the meta-labeling and bet-sizing APIs on SPY. ## Calendar-Aware Labels @@ -511,7 +508,7 @@ stats = compute_label_statistics(df, label_col="label") # "positive_ratio", "negative_ratio", "neutral_ratio"} ``` -> **Book**: Ch7 `03_label_methods.py` demonstrates sequential bootstrap applied to triple-barrier labels, showing how it reduces effective sample size while improving independence. +> **Book**: [Label Engineering Methods](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/03_label_methods.ipynb) calls the sample-weighting and sequential-bootstrap APIs. ## Time-Based Durations @@ -549,11 +546,10 @@ from ml4t.engineer.labeling.utils import ( Time-based horizons require a `timestamp_col` in the input DataFrame. -## Performance +## Runtime behavior -- **Speed**: ~50,000 labels/second (Numba-accelerated) -- **Memory**: Efficient vectorized implementation via Polars -- **Accuracy**: Exact match with AFML reference (validated at 1e-10 tolerance against mlfinpy) +Benchmark with your row count, horizon, labeling method, and hardware. Runtime and +memory use depend on those inputs. ## Best Practices @@ -569,10 +565,9 @@ Time-based horizons require a `timestamp_col` in the input DataFrame. ## See It In The Book -- Ch7 `03_label_methods.py` for the full comparison of labeling methods -- Ch7 `04_minimum_favorable_adverse_excursion.py` for barrier behavior analysis -- Case-study `02_labels.py` workflows, especially CME Futures for ATR barriers -- [Book Guide](../book-guide/index.md) for the full chapter and case-study map +- [Label Engineering Methods](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/03_label_methods.ipynb) calls Engineer's labeling APIs. +- [ETFs: Label Engineering](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/case_studies/etfs/02_labels.ipynb) is a related workflow that does not call Engineer. +- [Book Guide](../book-guide/index.md) records the pinned revision and task mapping. ## Next Steps diff --git a/docs/user-guide/ml-readiness.md b/docs/user-guide/ml-readiness.md index 231ffa3..fa5e666 100644 --- a/docs/user-guide/ml-readiness.md +++ b/docs/user-guide/ml-readiness.md @@ -2,7 +2,7 @@ This guide explains the `normalized` field in feature metadata and how to prepare features for machine learning models. -> **Book**: *ML for Trading, 3rd ed.* — Ch8 `01_price_volume_features.py` compares normalized vs non-normalized features on real ETF data, including preprocessing strategies for each type. +[Price and Volume Feature Families](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/08_financial_features/01_price_volume_features.ipynb) calls Engineer's registry and selected feature functions, then combines them with manual teaching implementations. The [Book Guide](../book-guide/index.md) records the pinned revision and relationship. Use this page when you need to decide which features can go straight into a model and which ones should pass through a preprocessing step first. diff --git a/docs/user-guide/preprocessing.md b/docs/user-guide/preprocessing.md index 6449c44..d0e5c10 100644 --- a/docs/user-guide/preprocessing.md +++ b/docs/user-guide/preprocessing.md @@ -187,9 +187,9 @@ X_train, X_test, y_train, y_test = builder.train_test_split(train_size=0.8) ## See It In The Book -- Ch7 `02_preprocessing_pipeline.py` for split-aware preprocessing -- [ML Readiness](ml-readiness.md) for deciding which features need scaling first -- [Book Guide](../book-guide/index.md) for the full Chapter 7 workflow map +- [Preprocessing Pipeline](https://github.com/stefan-jansen/machine-learning-for-trading/blob/d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb/07_defining_the_learning_task/02_preprocessing_pipeline.ipynb) calls Engineer's `StandardScaler` and teaches the wider cleaning workflow. +- [ML Readiness](ml-readiness.md) explains which features need scaling. +- [Book Guide](../book-guide/index.md) records the pinned revision and task mapping. ## Next Steps