Single place for setup, training, and running the dashboard/demo. Keep this open while you work; no need to hunt other READMEs.
- A GPU market simulator (
env/gpu_market_env.py) with price-sensitive demand and queues. - Two policies: rule-based baseline (
policies/baseline.py) and GA-learned (policies/ga_policy.py+ JSON params). - An objective file (
config/objective.yaml) that steers reward weighting (revenue, waits, fairness, market activity). - A Streamlit dashboard (
dashboard/side_by_side.py) to watch baseline vs GA under identical demand.
- Python 3.10+ recommended.
- Install deps once:
pip install -r requirements.txt
- Edit
config/objective.yamlto set weights and constraints. Typical keys:weights.revenue,weights.low_wait_time,weights.fairness,weights.market_activityconstraints.min_price,constraints.max_price,constraints.max_avg_wait
- Scenario presets live in
config/scenarios.yamland are selectable from the dashboard sidebar.
- Train v2 (5-parameter policy):
python -m policies.train_ga v2 --generations 50 --population_size 30 --output policies/ga_best_params.json
- Output:
policies/ga_best_params.json(orga_best_params_gen*.jsonif you set a different name). The dashboard auto-loads the most recentga_best_params*.json. - If you just want an untrained/random policy, set
--generations 0or usepolicies/ga_untrained.json.
- Start:
python -m streamlit run dashboard/side_by_side.py
- In the sidebar:
- Pick a scenario (loads weights into
objective.yaml). - Choose GPUs, base arrival rate, episode length.
- Start to run baseline and GA in lockstep (shared seed).
- Optional: enable logging to CSV (
logs/side_by_side_*.csv). - Demand triggers: surge/shock hit both systems at once.
- Pick a scenario (loads weights into
- On completion, you see:
- Performance Summary (active periods only): averages of reward/queue/wait/price/util, plus revenue per step from averages.
- Current State: run averages with deltas to the last step; totals for reward and revenue; market-death indicators; GA params displayed.
- Charts: price, queue, utilisation, wait time, cumulative reward for both policies.
- Revenue is price × running jobs (always positive, sums to Total Revenue).
- Reward is the weighted objective: revenue term (normalised) minus penalties (waits, fairness, market death), plus small activity bonuses. It can go negative. In the side-by-side view, reward accumulation halts after market death to make the freeze obvious.
- Quick demo:
pip install -r requirements.txtpython -m streamlit run dashboard/side_by_side.py- Select a scenario (e.g., revenue-heavy), start, and watch baseline vs GA.
- Retrain for a new scenario:
- Adjust
config/objective.yaml(or pick a scenario in the dashboard and click “Apply Scenario & Retrain GA”). python -m policies.train_ga v2 --generations 50 --population_size 30 --output policies/ga_best_params.json- Rerun the dashboard; it will pick up the latest params.
- Adjust
dashboard/side_by_side.py— demo entrypoint.env/gpu_market_env.py— simulator and reward.policies/ga_policy.py— GA policy logic.policies/baseline.py— rule-based baseline.policies/ga_best_params*.json— trained GA params.config/objective.yaml,config/scenarios.yaml— objective weights and presets.
- The env tracks consecutive zero arrivals; penalties ramp after one zero.
- In the dashboard, once market death is flagged, reward accumulation is frozen so the cumulative line flattens.
pytest tests/ -v- Logs:
logs/side_by_side_*.csv(enable from sidebar). Safe to prune.
pip install -r requirements.txt
python -m policies.train_ga v2 --generations 50 --population_size 30 --output policies/ga_best_params.json
python -m streamlit run dashboard/side_by_side.pyThe evaluator is run inside a zkVM to produce cryptographic proofs that a given policy achieved a given fitness under a fixed config and seed. This enables verifiable on-chain claims without re-running the simulation. Requires Rust, SP1, and Linux/WSL.
→ zk/README.md — Build, prove, verify, and tests.
- docs/end-to-end-testing.md — ZK prove/verify demo flow (WSL)
- docs/PROJECT_DETAILS.md — Full PoC overview (market + GA + ZK)
- docs/tickets.md — AIBlock L1 integration backlog
- docs/ — Architecture, GA, reward, validation, dashboard guides
- SCALE: 1_000_000 fixed-point for core maths.
- RNG:
random.Random(seed)passed into env/GA eval; Poisson via deterministic Knuth sampler. - Tie-breaks: GA selection sorted by (fitness desc, param-hash asc).
- Serialization: traces logged as canonical JSONL (sorted keys) when enabled.
- Bounds: price clamp uses
objective.yamlmin/max; GA param bounds are enforced in policy/train code.