Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

UTMM-Lite GPU Market (PoC)

Single place for setup, training, and running the dashboard/demo. Keep this open while you work; no need to hunt other READMEs.

What this is

  • A GPU market simulator (env/gpu_market_env.py) with price-sensitive demand and queues.
  • Two policies: rule-based baseline (policies/baseline.py) and GA-learned (policies/ga_policy.py + JSON params).
  • An objective file (config/objective.yaml) that steers reward weighting (revenue, waits, fairness, market activity).
  • A Streamlit dashboard (dashboard/side_by_side.py) to watch baseline vs GA under identical demand.

Prereqs

  • Python 3.10+ recommended.
  • Install deps once:
    pip install -r requirements.txt

Configure the objective

  • Edit config/objective.yaml to set weights and constraints. Typical keys:
    • weights.revenue, weights.low_wait_time, weights.fairness, weights.market_activity
    • constraints.min_price, constraints.max_price, constraints.max_avg_wait
  • Scenario presets live in config/scenarios.yaml and are selectable from the dashboard sidebar.

Train the GA

  • Train v2 (5-parameter policy):
    python -m policies.train_ga v2 --generations 50 --population_size 30 --output policies/ga_best_params.json
  • Output: policies/ga_best_params.json (or ga_best_params_gen*.json if you set a different name). The dashboard auto-loads the most recent ga_best_params*.json.
  • If you just want an untrained/random policy, set --generations 0 or use policies/ga_untrained.json.

Run the dashboard (side-by-side demo)

  • Start:
    python -m streamlit run dashboard/side_by_side.py
  • In the sidebar:
    • Pick a scenario (loads weights into objective.yaml).
    • Choose GPUs, base arrival rate, episode length.
    • Start to run baseline and GA in lockstep (shared seed).
    • Optional: enable logging to CSV (logs/side_by_side_*.csv).
    • Demand triggers: surge/shock hit both systems at once.
  • On completion, you see:
    • Performance Summary (active periods only): averages of reward/queue/wait/price/util, plus revenue per step from averages.
    • Current State: run averages with deltas to the last step; totals for reward and revenue; market-death indicators; GA params displayed.
    • Charts: price, queue, utilisation, wait time, cumulative reward for both policies.

Interpreting rewards vs revenue

  • Revenue is price × running jobs (always positive, sums to Total Revenue).
  • Reward is the weighted objective: revenue term (normalised) minus penalties (waits, fairness, market death), plus small activity bonuses. It can go negative. In the side-by-side view, reward accumulation halts after market death to make the freeze obvious.

Typical workflows

  • Quick demo:
    1. pip install -r requirements.txt
    2. python -m streamlit run dashboard/side_by_side.py
    3. Select a scenario (e.g., revenue-heavy), start, and watch baseline vs GA.
  • Retrain for a new scenario:
    1. Adjust config/objective.yaml (or pick a scenario in the dashboard and click “Apply Scenario & Retrain GA”).
    2. python -m policies.train_ga v2 --generations 50 --population_size 30 --output policies/ga_best_params.json
    3. Rerun the dashboard; it will pick up the latest params.

Key files (runtime)

  • dashboard/side_by_side.py — demo entrypoint.
  • env/gpu_market_env.py — simulator and reward.
  • policies/ga_policy.py — GA policy logic.
  • policies/baseline.py — rule-based baseline.
  • policies/ga_best_params*.json — trained GA params.
  • config/objective.yaml, config/scenarios.yaml — objective weights and presets.

Notes on market death

  • The env tracks consecutive zero arrivals; penalties ramp after one zero.
  • In the dashboard, once market death is flagged, reward accumulation is frozen so the cumulative line flattens.

Tests

pytest tests/ -v

Housekeeping

  • Logs: logs/side_by_side_*.csv (enable from sidebar). Safe to prune.

Minimal commands to remember

pip install -r requirements.txt
python -m policies.train_ga v2 --generations 50 --population_size 30 --output policies/ga_best_params.json
python -m streamlit run dashboard/side_by_side.py

ZK proofs

The evaluator is run inside a zkVM to produce cryptographic proofs that a given policy achieved a given fitness under a fixed config and seed. This enables verifiable on-chain claims without re-running the simulation. Requires Rust, SP1, and Linux/WSL.

→ zk/README.md — Build, prove, verify, and tests.

Documentation

Protocol constants (determinism / ZK prep)

  • SCALE: 1_000_000 fixed-point for core maths.
  • RNG: random.Random(seed) passed into env/GA eval; Poisson via deterministic Knuth sampler.
  • Tie-breaks: GA selection sorted by (fitness desc, param-hash asc).
  • Serialization: traces logged as canonical JSONL (sorted keys) when enabled.
  • Bounds: price clamp uses objective.yaml min/max; GA param bounds are enforced in policy/train code.

About

UTMM-Lite GPU Market – simulator, GA policies, ZK proofs

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages