Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

GPU Monitor — ML Training Dashboard

Real-time web dashboard for monitoring GPU/system metrics and ML training processes.

Python FastAPI React License

Architecture

┌──────────────────────────────────────────────────────┐
│                   Browser (React)                     │
│  ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌─────────┐ │
│  │ CPU/RAM  │ │   GPU    │ │ Process  │ │ Status  │ │
│  │  Charts  │ │  Charts  │ │  Table   │ │  Badge  │ │
│  └────┬─────┘ └────┬─────┘ └────┬─────┘ └────┬────┘ │
│       └─────────────┴────────────┴─────────────┘     │
│                        ▲ WebSocket (JSON)             │
└────────────────────────┼─────────────────────────────┘
                         │
┌────────────────────────┼─────────────────────────────┐
│  FastAPI Backend       │                              │
│  ┌─────────────────────┴──────────────────┐          │
│  │         WebSocket Broadcast Loop       │          │
│  │         (every 2s, configurable)       │          │
│  └────┬──────────────────────┬────────────┘          │
│       ▼                      ▼                       │
│  ┌──────────┐         ┌────────────┐                 │
│  │collector │         │  notifier  │                 │
│  │ psutil   │         │   httpx    │──► Discord      │
│  │ pynvml   │         │  webhooks  │──► Telegram     │
│  └──────────┘         └────────────┘                 │
└──────────────────────────────────────────────────────┘

Features

  • Live Streaming — CPU, RAM, GPU util, VRAM, temperature, power via WebSocket
  • ML Process Tracker — Detects python, torch, tensorflow, jupyter processes
  • Auto GPU Fallback — Gracefully switches to Mock Mode when no NVIDIA GPU detected
  • Threshold Alerts — Configurable webhook alerts to Discord/Telegram
  • Dark Mode UI — Glass-morphism cards, animated gauges, responsive layout
  • Auto-Reconnect — Exponential backoff WebSocket reconnection

Quick Start

Prerequisites

  • Python 3.10+
  • Node.js 18+
  • NVIDIA GPU + drivers (optional — falls back to mock mode)

Backend

cd backend
pip install -r requirements.txt
uvicorn app.main:app --reload --host 0.0.0.0 --port 8000

Frontend

cd frontend
npm install
npm run dev

Open http://localhost:5173 — Vite proxies WebSocket and API calls to the backend.

Configuration

Edit backend/config.json:

{
  "polling_interval_seconds": 2,
  "alert_cooldown_seconds": 60,
  "thresholds": {
    "gpu_util_percent": 95,
    "vram_percent": 90,
    "gpu_temp_celsius": 85,
    "cpu_percent": 95,
    "ram_percent": 90
  },
  "webhooks": {
    "discord_url": "https://discord.com/api/webhooks/...",
    "telegram_bot_token": "123456:ABC-DEF...",
    "telegram_chat_id": "987654321"
  }
}

Config is hot-reloadable — changes take effect on the next polling cycle.

WebSocket Payload

Connect to ws://localhost:8000/ws to receive JSON every 2 seconds:

{
  "timestamp": 1700000000.0,
  "cpu": {
    "cpu_percent": 23.5,
    "cpu_count": 8,
    "ram_used_gb": 12.3,
    "ram_total_gb": 32.0,
    "ram_percent": 38.4
  },
  "gpus": [{
    "index": 0,
    "name": "NVIDIA RTX 4090",
    "gpu_util_percent": 87,
    "vram_used_mb": 20480.0,
    "vram_total_mb": 24576.0,
    "vram_percent": 83.3,
    "temperature_c": 72,
    "power_w": 285.0,
    "mock": false
  }],
  "processes": [{
    "pid": 12345,
    "name": "python",
    "cmdline": "python train.py --epochs 100",
    "cpu_percent": 45.2,
    "ram_mb": 4096.0,
    "uptime_seconds": 3600
  }],
  "gpu_available": true,
  "gpu_count": 1
}

REST Endpoints

Endpoint Method Description
/api/status GET GPU availability and server status
/api/snapshot GET One-shot metrics snapshot
/ws WS Streaming metrics (JSON every 2s)

Project Structure

├── backend/
│   ├── app/
│   │   ├── __init__.py
│   │   ├── main.py          # FastAPI + WebSocket broadcast
│   │   ├── collector.py     # psutil + pynvml metrics
│   │   └── notifier.py      # Webhook alert dispatcher
│   ├── config.json           # Thresholds + webhook URLs
│   └── requirements.txt
├── frontend/
│   ├── src/
│   │   ├── components/       # MetricCard, Charts, ProcessTable, StatusBadge
│   │   ├── hooks/            # useWebSocket (auto-reconnect)
│   │   ├── App.jsx
│   │   ├── main.jsx
│   │   └── index.css
│   ├── package.json
│   └── vite.config.js
└── README.md

Cross-Platform

  • Linux: Full GPU support via NVIDIA drivers + pynvml
  • Windows: Same — requires NVIDIA drivers. Falls back to Mock Mode if not present.
  • macOS: CPU/RAM only (no NVIDIA support), automatic Mock Mode

License

MIT

About

Lightweight, real-time web dashboard and alert system for monitoring GPU/VRAM utilization, CPU metrics, and ML process states during model training jobs. Supports NVIDIA GPUs with CPU mock fallback.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages