TensorGrid
GPU cluster observability for ML training — live scheduler, thermal model, alerting.
Cluster heatmap
Per-GPU utilization across 4 nodes × 8 GPUs. Click any cell for node detail. Sim time 0:00.
GPU 0
GPU 1
GPU 2
GPU 3
GPU 4
GPU 5
GPU 6
GPU 7
node-0
node-1
node-2
node-3
idle100% util
Power & thermals
Cluster power draw and average die temperature, last 120 ticks.
Power (kW)Avg temp (°C)
Job queue
Waiting jobs in scheduling order (priority, then FIFO). 3 queued.
| Job | GPUs | Priority | Wait | |
|---|---|---|---|---|
| sdxl-distill-2 | 1 | P2 | 0:00 | |
| llama-3-70b-sft-0 | 1 | P4 | 0:00 | |
| whisper-large-finetune-1 | 4 | P4 | 0:00 |
Running jobs
0 jobs training now.
Nothing running — the queue is empty and no jobs are training.
GPUs allocated
0/32
0% of cluster
Avg temp
58.9°C
peak 59.4°C
Power draw
16.0 kW
cluster total
Throughput
0 tok/s
0 jobs training
Simulation controls
Synthetic training jobs arrive automatically.
Fills every node with P2 jobs — watch thermals and queue pressure react.
Submit job
Enqueues into the live scheduler.
4
180s
Alerts
0 active · 0 total fired
No alerts yet. Saturate the cluster to trigger thermal and queue-pressure alerts.
Recently completed
0 jobs finished this session.
No completions yet.