TensorGrid

GPU cluster observability for ML training — live scheduler, thermal model, alerting.

Cluster heatmap

Per-GPU utilization across 4 nodes × 8 GPUs. Click any cell for node detail. Sim time 0:00.

GPU 0
GPU 1
GPU 2
GPU 3
GPU 4
GPU 5
GPU 6
GPU 7
node-0
node-1
node-2
node-3
idle100% util

Power & thermals

Cluster power draw and average die temperature, last 120 ticks.

Power (kW)Avg temp (°C)

Job queue

Waiting jobs in scheduling order (priority, then FIFO). 3 queued.

JobGPUsPriorityWait
sdxl-distill-21P20:00
llama-3-70b-sft-01P40:00
whisper-large-finetune-14P40:00

Running jobs

0 jobs training now.

Nothing running — the queue is empty and no jobs are training.

GPUs allocated
0/32
0% of cluster
Avg temp
58.9°C
peak 59.4°C
Power draw
16.0 kW
cluster total
Throughput
0 tok/s
0 jobs training

Simulation controls

Synthetic training jobs arrive automatically.

Fills every node with P2 jobs — watch thermals and queue pressure react.

Submit job

Enqueues into the live scheduler.

4
180s

Alerts

0 active · 0 total fired

No alerts yet. Saturate the cluster to trigger thermal and queue-pressure alerts.

Recently completed

0 jobs finished this session.

No completions yet.