GPUscale

gpuscale.net
v4.7
Local Open-Source GPU VRAM calculator for self-hosted LLM fleets.
Private by default: projects save to this browser only
Results · computed live from your inputs
Summary
Memory ledger
Time to first token
0ms
Per-user speed
0tok/s
Aggregate throughput
0tok/s
Mean user latency
0s
SLO compliance
Deployment topology & fleet map · every model on every GPU
Resilience buys survival, not speed: the guaranteed figures are what remains during the failure you chose to cover. Idle spares add nothing day-to-day; active/active sites add burst headroom you should not plan sustained load on.
Speed per user as demand growsmore simultaneous calls = slower streams
per-user tok/s aggregate tok/s memory limit (red dashes) speed target (amber dashes) below target meets target operating point
Speed per user vs conversation lengthlonger context = slower streams
per-user tok/s speed target (amber dashes) below target meets target current context
Request latency anatomy
Reading of this configuration
Recommendations