GPU allocation (accelerator hosts)
GPUs are the expensive, scarce resource everyone fights over. AI labs and render studios will pay almost anything to get one -- but you only have so many cards, and a job that can't find a free GPU simply waits in line. That queue, not slowness, is what makes these customers anxious.
GPU demand is spiky and deadline-driven: a render shop is quiet for weeks, then needs every card you own the night before a delivery. Buy too few and good customers wait and grumble; buy too many and pricey silicon sits idle. The squeeze gets worse during a GPU-shortage event, when the labs and studios all grow at once.
Detailed explanation
Accelerator hosts (GPU 2x / 4x / 8x) pack differently from general compute: capex is dominated by the card, power-per-U is high, and the workload is compute-bound (rare, fat requests) rather than NIC-bound. A VM with required_accelerator=Gpu only fits on a host with a free GPU; there's no oversubscribing a physical card the way you oversell vCPU.
Over-commit GPUs and the wait queue -- not latency -- becomes the thing customers leave over. Demand spikes during a GPU-shortage macro event, when GPU-hungry archetypes (AI Lab, VFX Studio, Genomics Lab) grow faster. Watch the host-category mix: GPU hosts are a deliberate, capital-heavy bet, sized to the accelerator-demanding tenants you've actually signed.