Skip to main content

Capacity availability

Dedicated endpoint deployment depends on available capacity across different regions. Not all listed GPU types may be available at all times. Availability depends on current regional capacity.

Capacity after stops, incidents, and maintenance

Self-service dedicated endpoints are provisioned on Nebius AI Cloud on-demand (pay-as-you-go) capacity. This capacity is drawn from a shared pool and is allocated to your endpoint only while it is active — it is not a long-term reservation, and the minimum replicas guarantee applies only while the endpoint is running. When an endpoint stops, its GPU capacity is released back to the shared pool. This applies regardless of why the endpoint stopped:
  • You stop the endpoint manually
  • The endpoint is interrupted by an incident
  • The endpoint is stopped for scheduled or emergency maintenance
When the endpoint starts again, capacity is requested from the pool anew and allocated based on availability at that moment. There is no mechanism that holds your previous GPUs for the endpoint or returns them after an interruption. As a result:
  • After a restart, the endpoint may stay in Starting / Not ready, or reach only Partially ready, until sufficient capacity becomes available. See Lifecycle & Readiness Status.
  • During periods of high demand, or for scarce GPU types, restarting may take significantly longer than the initial deployment, and capacity may be temporarily unavailable.
This is the standard behavior of on-demand capacity in Nebius AI Cloud, on which self-service dedicated endpoints run: without a reservation, GPU resources are taken from a shared pool and returned to it when they stop — whether stopped by you or by a maintenance event — and recovery after maintenance or an incident does not restore the capacity your endpoint used before. For details on how on-demand and reserved capacity work, see Capacity reservations for Compute virtual machines in the Nebius AI Cloud documentation.
If your workload requires guaranteed capacity across restarts, incidents, and maintenance windows, contact our Sales team to reserve dedicated GPU capacity

Minimum replicas guarantee

Minimum replicas are reserved and non-preemptible and remain allocated to your deployment for as long as the endpoint is active. If the endpoint stops for any reason, this allocation is released — see the section above.

Maximum replicas scaling

Scaling above minimum replicas depends on available burst capacity and is not guaranteed indefinitely. Additional replicas may be reclaimed after scale-down and may not always be available again without sufficient capacity.

SLA

Self-service dedicated endpoints do not include a formal SLA unless covered by contract. Historical average monthly request success rate has been approximately 99.9%.
If you encounter capacity errors or are unable to scale beyond the minimum number of replicas, contact our Sales team to reserve dedicated GPU capacity