Region
Choose where your deployment runs to optimize latency and meet data residency requirements.
GPU configuration
Select GPU type and GPUs per replica to match your performance and throughput needs.
Autoscaling
Set minimum and maximum replicas to automatically scale capacity with traffic.
Lifecycle management
Create, update, stop, and delete deployments as your workloads evolve.
Key use cases:
- predictable capacity
- finetuned base model with custom weights
- compliance / private infra
- bigger control over deployment