Skip to main content
Dedicated endpoints provide isolated, configurable deployments of supported models and their performance templates. Use the control plane to create and manage deployments, and the data plane to run inference through OpenAI-compatible APIs. With dedicated endpoints, you control:

Region

Choose where your deployment runs to optimize latency and meet data residency requirements.

GPU configuration

Select GPU type and GPUs per replica to match your performance and throughput needs.

Autoscaling

Set minimum and maximum replicas to automatically scale capacity with traffic.

Lifecycle management

Create, update, stop, and delete deployments as your workloads evolve.

Key use cases:

  • predictable capacity
  • finetuned base model with custom weights
  • compliance / private infra
  • bigger control over deployment

Dedicated vs Public Endpoints Comparison

Learn more

Deploy via API

Deploy in UI

FAQ & Troubleshooting