Skip to main content
This endpoint is not yet available. It is planned for a future release.
Dedicated GPU deployments with autoscaling, scale-to-zero, and per-deployment configuration. For enterprises that need guaranteed throughput, custom models, or data residency guarantees.

Planned endpoints

Planned autoscaling options

Serverless inference is the right choice for most developers. Deployments are for enterprises with guaranteed throughput requirements.
  • Models — Available models for deployment
  • Roadmap — feature timeline