Skip to main content
Dedicated Inference reserves capacity for your organization. Send inference requests to a stable endpoint. The endpoint routes each request to a deployment. Use Dedicated Inference when you need predictable performance in production. Examples include real-time applications, customer-facing products, and high-volume pipelines. See supported models.

Key Benefits

Your deployments run on capacity reserved for your organization.
Reserved capacity provides predictable latency and throughput under load.
Deploy fine-tuned weights alongside standard model versions.
Tune capacity, draft models, model configuration, and quantization for your workload.
Use every Shared Inference capability, plus fine-tuning, weight management, and service tier controls.

Supported Models

Dedicated Inference supports many model families, parameter sizes, and variants, including -instruct and -thinking. You can also deploy custom weights and tune each deployment for your performance goals.
Multimodal models, including Gemma 4, accept text and image inputs through Dedicated Inference. See Image Inputs.

Features

Dedicated Inference includes all Shared Inference capabilities, plus:
  • Fine-tuning: Deploy custom model weights behind your endpoint.
  • Management API: Manage model versions, deployments, capacity, and endpoints through the API.
  • Batch API: Run large asynchronous workloads on reserved capacity.
  • Predicted Outputs: Reduce latency by supplying expected output content.
  • Service tiers: Prioritize requests to meet your SLA requirements.
  • Metrics: Monitor endpoint requests, tokens, latency, and health with Prometheus-compatible metrics.

Resource Model

Use these terms consistently when working with Dedicated Inference: An inference request uses these resources in order: API route → endpoint → deployment → replica.
For compatibility with OpenAI clients, the inference API uses the model field to select an inference target. For Shared Inference, use a model ID. For Dedicated Inference, use an endpoint ID.

Get Started

Dedicated Inference is available to enterprise customers. Contact us to discuss your workload.