This feature is in Private Preview. For access or more information, contact us or reach out to your account representative.
Key Concepts
- Model Architecture: The serving compatibility class for a model, such as
gpt-oss-120borzai-glm-4.7. Each endpoint supports a specific architecture, and custom weights must be compatible with that architecture. - Model Version: An immutable snapshot of model weights. Each upload creates a new version identified by an auto-incrementing integer, such as
1,2, or3. Versions can also have aliases likeproductionorv1-stable. - Deployment: A running model version and its serving configuration. A deployment belongs to an endpoint and can be replaced without changing the endpoint ID used by applications.
- Endpoint: A stable inference target identified by a unique endpoint ID, such as
my-org-gpt-oss-120b. The endpoint routes requests to one or more deployments. - Replica: One running copy of a deployment. Replica count determines the capacity available to the deployment.
The inference API uses the OpenAI-compatible
model field as a target selector. For Dedicated Inference, pass the endpoint ID in this field. In the Management API, model refers to a model version resource name where noted.Typical Workflow
- Upload model weights — Push your fine-tuned weights from S3 to Cerebras. The upload is asynchronous; you’ll receive a version ID to track progress.
- Check upload status — Poll the version status until the sync completes.
- Create a deployment: Once the upload is complete, deploy the version behind your endpoint.
-
Run inference: Make requests to your endpoint using the endpoint ID as the
modelfield. -
Iterate: Upload new versions as you fine-tune, assign aliases like
productionto track releases, and replace deployments without changing the endpoint ID.
Authentication
The Management API uses a separate API key from the standard inference API. You can find your Management API key under Management API keys on the API keys page in the Cerebras Cloud console.S3 Bucket Setup
Before uploading model weights, you need an S3 bucket configured with cross-account access to Cerebras. Your Cerebras representative will provide specific instructions, but the bucket policy must grant cross-account access:<your-bucket-name> with your S3 bucket name. The IAM role ARN (<cerebras-provided-iam-role-arn>) will be provided by Cerebras and enables secure cross-account access to your model weights.
