Skip to main content
Models available with Cerebras Shared Inference can be used on the Free Trial and Pay as You Go tiers, subject to rate limits and pricing. For additional model families, reserved capacity, higher throughput, and production SLAs, see Dedicated Inference.
New here? Follow the Quickstart to make your first API call. To pick a model by use case, see the model selection guide. Select any model name below for full specs, capabilities, and per-tier limits.

Available Models

Looking for more models? Many additional model families are available through Dedicated Inference.

Model Compression

This section provides transparency about the compression state of each model available on our platform. We host a variety of open-source models from the community. We do not currently host pruned models with Shared Inference. All models served through Shared Inference are the original, unpruned versions. While we conduct research on pruning techniques like REAP (Router-weighted Expert Activation Pruning), these pruned models are shared with the research community on Hugging Face but are not available through Shared Inference. You can read more about REAP in our research blog. All Shared Inference models are unpruned. Cerebras uses selective weight-only quantization only during storage to preserve maximal quality. This means that the weights are stored in partial 16-bit / 8-bit / 4-bit, in-line with industry standards. For quality, sensitive layers are stored at full precision with dequantization on the fly, so operations are done in high precision. The activations, attention, and kv cache remain in full precision and unquantized.

Frequently Asked Questions

No. We are committed to serving the original weights for existing model IDs without modification. We do not alter model architectures via pruning on our hosted portfolio. If we explore additional compression techniques like pruning in the future, we will offer them as separate model variants with distinct model IDs. This keeps the difference transparent and lets you choose the variant that best fits your needs.
Our REAP pruned models are available on Hugging Face for research and experimentation purposes: Cerebras REAP Collection. These models demonstrate our pruning research but are not served through our production API.
Compression is an umbrella term for techniques that reduce model size or computational requirements. Common compression techniques include:
  • Quantization: Reducing the precision of numbers used to represent model weights (e.g., converting from FP16 to FP8). This reduces memory usage without changing the model’s architecture.
  • Pruning: Permanently removing parts of a model, like layers or experts, to reduce model size. This changes the model’s architecture and creates a different model.