Skip to main content
Get started with the world’s fastest inference. Already familiar with LLM APIs? Skip straight to the API reference or try the playground.
1

Set up your API key

Visit the Cloud Console and sign up or log in. Navigate to API Keys in the left nav bar to create a key. See API Keys for details.Set your API key as an environment variable so you don’t have to pass it with every request:
Confirm the variable is set:
export and $env: set the variable for the current shell only. setx on Windows persists the variable, but you must open a new terminal window for it to take effect. To persist on macOS or Linux, add the export line to your ~/.zshrc, ~/.bashrc, or equivalent shell profile.
2

Install the SDK

Install the Cerebras SDK for your language of choice. You can also call the API directly with cURL (see Step 3).
3

Make your first API request

Run the following code to send a chat completion request:
You should see a response like:

Common Errors

  • 401 UnauthorizedCEREBRAS_API_KEY isn’t set in the shell running your code. Re-run the echo command above in the same terminal to confirm. On Windows after setx, open a new terminal.
  • 404 model not found — The model ID is misspelled, deprecated, or not available on your account. See the full list of public models on the Models page.
  • 429 Too Many Requests — You’ve hit a rate limit. Free accounts have lower per-minute limits than paid accounts. See Rate limits for current quotas and how to request an increase.
For all status codes, see the error reference.

Next Steps