Migrate from Closed Models
If you’re moving from Claude, GPT, or Gemini, here are open-source alternatives available on Cerebras.
Explore the full model catalog at Dedicated Endpoints, or get started with the Quickstart.
Documentation Index
Fetch the complete documentation index at: /llms.txt
Use this file to discover all available pages before exploring further.
Find the right open-source model for your workload on Cerebras, including alternatives for Claude, GPT, and Gemini.
| Category | Use Case | Large (>200B) | Medium (20B–200B) | Small (<20B) | Why Cerebras? |
|---|---|---|---|---|---|
| Code & Development | Code generation & reasoning | Kimi K2.7 Code GLM 5.1 MiniMax M2.5 | Qwen 3.8 27B (Public) Gemma 4 31B | Reason through entire codebases, including complex requirements, dependencies, and edge cases, without disrupting developer workflows. | |
| Code completion & bug fixing | MiniMax M2.5 Kimi K2.7 Code | Qwen 3.8 27B (Public) Gemma 4 31B | Generate, critique, and repair code at the speed you type. | ||
| Terminal tasks | GLM 5.1 Kimi K2.7 Code MiniMax M2.5 | Qwen 3.8 27B (Public) GPT OSS 120B (Public) | Agents can reason between commands, inspect results, and continue acting while the experience remains interactive. | ||
| Long-horizon autonomous coding | Kimi K2.7 Code GLM 5.1 MiniMax M2.5 | Run long agent loops with strong open-source models, reducing hours of work to minutes. | |||
| AI-Powered Apps | Agents with tool use | Kimi K2.7 Code GLM 5.1 MiniMax M2.5 | Qwen 3.8 27B (Public) GPT OSS 120B (Public) | Make more tool calls, complete more plan, act, and observe loops, and attempt more recoveries in a single user turn. | |
| Professional workflows | Kimi K2.6 GLM 5.1 | Qwen 3.8 27B (Public) Gemma 4 31B | Complete more planning, comparison, and verification steps in the same time window. | ||
| Summarization | Kimi K2.6 GLM 5.1 | Gemma 4 31B GPT OSS 120B (Public) | Process longer contexts and synthesize information without making users wait. | ||
| Conversational chat | MiniMax M2.5 | GPT OSS 120B (Public) | Create responsive, natural conversations for consumer-facing assistants. | ||
| Low-latency NLU & extraction | Gemma 4 31B | Extract and validate structured data fast enough for production workflows. | |||
| Vision & Multimodal | Vision & document understanding | Kimi K2.7 Code Kimi K2.6 | Qwen 3.8 27B (Public) Gemma 4 31B | Richer reasoning across text, image, and other inputs while keeping multimodal workflows responsive. |
| Provider | Closed Source | Use Case | Open Source Alternatives |
|---|---|---|---|
| Claude | Claude Opus 4.8 | Multi-hour coding agents, complex multi-file refactors, and difficult reasoning tasks where end-to-end correctness is critical | Kimi K2.6 Kimi K2.7 Code GLM 5.1 |
| Claude Sonnet 5 | Daily coding and IDE-adjacent agent loops that balance speed and intelligence | Primary: GLM 5.1 Fallbacks: Kimi K2.7 Code, Qwen 3.8 27B (Public) | |
| Claude Haiku 4.5 | Customer support, classification, extraction, short-form generation, and subagents in multi-agent systems | Gemma 4 31B GPT OSS 120B (Public) MiniMax M2.5 | |
| OpenAI GPT | GPT 5.6 Terra | Balanced reasoning and coding for subagents in agentic systems | Kimi K2.6 Kimi K2.7 Code GLM 5.1 |
| GPT 5.6 Luna | Classification, extraction, ranking, and low-cost coding subagents | Qwen 3.8 27B (Public) | |
| GPT 5.4 Nano/Mini | Balanced reasoning and coding, subagents in agentic systems, and structured tasks | MiniMax M2.5 Gemma 4 31B GPT OSS 120B (Public) | |
| Gemini | Gemini 3.1 Pro | Image understanding for coding, document analysis, and scientific reasoning | Kimi K2.6 GLM 5.1 Qwen 3.8 27B (Public) |
| Gemini 3.5 Flash Lite | Low-latency multimodal chat and tool use for real-time experiences | Gemma 4 31B GPT OSS 120B (Public) MiniMax M2.5 Qwen 3.8 27B (Public) |
Was this page helpful?
