Cerebras’s public inference documentation lists two models on offer to both trial users and customers paying by usage: OpenAI’s GPT OSS 120B and Alibaba’s Qwen 3.8 27B. Cerebras says GPT OSS 120B serves a 65,000 token context window on the free tier and 131,000 tokens on paid access, at a published throughput near 3,000 tokens per second. Qwen 3.8 27B gets 64,000 tokens free and 128,000 paid, running near 1,500 tokens per second by Cerebras’s own figures.

Cerebras states both models run unpruned, the full architecture as released, though it applies weight-only quantization down to 4-bit during storage. Sensitive layers stay at full precision with on-the-fly dequantization, and activations, attention and the KV cache remain unquantized throughout inference, per the documentation.

Larger model families, reserved capacity and production service-level agreements sit behind Cerebras’s separate Dedicated Endpoints tier, not the public catalog.

Cerebras’s own inference documentation, accessed September 4, 2026.