Cerebras

Wafer-scale chip inference β€” among the fastest tokens-per-second in the industry.

Cerebras runs inference on their custom Wafer-Scale Engine, delivering some of the highest tokens-per-second throughput available on Llama and GPT-OSS class models. US-hosted. The default pick when output throughput dominates the workload and US hosting is acceptable.

1 route2 modelsUSπŸ‡ΊπŸ‡Έ HQ United States
cerebras.ai

Models on Cerebras

Every model we route through Cerebras. Compare residency, zero retention, training posture, and price at a glance β€” full data-handling detail per route below.

ModelRegionZero data retentionTrainingContextInputOutput
USZero data retentionNo128K$2.25$2.75
USZero data retentionNo131K$0.35$0.75

Data handling per route

Cerebras hosts on 1 route. Each route has its own privacy posture, residency, and no-KYC terms. Postures are maintained by qynio with a last-verification timestamp.

United StatesπŸ‡ΊπŸ‡Έ

Zero data retention is on by default on Pay-as-you-go β€” no action required. No training on customer data. US; wallets; deposit available.

Zero data retention
On by default on Pay-as-you-go. Derived from the logging and moderation facts.
Training
No training on customer data.
Logging
None
Moderation
Not established
Caching
Not established
Subprocessor access
Not established
no-KYC deposit
deposit available
Transfer mechanism
wallets

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation