Wafer

Tuned serving for open-weight models β€” GLM, Kimi and DeepSeek.

Wafer serves open-weight models on a stack it tunes per workload, profiling traffic to pick the serving engine, kernels and hardware for each deployment. US-hosted. The route to pick when you want tuned throughput on GLM, Kimi or DeepSeek and US hosting is acceptable.

1 route5 modelsUSπŸ‡ΊπŸ‡Έ HQ United States
wafer.ai

Models on Wafer

Every model we route through Wafer. Compare residency, zero retention, training posture, and price at a glance β€” full data-handling detail per route below.

ModelRegionZero data retentionTrainingContextInputOutput
USZero data retentionNo1M$3.00$12.75
USZero data retentionNo1M$1.19$4.40
USZero data retentionNo1M$0.10$0.35
USZero data retentionNo1M$1.26$3.96
USZero data retentionNo1M$0.10$0.25

Data handling per route

Wafer hosts on 1 route. Each route has its own privacy posture, residency, and no-KYC terms. Postures are maintained by qynio with a last-verification timestamp.

United StatesπŸ‡ΊπŸ‡Έ

Zero data retention is on by default on Pay-as-you-go β€” no action required. No training on customer data. US; wallets; deposit available.

Zero data retention
On by default on Pay-as-you-go. Derived from the logging and moderation facts.
Training
No training on customer data.
Logging
None
Moderation
Not established
Caching
Not established
Subprocessor access
Not established
no-KYC deposit
deposit available
Transfer mechanism
wallets

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation