Skip to main content

GLM 5.2 Fast via Wafer now available on AI Gateway

By Steven Van ·

Vercel adds GLM 5.2 Fast on a new provider, Wafer, which it says roughly doubles serverless throughput over other providers.

GLM 5.2 Fast is now available on Vercel's AI Gateway, served through a new provider called Wafer. Vercel's own benchmarking across small-context, large-context, and tool-call scenarios found Wafer delivers twice the throughput of other providers serving GLM-5.2 on serverless, leading on decode and end-to-end speed for sustained generation.

In that testing, GLM 5.2 Fast on Wafer reached 170+ tok/s on small-context requests and 200+ tok/s on large-context ones. Developers can use it by setting the model to "zai/glm-5.2-fast" in the AI SDK. As with other AI Gateway models, pricing reflects the provider's own rates with no markup, and there is no platform fee on inference, including Bring Your Own Key requests.

Vercel
Vercel
The platform for frontend developers — deploy, preview, and scale web apps and AI agents with zero config.
View Vercel →

Read the original announcement →

Read GLM 5.2 Fast via Wafer now available on AI Gateway on Creators Toolbox