Nemotron 3 Ultra now available on AI Gateway
By Steven Van ·
The open Nvidia model offers a 1M token context window, up to 350 tokens per second, and up to 30% lower cost on agentic tasks.
Nvidia's Nemotron 3 Ultra is now available through Vercel's AI Gateway. It's an open Mixture-of-Experts reasoning model with a 1M token context window, built for multi-turn agent workflows: planning, tool use, sub-agent delegation, and error recovery.
Nvidia puts throughput at up to 350 tokens per second, with up to 30% lower cost on agentic tasks. To use it, set the model to nvidia/nemotron-3-ultra-550b-a55b in the AI SDK. Usage across models can be tracked on the AI Gateway's model leaderboard.

Vercel
The platform for frontend developers — deploy, preview, and scale web apps and AI agents with zero config.
View Vercel →