Skip to main content

Nemotron 3 Ultra now available on AI Gateway

By Steven Van ·

The open Nvidia model offers a 1M token context window, up to 350 tokens per second, and up to 30% lower cost on agentic tasks.

Nvidia's Nemotron 3 Ultra is now available through Vercel's AI Gateway. It's an open Mixture-of-Experts reasoning model with a 1M token context window, built for multi-turn agent workflows: planning, tool use, sub-agent delegation, and error recovery.

Nvidia puts throughput at up to 350 tokens per second, with up to 30% lower cost on agentic tasks. To use it, set the model to nvidia/nemotron-3-ultra-550b-a55b in the AI SDK. Usage across models can be tracked on the AI Gateway's model leaderboard.

Vercel
Vercel
The platform for frontend developers — deploy, preview, and scale web apps and AI agents with zero config.
View Vercel →

Read the original announcement →

Read Nemotron 3 Ultra now available on AI Gateway on Creators Toolbox