Hosting Qwen3 235B on Blackwell
Perplexity published research on serving Qwen3 235B MoE models on NVIDIA GB200 NVL72 racks — disaggregated prefill/decode, tensor parallelism, and NVLink t
Perplexity published research on serving Qwen3 235B MoE models on NVIDIA GB200 NVL72 racks — disaggregated prefill/decode, tensor parallelism, and NVLink throughput wins. GB200 is a major step up over Hopper for high-throughput inference on large MoE models.