Gemini 3.1 Flash-Lite hits general availability
By Steven Van ·
The fastest, cheapest Gemini 3 model targets high-volume agent, translation, and data tasks; JetBrains and Gladly are already using it in production.
Google made Gemini 3.1 Flash-Lite generally available, calling it the fastest and most cost-efficient model in the Gemini 3 family. It's aimed at high-volume agent tasks, translation, and simple data processing.
Google cites early production use at JetBrains (real-time code completion in its IDE assistant and Junie agent), Gladly (customer service across SMS, WhatsApp, and Instagram, at roughly 60% lower cost than thinking-tier models, with p95 latency around 1.8 seconds for full replies), Astrocade (multimodal safety checks, comment translation, and thumbnail prompt refinement for its game-generation platform), and krea.ai (prompt expansion in its Nodes tool). Ramp, AlphaSense, and OffDeal are also using it for latency-sensitive and high-volume workflows. Full details are in the announcement.
