Perplexity News
Portable Computer Comes to Windows RTX PCs
September 14, 2026
Perplexity's Portable Computer is now available in the Windows app on compatible NVIDIA RTX PCs. Run the agent harness, models, and tools locally so work with files and connected apps stays on your PC, while frontier cloud models can be used when needed.
View on X
Tag Coworkers in Computer Sessions
September 11, 2026
Perplexity Computer users can now tag coworkers in sessions with @ mentions and share the session with them. The feature is available on the web for all Computer users.
View on X
Desktop and Mobile Previews for Computer Websites
September 9, 2026
Websites created with Computer now include desktop and mobile previews in the artifact side pane. Switch views, refresh the preview, edit the site, or leave a comment without leaving the artifact.
View on X
Custom connectors: Bring your own MCP server
September 1, 2026
Register a remote MCP server once on your Project connectors page. Perplexity stores the server's credential, so your application does not need to store or send it with each request. Use the generated connector ID with type: "connector" in Agent API requests. Custom connectors are available to all Projects and support API-key or no authentication, with Streamable HTTP or SSE transport. See Add a custom connector.
Sign in with Perplexity for the remote MCP server
September 1, 2026
The remote Perplexity MCP Server now supports OAuth. Add https://api.perplexity.ai/mcp to any MCP client that supports OAuth, such as claude.ai, Claude Code, Cursor, or VS Code, and sign in with your Perplexity account instead of pasting an API key. You choose which API organization to bill during sign-in. API keys continue to work for clients without OAuth support.
Gemini 3.8 Flash
September 1, 2026
The Agent API now supports google/gemini-3.8-flash. Promotional pricing through December 31, 2026 is \$0.75 per million uncached-input tokens, \$0.075 per million cached-input tokens, and \$3.75 per million output and reasoning tokens. See the Agent API Models reference.
GLM 5.3 Flash
September 1, 2026
The Agent API and Router API now support perplexity/glm-5.3-flash at \$0.15 per million uncached-input tokens, \$0.03 per million cached-input tokens, and \$0.50 per million output tokens. See the Agent API Models reference or the Router model catalog.
Optimizing On-Device Inference for Apple Silicon
September 1, 2026
Perplexity describes Lily, a lightweight local inference engine built for Apple silicon and Qwen3.6-35B-A3B, with separate optimizations for prefill and decode. The post walks through how Qwen's architecture opens model-specific optimizations on Apple silicon, where further tuning stops paying off, and an end-to-end comparison against MLX-LM. Lily is to be open-sourced.
GLM 5.3 in Perplexity Computer
August 28, 2026
GLM 5.3 is available in Perplexity Computer, aimed at long-context multimodal agent work. Perplexity says it beat GLM 5.2 on WANDR, its own benchmark for large-scale research with evidence attached.
View on X
GPT-5.6 Terra and Luna in Perplexity Computer
August 6, 2026
GPT-5.6 Terra and Luna are now live in Perplexity Computer. Terra is the new default model for all Computer subagents and is also available as an orchestrator model, while Luna serves as the primary model for scheduled automations.
View on X
Voice Mode Isolates Your Voice
August 6, 2026
Perplexity's voice mode now isolates your voice from everything else. A new audio focus model filters out background noise, music, and other voices, so you can talk to Perplexity in a noisy environment like a crowded coffee shop. Live now for voice conversations and dictation on macOS and Windows for all Computer users.
View on X
GLM 5.3
August 1, 2026
The Agent API and Router API now support perplexity/glm-5.3 at \
ews: [.40 per million uncached-input tokens, \$0.26 per million cached-input tokens, and \$4.40 per million output tokens. See the Agent API Models reference or the Router model catalog.
Prompt caching for presets
August 1, 2026
Agent API presets now use stable prompt cache keys automatically, allowing independent requests with the same preset to reuse the shared prompt prefix (system prompt and tool definitions). No request changes are required, and an explicit prompt_cache_key still overrides the preset default. This can reduce costs by about 5% for applications that use presets frequently, depending on cache utilization. The current preset values include each key for frozen configurations.
Fast preset updated
August 1, 2026
The Agent API fast preset now uses openai/gpt-5.6-luna with minimal reasoning effort and priority processing. Dynamic fast preset requests pick up the change automatically. If you use a frozen configuration, update the model and reasoning effort and set service_tier to priority. Priority processing uses 2× the model's standard token prices.
Gemini 3.7 Flash
August 1, 2026
The Agent API and Router API now support google/gemini-3.7-flash at launch pricing of \$0.375 per million input tokens, \$0.0375 per million cached-input tokens, and \
ews: [.875 per million output tokens. See the Agent API Models reference.
Grok 4.6
August 1, 2026
The Agent API now supports xai/grok-4.6, xAI's latest flagship reasoning and agentic model. See pricing in the Agent API Models reference.
NVIDIA Nemotron 3 Ultra
August 1, 2026
The Agent API and Router API now support perplexity/nemotron-3-ultra-550b-a55b at \$0.25 per million input or cached-input tokens and \$2.50 per million output tokens. See the Agent API Models reference or the Router model catalog.
NVIDIA Nemotron 3.5 Lightning
August 1, 2026
The Agent API and Router API now support perplexity/nemotron-3.5-lightning-30b-a3b, a fast, efficient open-weight reasoning model, at \$0.0115 per million input tokens, \$0.00115 per million cached-input tokens, and \$0.17 per million output tokens. See the Agent API Models reference or the Router model catalog.
DeepSeek V4 Flash 0731
August 1, 2026
The Agent API and Router API now support perplexity/deepseek-v4-flash-0731, a fast, efficient open reasoning model with a 1M-token context window. See pricing in the Agent API Models reference or the Router model catalog.
Introducing Projects
July 29, 2026
Perplexity launched Projects, an evolution of Spaces. Projects are hubs for ongoing work in Computer — a single place to manage, create, and collaborate on tasks with a shared file system and persistent memory.
View on X
Claude Opus 5 in Perplexity + Computer
July 24, 2026
Claude Opus 5 is now available in Perplexity and Perplexity Computer. Perplexity evaluated it against six other models on WANDR. It outperformed all but Fable 5, while being 57% cheaper.
View on X
Perplexity CLI: Web Search for Coding Agents
July 24, 2026
The Perplexity CLI is now available, giving coding agents the ability to search the web. Install it as an agent skill by pointing your agent at the pplx-cli SKILL.md on GitHub.
View on X
Agent API: New Models
July 1, 2026
The Agent API added support for several new models this month, all with direct first-party token pricing. See the full list in the Agent API Models reference. Claude Opus 5 The Agent API now supports anthropic/claude-opus-5. GPT-5.6 Family The Agent API now supports openai/gpt-5.6-sol, openai/gpt-5.6-terra, and openai/gpt-5.6-luna, giving you access to the latest GPT-5.6 models. Gemini Flash Models The Agent API now supports google/gemini-3.6-flash and google/gemini-3.5-flash-lite, adding Google's latest fast and cost-efficient models. Grok 4.5 The Agent API now supports xai/grok-4.5, xAI's flagship coding and agentic model. Kimi K3 The Agent API now supports perplexity/kimi-k3, Moonshot AI's flagship reasoning model. See the Agent API Models reference for pricing and model-specific settings. Retired: Gemini 3.1 Flash Lite Preview google/gemini-3.1-flash-lite-preview has been retired: Google removed the underlying preview model from its API. Requests for this id now return a model not supported error. Use google/gemini-3.1-flash-lite instead — the stable successor at the same pricing. See the Agent API Models reference.
GPT-5.6 price cuts and Sol Fast mode
July 1, 2026
GPT-5.6 Luna now costs \$0.20 per million input tokens and \
ews: [.20 per million output tokens. GPT-5.6 Terra now costs \$2 per million input tokens and \
ews: [2 per million output tokens. GPT-5.6 Sol now supports Fast mode at 2× standard token pricing; send service_tier: "priority" to use it.
New: Router API
July 1, 2026
Unified access to open-weight models hosted by Perplexity through a single endpoint — with your existing Perplexity API key. Key features: OpenAI Chat Completions and Anthropic Messages compatibility: switching is a base-URL change Automatic health-based routing and failover across model deployments Per-token pricing at each model's published rates, with no per-request fees Get started with the Router API →
Remote MCP Server
July 1, 2026
The Perplexity MCP Server is now available as a remote server hosted by Perplexity at https://api.perplexity.ai/mcp. Connect any MCP client that supports Streamable HTTP using your API key as a bearer token, with no local installation and nothing to update. Tools and behavior are identical to the local server, and usage is billed to your API key at standard API pricing.
Low preset updated
July 1, 2026
The Agent API low preset now uses openai/gpt-5.6-luna with minimal reasoning effort and a 32,768-token maximum output. If you use a frozen configuration, update these values to match the current preset. Dynamic low preset requests pick up the change automatically.
Inline citations for research presets
July 1, 2026
The Agent API's search-backed presets include inline citations: fast cites claims drawn from search results with numbered citations such as [1], while low, medium, and high now cite claims drawn from tool results or provided source artifacts with source-typed citations such as [web:1]. After a successful tool call, the low, medium, and high presets include at least one citation in the final answer. The current preset values include the system prompts for frozen configurations.
MCP Server 1.0: now backed by the Agent API
July 1, 2026
The Perplexity MCP Server v1.0.0 moves its model-backed tools from the legacy Sonar models to Agent API presets: perplexity_ask uses the fast preset, perplexity_reason uses medium, and perplexity_research uses high. Tool names and response shapes are unchanged, and all three tools are faster and cheaper on average than their Sonar predecessors. Long-running research now streams progress to MCP clients that request it, and cancelling an MCP request cancels the underlying run. The strip_thinking and reasoning_effort parameters were removed from the tool schemas (the Agent API emits no think tags, and presets manage reasoning effort); clients still sending them are ignored gracefully.
Sonar Chat Completions migration
July 1, 2026
More models, tools, and research-backed presets: Sonar Chat Completions is now Agent API. Migration guide here.
Claude Fable 5 as Orchestrator
July 1, 2026
Claude Fable 5 is once again available in Perplexity as an orchestrator model.
View on X
Computer for Counsel
June 24, 2026
Perplexity introduced Computer for Counsel — Computer now connects the research databases, document tools, and matter-management systems lawyers use daily, pulling citable sources from midpage, LegalZoom, Docusign, NetDocuments, and more. Available for Pro and Max subscribers.
View on X
Brain in Computer
June 21, 2026
Introducing Brain in Computer — a continuously learning memory system. Every task on Computer plugs into a context graph built by Brain, making Computer more stateful with every run. Available as a research preview for all Perplexity Max subscribers.
View on X
Nemotron 3 Ultra
June 6, 2026
Nemotron 3 Ultra is now available for Pro and Max subscribers on Perplexity and Computer — NVIDIA's new open model built for long-running agents.
View on X
Main Street AI Accelerator
June 6, 2026
Perplexity launches the Main Street AI Accelerator with the U.S. SBA — committing $25M in Computer credits ($250 each for up to 100,000 eligible companies).
View on X
Computer Windows
June 6, 2026
Personal Computer is coming to Windows — orchestrates across apps and files on your machine, rolling out first to paying Max and Enterprise Max subscribers.
View on X
Hybrid Agentic Inference
June 6, 2026
Hybrid agentic inference is coming to Perplexity Computer — splits tasks between a local model on your machine and frontier models in the cloud.
View on X
Search as Code
June 6, 2026
Introducing Search as Code — a new search architecture for AI agents that writes Python to call the search stack directly instead of looping through function calls.
View on X
Microsoft Office Integration
June 6, 2026
Perplexity Computer is now available inside Microsoft Excel, Word, PowerPoint, and Outlook — orchestrate across work in the side panel.
View on X
Perplexity Computer Connects to Snowflake
May 15, 2026
Computer now connects to Snowflake. Run end-to-end work against live warehouse data and get answers with SQL, source tables, filters, and metrics. Build dashboards and automations for pipeline analysis, product usage, and customer segments.
View on X
Hosting Qwen3 235B on Blackwell
May 12, 2026
Perplexity published research on serving Qwen3 235B MoE models on NVIDIA GB200 NVL72 racks — disaggregated prefill/decode, tensor parallelism, and NVLink throughput wins. GB200 is a major step up over Hopper for high-throughput inference on large MoE models.
Personal Computer in Perplexity Mac App
May 8, 2026
Perplexity rolled out Personal Computer to all users in a new Mac app — runs tasks across local files, native apps, the web, and Perplexity servers.
View on X
Finance Search in the Agent API
May 6, 2026
Perplexity introduced Finance Search in the Agent API — programmatic access to live financial data, filings, and market context for agentic workflows.
Computer for Professional Finance
May 6, 2026
Perplexity launched Computer for Professional Finance — 35 dedicated finance workflows on top of Computer.
View on X
Perplexity Computer in Microsoft Teams
May 5, 2026
Perplexity Computer is now available inside Microsoft Teams. Run research, analysis, and document creation directly in your Teams workspace with the same capabilities as the standalone Computer.
View on X
GPT-5.5 in Perplexity
April 25, 2026
GPT-5.5 is now available on Perplexity for Max subscribers, and is rolling out as the default orchestration model in Computer for both Pro and Max subscribers.
View on X
Opus 4.7 Now Default in Perplexity
April 18, 2026
Claude Opus 4.7 is now the default orchestration model powering Computer. Also available for Max subscribers on web, iOS, and Android.
View on X
Personal Computer
April 17, 2026
Perplexity releases Personal Computer — integrates with the Perplexity Mac App for secure orchestration across your local files, native apps, and browser. Rolling out to all Max subscribers and waitlist members.
View on X