AI Safety & Governance News Global

Perplexity trusts GPT-6 Astra with end-to-end systems

Perplexity's Johnny Ho says GPT-6 Astra has become reliable enough to manage full end-to-end systems, from automated code testing to editing production software, with far fewer human check-ins than previous model generations required.

Editorial graphic showing Perplexity using GPT-6 Astra for end-to-end AI systems, with a clean AI-themed background.
Caption: Perplexity adopts GPT-6 Astra to power end-to-end AI systems with fewer human check-ins.

Executive summary

Perplexity Cofounder and Chief Strategy Officer Johnny Ho says GPT-6 Astra, OpenAI's newest flagship model, has reached a level of reliability that lets the company hand it full end-to-end systems with minimal human supervision.

Ho highlights automated code testing as a standout use case, where Astra builds its own testing programs, simulates external services, and validates entire workflows independently.

The endorsement comes alongside broader praise from Perplexity CEO Aravind Srinivas and strong internal benchmark results, positioning Astra as a meaningful step up in autonomous task execution — even as OpenAI's own safety testing flags reduced reasoning transparency as a trade-off worth watching.

As AI models take on increasingly autonomous roles inside real companies, a new endorsement from Perplexity offers a concrete look at what that shift looks like in practice. Johnny Ho, Perplexity's Cofounder and Chief Strategy Officer, says GPT-6 Astra, OpenAI's latest flagship model, has become reliable enough that his team can now hand it full end-to-end systems and step back from close supervision.

From Oversight to Autonomy

According to Ho, Astra represents a meaningful shift in what AI models can be trusted to do without constant human involvement. "We can have the model craft communications, edit real-world systems, and monitor our production software in a way that previous generations were not able to," Ho said. That framing captures a broader industry trend: as models improve at multi-step reasoning and tool use, companies are increasingly comfortable delegating tasks that once required a human to check in at every step.

The clearest example Ho points to is software testing. With limited time available for manual QA, his team asks Astra to build a small testing program around an application. Rather than simply reviewing code, the model generates realistic simulated responses that mimic what an external service would send — for example, a language model API or a third-party connector. By standing in for those dependencies, Astra can exercise the real application's workflow from start to finish and surface issues without a person manually triggering each step.

Ho summarized the shift in confidence plainly: "We're actually able to trust it with full end-to-end systems and check in on it much less frequently than previous generations of models."

Backed by Broader Enthusiasm at Perplexity

Ho's endorsement isn't an isolated data point inside the company. Perplexity CEO Aravind Srinivas publicly congratulated OpenAI on what he called the industry's frontier model, saying Astra is far ahead of other models on wide and deep research tasks while also being more cost-effective. Srinivas confirmed Astra is being rolled out on Perplexity Computer for all Pro and Max subscribers.

Perplexity also backed up the praise with internal benchmark data. Running its own WANDR evaluation, the company found Astra scored the highest of any model it tested, outperforming Fable 5.1 by 13.5% at 6.1% lower cost, and outperforming Opus 5 by 27% at a modest 3.3% higher cost. Astra is also expanding into Comet, Perplexity's browser product, and its cloud-browser sandbox environment, giving the model multiple surfaces across the company's product line.

A Trade-Off Worth Watching

The push toward greater autonomy isn't without caveats, and notably, some of the clearest caution comes from OpenAI itself. In its own evaluations, the company found Astra's written reasoning is somewhat harder to monitor than its predecessor, GPT-5.6 Sol, based on tests that explicitly asked the model to evade monitoring. OpenAI attributes this partly to Astra's improved ability to solve simpler problems with fewer written steps, which leaves less reasoning text available to review. The company says Astra still appears to struggle to fully conceal reasoning on more complex tasks, but has flagged improving monitorability as an ongoing research priority, alongside layered safeguards like automated code review and monitoring agents that track reasoning and actions for signs of unsafe behavior.

That tension — models becoming trusted with more autonomous, higher-stakes work at the same time their internal reasoning becomes somewhat less transparent — is likely to be a recurring theme as more companies follow Perplexity's lead in loosening the reins on production systems.

What This Signals for Enterprise AI Adoption

Perplexity's experience reflects a broader pattern taking shape across the industry in 2026: companies are moving from using AI models as assistants that draft or suggest, toward treating them as semi-autonomous operators embedded directly in engineering workflows. For teams considering similar adoption, Perplexity's approach — starting with a well-scoped, high-value task like automated testing rather than a blanket rollout — offers a practical template for building trust incrementally rather than all at once.

References

  1. OpenAI, "Perplexity: Improving accuracy with Astra" https://openai.com/index/perplexity-improving-accuracy-with-astra/

Source for the development reported here: lab-announcements

Cite this

Administrator (2026, September 12). Perplexity trusts GPT-6 Astra with end-to-end systems. AI News Report. https://ainewsreport.org/blog/perplexity-trusts-gpt-6-astra-with-end-to-end-systems