Palo Alto Networks has launched an always-on offensive security service that uses Anthropic’s Claude Mythos 5, OpenAI’s GPT-5.6-Cyber and open-weight models to find vulnerabilities, test whether they can be exploited and recommend fixes across enterprise environments.
Unit 42 Continuous Frontier AI Defense, announced September 22 and available worldwide by annual subscription, extends the company’s earlier point-in-time Frontier AI Exposure Analysis into continuous testing. It covers first- and third-party web applications, APIs, cloud infrastructure, source-code repositories, identities and network assets as those environments change.
The most revealing part of the launch is not that artificial intelligence can automate penetration testing. It is Palo Alto Networks’ finding that no single model caught more than 40% of vulnerabilities in a complex environment. Mythos 5 and GPT-5.6-Cyber overlapped on fewer than 10% of the exposures they identified, according to the company’s evaluation. In practical terms, the models behaved less like interchangeable engines and more like specialists with sharply different blind spots.
How the Unit 42 service works
The service begins with a full-estate baseline and then keeps testing as applications and infrastructure change. A proprietary harness routes individual offensive-security tasks to different models, combining their output with Unit 42 threat intelligence and human expertise. The objective is to move past a list of possible weaknesses and validate complete attack paths.
That distinction matters. A low-severity identity mistake, a skipped one-time-password check and a session-routing flaw may not look urgent in separate tickets. Unit 42 described one financial-sector engagement in which those three conditions could be chained into account takeover and payment fraud without action from the victim. Continuous testing is meant to find that relationship, show that the path is exploitable and give the affected teams a prioritized remediation plan.
The service can test attack paths across application code, exposed APIs, cloud resources and network assets. It produces code-level guidance and virtual-patch recommendations, and customers can pair it with Palo Alto Networks’ separate Frontier Virtual Patching service when a vendor patch does not yet exist. Palo Alto Networks says customer source code and telemetry are handled through zero-data-retention architectures and are not retained or used to train public models.
Why Palo Alto Networks uses several models
Security scanners are usually judged by coverage, accuracy and the work required to investigate their findings. Frontier models add another variable: two capable systems can inspect the same environment and discover different classes of weaknesses.
Palo Alto Networks’ reported overlap rate makes the multi-model design central to the product rather than a branding exercise. If each model sees a different slice of the attack surface, the orchestration layer must decide which model receives a task, how results are verified, when another model should challenge a finding and how duplicate or contradictory output is reconciled. Those choices will determine whether the service improves coverage or merely creates more findings for security teams to triage.
The company says it invested $17 million and spent six months developing and validating the method across its own systems and more than 100 Unit 42 customer engagements. During its internal deployment, Mythos-based scanning produced what Palo Alto Networks characterized as a year’s worth of traditional penetration-testing results in three weeks. It also reported 3.2 times more high- and critical-severity findings per product than legacy testing and a 51% reduction in mean time to remediate.
Those are vendor-reported results, not an independent benchmark. The comparison also needs context: a recurring service has more opportunities to inspect changing systems than a scheduled engagement, and faster remediation depends on customer staffing, ownership and release processes as much as discovery speed.
The findings challenge CVE-first security programs
Across customer assessments, Unit 42 says the service found exposures in every environment tested, with 37% rated high or critical. More than two-thirds of validated exposures in third-party applications had no known CVE.
That does not mean vulnerability databases are obsolete. CVEs remain essential for tracking affected products, patches and known exploitation. But a CVE-centered program can miss business-logic failures, unsafe combinations of legitimate features, identity gaps, misconfigurations and attack chains that do not map neatly to one published software defect.
This is where continuous adversary simulation could add useful evidence. Instead of ranking a weakness only by a generic severity score, a security team can ask whether it is reachable in its environment, what identity an attacker would gain, which systems become accessible next and what control breaks the chain most efficiently.
What buyers should ask before giving an AI tester access
An offensive testing service needs broad visibility to be useful. That creates a governance problem: the system may inspect proprietary source code, authentication flows, cloud configurations, internal APIs and sensitive telemetry. Palo Alto Networks’ zero-data-retention claim addresses one part of that risk, but prospective customers still need detailed answers about execution boundaries and oversight.
- Where does testing run? Buyers should identify which data and artifacts leave their environment, where prompts and results are processed, and what is logged for audit and incident review.
- What actions are permitted? The contract and technical controls should separate passive analysis, exploit validation and any action capable of changing production systems or touching customer data.
- How are models selected and changed? A multi-model service should disclose how model updates affect coverage, data handling, repeatability and previously accepted risk.
- How are false positives and unsafe tests contained? Human review, scoped credentials, rate limits, allowlists, maintenance windows and stop conditions should be visible parts of the engagement.
- Who owns remediation? Continuous discovery is valuable only when findings reach a named engineering or infrastructure owner, include reproducible evidence and are retested after the fix.
Palo Alto Networks has not published standard pricing; subscriptions vary by the OpenAI, Anthropic and open-weight models selected. That makes proof-of-value design especially important. A useful evaluation should compare verified attack paths, high-severity findings per tested asset, false-positive workload, remediation time and retest results against the organization’s current penetration-testing and exposure-management process.
Continuous pentesting changes the operating model
The service arrives as powerful cyber models are being distributed through gated partner programs rather than broad public access. Unit 42 is positioning itself as the controlled layer between those models and enterprise systems: it supplies the harness, threat context, human supervision and remediation workflow around capabilities that customers may not be able to access directly.
For security teams, the durable change is cadence. Annual or quarterly penetration tests produce a snapshot, while cloud environments, APIs, identities and application releases change daily. Always-on testing can narrow that gap, but only if discovery, validation and remediation operate as one system. Otherwise, continuous testing will simply create a continuous backlog.
The launch gives enterprises a concrete way to test whether frontier AI can improve offensive-security coverage without surrendering control of the process. The less flattering lesson from Palo Alto Networks’ own numbers is that no model is comprehensive. Buyers should evaluate the service as a managed testing and remediation program, not as an autonomous replacement for security engineers.