The Question

A large energy company runs Pentera across its enterprise network on a continuous basis. The automated pentesting platform tests lateral movement paths, Active Directory attack vectors, and credential exploitation across the full network scope every week. The security team has a well-maintained finding backlog, a clear picture of which attack paths are exploitable in the live environment, and strong MTTR metrics for the validation findings that Pentera surfaces.

The CISO presents to the board: the automated pentesting program has replaced the need for annual red team engagements. The platform covers more techniques, runs more frequently, and costs significantly less than a full red team. Budget for the manual engagement has been reallocated.

Six months later, a Mandiant red team hired as part of an acquisition due diligence process discovers a three-step attack path to the SCADA environment that Pentera never found. It involved a social engineering vector against an IT contractor, a misconfiguration in a cloud-to-OT gateway that was not in the automated pentesting scope, and a privilege escalation technique that leveraged a custom application's undocumented API endpoint. None of those three steps appears in Pentera's technique library.

The CISO's error was not deploying Pentera. It was concluding that continuous automated pentesting had made manual red team engagements redundant. The two tools operate at different discovery depths, on different time scales, and against different categories of attack paths.

Automated pentesting and red team engagements are not substitutes for each other — they operate at different time scales and discovery depths, and a CTEM program that uses only one is missing what the other uniquely provides.

Why This Matters Now

The 2025 MGM Resorts post-incident analysis — widely discussed in the security community following the company's public disclosures and subsequent litigation documents — illustrated this distinction at scale. The initial breach vector was a social engineering attack against the IT help desk: a threat actor called posing as an employee, obtained credential reset assistance, and used that access to begin lateral movement. No automated pentesting or BAS platform tests against live help desk agents.

Once inside, the threat actors used a combination of well-documented techniques — Okta session hijacking, Scatter Spider TTPs — that a mature BAS platform would have simulated. But the initial entry point was entirely outside the scope of any automated tool. The manual red team engagement MGM had conducted 18 months prior had, in fact, identified the help desk social engineering vector as a high-risk path. The finding was documented. The control improvement — enhanced identity verification for help desk requests — had been partially implemented but not fully operationalized before the breach.

The MGM case illustrates two realities simultaneously: manual red teams find attack paths that automated tools cannot, and the operational gap is not always in the discovery stage — sometimes it is in the mobilization of findings that the red team has already surfaced. Both realities are relevant to CTEM program design. The validation stage must include both automated and manual modalities, and the mobilization stage must ensure that red team findings receive the same remediation tracking discipline as VM findings.

What the CURVE™ Data Shows

The 2026 Stackcurve CTEM CURVE™ Report evaluated the validation stage of CTEM programs across 40 enterprise deployments, assessing the mix of automated pentesting, BAS, manual red team, and purple team modalities in use and their correlation with CTEM program maturity scores.

Programs rated Tier 1 in validation maturity universally combined continuous automated pentesting or BAS with at least one annual manual red team engagement. No Tier 1 program relied on automated tools alone. Programs that had replaced annual red teams entirely with BAS platforms rated Tier 2 or Tier 3 in validation maturity, primarily because the discovery depth for novel attack paths was insufficient.

Among automated pentesting vendors, Pentera rated Tier 1 in network and Active Directory simulation. Cymulate and AttackIQ rated Tier 1 in MITRE ATT&CK-aligned kill chain simulation. Among red team service providers evaluated, Mandiant (Google Cloud), CrowdStrike Services, Bishop Fox, and NCC Group rated Tier 1 in adversary emulation depth and CTEM integration quality — specifically, the structured handoff of red team findings into the CTEM remediation workflow.

Purple team service maturity was rated for its contribution to the mobilization stage. Bishop Fox, Mandiant, and Rapid7 Services rated Tier 1 in purple team exercise design that specifically builds security operations response capability against the attack paths CTEM has surfaced.

The full vendor rankings are in the 2026 Stackcurve CTEM CURVE™ Report — free to download.

The Gap Most Buyers Miss

The framing of "automated vs. manual" is itself a gap. The correct framing is "continuous and scalable" versus "deep and novel" — and a mature CTEM validation stage requires both.

What Automated Pentesting and BAS Provide

Continuous automated pentesting platforms like Pentera operate at a time scale and scope that no human red team can match. They test the full network scope, on a weekly or continuous cadence, against a library of known techniques derived from real attacker behavior. This produces three specific CTEM capabilities that manual engagements cannot replicate:

First, continuous validation: the attack surface changes daily. New vulnerabilities are disclosed. Cloud configurations drift. Developers deploy new services. Automated pentesting tests the current environment, not the environment as it existed during the last engagement six months ago.

Second, regression testing: when a finding is remediated, automated pentesting verifies that the remediation actually closed the path — not just that a patch was applied. This closes the loop that VM programs leave open.

Third, control validation at scale: BAS platforms test whether security controls are actually blocking the techniques they are configured to block, across the entire environment, continuously. This is the WAF syntax error detection capability discussed in Brief 12 — something no annual red team can provide with the same coverage frequency.

What Manual Red Teams Provide

A skilled adversary-emulating red team discovers attack paths that no automated tool would find, for a fundamental reason: human attackers combine technical, social, and physical vectors in ways that no technique library anticipates. The MGM help desk call is one example. Others include: a red teamer who notices that the badge reader on a server room door has a known firmware vulnerability and tests whether physical access enables a pivot to the internal network; a social engineering campaign that targets a specific executive's assistant to obtain calendar access and then uses that access to send a targeted phishing email that bypasses the email security gateway; a custom web application with an authentication bypass in a feature added by a development contractor that no CVE database covers because it was never assigned a CVE.

These attack paths are not theoretical. They are the attack paths that advanced threat actors use when they have moved beyond commodity exploitation. Red teams find them because they operate with the same goal-oriented, adaptive methodology as sophisticated attackers. Automated tools find what is in their libraries.

Purple Team: Building the Response Capability

Purple team exercises contribute to the CTEM mobilization stage in a way that neither VM nor BAS addresses. A purple team exercise has the red team execute attack techniques while the blue team defends in real time, with immediate feedback loops. This is not primarily a discovery exercise — it is a response capability development exercise. It builds the security operations team's ability to detect and respond to the attack paths that CTEM has identified as highest priority.

The CTEM program that lacks purple team exercises has identified attack paths and validated exploitability but has not tested whether security operations can actually contain a breach if those paths are traversed. That is a mobilization gap that VM, BAS, and red teams combined do not close.

Budget Architecture

The practical budget structure for an enterprise CTEM validation stage: a continuous automated pentesting or BAS platform in the range of $100K–$300K per year provides the continuous coverage layer. One annual red team engagement in the range of $150K–$500K depending on scope provides the novel path discovery layer. One to two purple team exercises per year at $50K–$150K each build the response capability. Total validation budget: $300K–$950K per year, depending on environment size and engagement scope. This is the architecture that consistently produces Tier 1 validation maturity in the CURVE™ research.

Questions Your Buying Team Should Be Asking

1. What is the scope of the automated pentesting platform's technique library, and how is that library updated when new adversary techniques emerge?

The value of an automated pentesting platform is directly proportional to the breadth and currency of its technique library. Ask specifically how the vendor tracks new attacker techniques — whether from MITRE ATT&CK updates, Mandiant M-Trends data, CrowdStrike Global Threat Report, or original research — and what the average latency is between a new technique's first documented use in the wild and its availability in the platform's simulation library.

2. How does the automated pentesting platform handle out-of-scope assets — specifically cloud-to-on-premises gateways, OT/IoT systems, and third-party contractor access paths?

Scope definition is one of the most common gaps in automated pentesting deployments. The assets most likely to provide novel attack paths to sensitive targets are often the ones that have been excluded from automated pentesting scope for operational reasons. Ask the vendor how it handles scope expansion and what the process is for including previously out-of-scope asset classes.

3. For red team engagements — what is your adversary emulation methodology, and how do you ensure the engagement tests against threat actors relevant to our industry and geography?

Red team quality varies significantly across service providers. The best engagements are designed against a specific threat profile — a named threat actor or a composite actor modeled on the groups that actually target the client's sector. Generic red team engagements that run standard kill chain scenarios without adversary-specific targeting produce findings that may not reflect the actual risk posture.

4. How do red team findings integrate into our CTEM remediation workflow — specifically, what deliverable format maps to our ITSM system, and how are findings tracked to closure?

Red team reports are frequently high-quality intelligence that disappears into a PDF and is never remediated. Ask service providers specifically what structured data format they deliver findings in and whether they support direct integration with your ITSM system for automated ticket creation and remediation tracking.

5. How does your purple team exercise design connect to our specific CTEM prioritization output — specifically, are you testing the attack paths our CTEM program has identified as highest priority?

The highest-value purple team exercises are not generic kill chain simulations — they test the specific attack paths that the CTEM program's prioritization and validation stages have identified as the most significant exposures. This requires the red team operator and the security operations team to work from the CTEM program's current output as the exercise design input.

The Stackcurve Take

The question "should we use automated pentesting or a manual red team?" is a false choice. Continuous automated pentesting and periodic manual red team engagements are complementary capabilities that cover different discovery depths and time scales. Replacing one with the other produces a validation program with a structural gap.

The operational model that the highest-maturity CTEM programs have converged on is a three-layer validation architecture: continuous automated pentesting for known-technique coverage and regression testing, annual red team for novel path discovery and adversary emulation, and periodic purple team for mobilization capability development. This architecture addresses the full spectrum of validation questions that a CTEM program needs to answer.

The budget implication is real — this three-layer architecture costs more than any single-tool approach. But the alternative is not lower cost; it is a validation program with specific, documented gaps that sophisticated attackers have learned to exploit.

The 2026 Stackcurve CTEM CURVE™ Report covers automated pentesting platforms, red team service providers, and purple team exercise design across the full CTEM validation stage. Download it free →


← Back to Research Library

Stackcurve Advisory Briefs are independent research. No vendor pays for placement, tier assignment, or editorial influence. The CURVE™ methodology is disclosed in full at stackcurve.net/research/methodology.