Consumer Protection Tuesday: AI-Powered Continuous Adversarial Testing at Coinbase

By Elias Abouzeid, Rick Roane, Tyler Almeida • 7min read

TL;DR: Introducing CAT (Continuous Adversarial Testing), an internal platform that uses autonomous AI security agents to test both existing Coinbase assets and newly discovered services, product changes, and deployments. New activity is evaluated continuously, while existing assets are also tested as needed. It covers web and mobile applications, backend services and infrastructure, Web2-to-smart-contract integrations, and internal AI tooling, augmenting our offensive security team’s ability to identify risk at a scale and cadence that were previously impossible.

Consumer Protection Tuesday: How to Spring Clean Your Digital Life
asdf

Why It Matters

Adversaries don't pause between pentest engagements. They probe continuously, they're better tooled than they've ever been, and they're now AI-augmented. The traditional security model of running a handful of point-in-time pentests each year was built for a pre-AI world, where the bottleneck was human capacity. To stay ahead, defenders need to operate at the same cadence and the same scale as the attackers.

Image 1

At Coinbase, our offensive security team made a deliberate bet to build a platform that gives every employee the leverage of a small army. We've spent the last several quarters building exactly that.

Meet CAT 

CAT is an internal platform that runs purpose-built AI agents. Findings from every agent flow into one validation pipeline, one triage queue, one audit trail, and one set of enforced Rules of Engagement. CAT has a number of already built out capabilities, and it continues to evolve and branch out into new domains:

8923847

New code scanning

  • PR security review: CAT scans every commit across every org as it lands, reviews again at merge, and rolls the results into a full product-level review at product launch. Most engineers resolve CAT advisories during their development flow, prior to them becoming vulnerabilities. All three stages run on frontier AI models at varying depth and breadth, providing top-tier security intelligence across the SDLC.

  • Product Launch reviews: Wired into the Product Launch pipeline: each cleared product gets an automatic security pass with no request needed. Agents read its design docs and surface design-level gaps.

  • Attack surface monitoring: Integrated with OffSec's homegrown attack-surface platform; diffs the estate every snapshot so any new host, service, repo, or endpoint is auto-discovered, queued, and tested the moment it appears — endpoints mapped back to their serving code, live hosts picked up for infra testing.

Existing code scanning

  • SAST & design hunts: Agents read the source code, reason about data flow across the codebase, and surface vulnerabilities.

  • Web2 ↔ Web3 connections: Agents map every point where backend code touches a smart contract and test each boundary, working repo → contract or contract address → every caller.

  • SHADE (Swarm Harness for Adversarial Discovery and Exploitation) — A swarm of dedicated hardware agents that each claim a repo, hunt for exploitable vulns, and report back, scaling coverage horizontally across the estate.

  • MCP Registry Scanner: Continuously evaluates adopted MCP servers against an MCP-specific threat model, rescanning on change and sweeping for unregistered "shadow" servers

  • Prompt Injection Review: Agents inspect code changes for injection paths into our own AI systems, covering the instructions, tool definitions, and untrusted content that reach a model.

  • Agent skill trust: Capabilities we grant internal AI agents are scanned before adoption and rescanned on change, so an agent's reach never quietly expands beyond what was approved.

  • MAST: Continuous iOS/Android assessment on every release: OWASP Mobile Top 10 static analysis, full MASVS conformance with manual verification, a privacy audit of permissions and third-party data flows, and MITM hostile-network resilience testing.

Dynamic testing

  • Infrastructure Testing — Continuous, non-destructive assessment of the network estate: agents enumerate reachable services, reason about exposure, and map every check to recognized methodology (NIST SP 800-115, PTES, MITRE ATT&CK). Testing is identify-only by default, with no exploitation and no credentialed access. Between assessments CAT diffs the estate and flags newly exposed services, so drift is caught as it happens rather than at the next scheduled review.

  • DAST — Agentic web-app testing that authenticates as a real user, crawls the app, and chains findings like a human attacker — from injection and access-control flaws to the business-logic bugs scanners structurally can't find.

  • Live Operative — An interactive pentesting surface where an engineer and an AI operative work the same target together. Rather than re-verifying what the automated scanners already found, Live Operative plans a fresh look: it takes the engagement's threat model as its primary input and treats existing findings as context to go deeper on, not results to repeat. It reasons over the full engagement picture, including design docs, operator notes, and the repos, endpoints and infrastructure in scope, then proposes adversarial scenarios for the engineer to approve, refine, or redirect. The engineer can steer mid-engagement, and the operative asks for help when it hits something it cannot resolve alone, such as an auth flow that needs a human decision. Every session is bounded by Rules of Engagement enforced outside the model, so the agent's autonomy never extends past what the operator authorized.

Guardrails the Model Cannot Talk its Way Past

Running autonomous agents against production infrastructure only works if the agent's authority is enforced somewhere the agent cannot reach. CAT's Rules of Engagement are a versioned, server-side policy applied at two independent layers: once when work is queued, and again structurally in the scanner before any packet leaves. Deny lists always override scope. Test windows bound when testing can run. Blast-radius caps bound how much can be touched at once. Fragile services are excluded from active probing, and a fleet-wide kill switch halts everything.

63453456

A separate guard then checks what every command would actually do before it runs. It always allows read-only actions, but blocks anything that would change data or state unless the target has been explicitly declared a non-production environment at scoping. This check happens at the tool execution layer, not in the system prompt, so it doesn’t matter what the agent is told or how it’s asked, production stays protected. 

Quality, not Just Quantity

Every CAT finding goes through a multi-stage validation pipeline before a human ever sees it. 

  • First, a preflight check confirms every tool the pipeline needs is actually reachable: code search, the LLM, ticketing, live infrastructure and telemetry. This catches degraded runs before a token is spent. 

  • An AI agent then reads the finding against the real source code and live system context and reaches an initial verdict. 

  • If it looks real, an independent code-level trace verifies the exact attack path from the entry point to the vulnerable code, then follows what it can reach downstream. Along the way it checks whether any controls sit on that path and whether they actually block the impact, only catch it after the fact, or don't affect it at all.

  • A second AI pass reviews everything adversarially. It sees the first pass's reasoning and the trace evidence, but its job is to independently re-check every claim and override the verdict if it finds something the first pass missed. If that second pass didn't dig deep enough to trust a negative finding, the pipeline automatically kicks off a third, more thorough pass. This escalation is deterministic and capped at three passes, not an open-ended retry loop.

  • Only confirmed findings move on. A likelihood check grounds exploitability in real production traffic. Separate privacy and operational reviews re-check data sensitivity and business impact against real data classifications, SLOs, and incident history. 

  • For internet-reachable bugs, a wild-exploitation check scans live production logs and WAF telemetry for signs the bug is already being probed. 

What comes out the other end is high-signal and ready for a human to act on. Confirmed issues flow straight into our tracking system with a full audit trail. That same trail (scope, methodology, severity rationale, regulatory mapping) is what CAT turns into audit-ready assessment reports, so evidence for auditors and regulators comes out of the testing itself instead of a quarter-end scramble.

Rating findings is its own problem. CVSS was not built for a world where a vulnerability's worst case is measured in customer funds, so we replaced it with a six-factor model tailored to our threat environment. Alongside likelihood factors like exploitation frequency, attack complexity, and required access, it scores three kinds of impact directly: funds, data, and operations. A finding that puts customer assets at risk is rated with that context, instead of being flattened into a generic severity band.

879897

Humans, Augmented (Not Replaced)

CAT isn't trying to replace our offensive security engineers. It's making them dramatically more effective by automating breadth work across the platform, freeing engineers to spend more time refining existing capabilities and building new ones that push discovery and defense-in-depth further. For investigations that need a human in the loop, CAT includes Live Operative.

567567

What's Next:

  • Evals — Measuring the agents themselves. We are building an evaluation pipeline that scores agent output against known-good datasets, so we can tell whether a prompt or model change actually improved detection instead of assuming it did.

  • Smarter Prioritization at Scale — As SHADE's coverage grows, the constraint shifts from how much we can test to what we should test next. We are building a prioritization engine that targets effort at the code and assets where a finding would matter most.

The Impact

Through internal use of CAT we've seen several measurable shifts in how our security team operates:

  • Continuous coverage replaces point-in-time engagements. CAT has completed more than 150,000 scans against Coinbase's production estate since mid-2026, including over 128,000 pull request reviews. Every change is reviewed when it lands, not at the next scheduled engagement.

  • Significantly more penetration test findings fixed each month. In 2024 and 2025, our offensive security program fixed a steady number of pentest findings each month. After our pentesters started using AI tooling in late 2025, that number went up sharply. CAT now finds most of them, and engineering teams have fixed a large share of what it has surfaced. Many are caught at the merged pull request, before the code ships to production.

  • Frontier models, pointed at the whole estate. Running the most capable models available against every pull request and every repository is normally a cost problem, not a capability problem. CAT solves it with per-scan budget enforcement, aggressive prompt caching, and model selection tuned per workload, which means we can afford to put frontier-class reasoning on routine changes rather than reserving it for the crown jewels. Better models make the platform better without a redesign.

  • Engineers focused on the hardest problems. Surface-level issues are caught, validated and routed autonomously, so the offensive security team's hours go toward novel attack research, threat modeling, and the deep multi-step investigation that only a human can do.

Looking Ahead

Currently CAT is internal only. As we mature the platform, we're evaluating which components could be made available externally.

We're also hiring. If building systems that continuously break our own products before adversaries can sounds like the kind of work that makes you reach for coffee in the morning, we'd love to talk.

Disclaimer: CAT does not replace human-led offensive security work. It augments our team's reach and lets them focus their expertise where it matters most. Deep, complex, multi-step attacks still require experienced human researchers, and always will.

Recent stories

Disclaimers: Derivatives trading through the Coinbase Advanced platform is offered to eligible EEA customers by Coinbase Financial Services Europe Ltd. (CySEC License 374/19). In order to access derivatives, customers will need to pass through our standard assessment checks to determine their eligibility and suitability for this product.