Why Are Model Benchmarks Always Based On Offensive Security?
If you make an observation of any recent security evaluation for an AI model, there's a good chance it looks like a CTF scoreboard. Solve the challenge, capture the flag, pop the box, move up the leaderboard. Cybench, NYU CTF, BountyBench, etc, It’s all based on offense.
Now try to find the defensive equivalent. Is there a benchmark for triaging a thousand alerts at 3 am? For catching a phishing campaign before it lands? For correctly deciding that the weird thing in the logs is actually benign?
It barely exists. And that gap is not an accident. There are real structural reasons why the entire leaderboard is red, and I think they're worth spelling out, because the gap quietly shapes what gets built next.
Offense has a scoreboard. Defense has a backlog.
The simplest reason why benchmarking is based on offense is that attack grading is easy. Either the model got the flag, or it didn't. The exploit runs, or it doesn't. That's a binary win condition you can verify automatically, at scale, across thousands of runs. No human judgment involved. Offense is also episodic. A CTF challenge is a few hours of self-contained work, which fits neatly inside a benchmark run. And there's a scaling story benchmark authors love. A more capable model pops more boxes. Clean, monotone, publishable.
Defense doesn't work like that. A successful defense is a breach that never happened. How do you score a null result? "Nothing bad occurred" could mean the model was excellent, or the scenario was too easy, or it got lucky. There's no flag to diff against. Defense is also posture maintained over weeks and months, while a benchmark run is a snapshot, and a snapshot tells you almost nothing about how well a model holds a perimeter for ninety days. Defensive quality doesn't scale cleanly either. It plateaus, or it jumps in weird steps, and nobody wants a graph that goes sideways.
Twenty-five years of CTF culture, zero years of shared SOC data
Even if you wanted to build a defensive benchmark, what would you build it from?
Offense normally comes pre-packaged. With decades of CTF culture, it gave the world structured challenges, public writeups, and standardized tooling. Nmap is nmap everywhere. Metasploit is Metasploit everywhere. A skillful professional can stand up a plausible attack target in a Docker container in a single afternoon.
The defensive equivalent doesn't exist in public, and honestly, I feel it can't. Incident reports, alert queues, EDR telemetry, phishing mailboxes- this material is proprietary and soaked in sensitive data. Nobody publishes their SOC's daily grind the way CTF players publish writeups. The closest thing defense has to a write-up is an anonymized case study behind an NDA.
There's also a global circularity problem. To benchmark defense, one would need a realistic adversary generating attacks against a simulated enterprise. That adversary is probably a model. So now the quality of your defense benchmark depends on an attack model you'd also want to benchmark. The offensive side people never face this, their ground truth is a flag file or a vulnerability buried somewhere very deep.
The cost asymmetry
An offensive benchmark environment is a container. A defensive one is a simulated enterprise/organization. You need an identity provider, an EDR fleet, a SIEM pipeline, and telemetry that flows continuously and looks like a real company. The moment you build it, it starts decaying, because threats keep changing and realistic telemetry drifts out of date. Attack challenges don't decay the same way. An exploit that worked stays working.
The same asymmetry shows up in how AI agents actually operate. Offense chains beautifully. Recon leads to a vulnerability, the vulnerability leads to an exploit, the exploit leads to post-exploitation, and every step feeds the next step in a clean toolchain. Agents are very good at that kind of chaining. Defense doesn't chain. The building blocks - identity, infrastructure, network belong to different systems and different teams, and there's no tidy input and output handoff for a model to walk down.
You can't screenshot an attack that didn't happen
Now the incentive layer, which might be the strongest driver of all.
Assume your agent finds a twenty-year-old vulnerability in a popular open-source library. That comes with a CVE number, an advisory, a disclosure timeline. It's verifiable, it's memorable, and it writes its own press release. Nobody writes a press release about “3% fewer false positives”, even when it saves a company real money.
Marketing follows artifacts, and offense produces artifacts. A prevented breach produces nothing you can point at, and the best defense gets you is a quiet quarter.
The same logic shapes how labs position themselves with enterprises. A pentest is bought as a line item. Procurement signs, the agent runs, you get a report, and nothing about the org has to change. A defensive agent is the complete opposite sale. It asks a SOC to restructure part of its operation around a model. One is a product purchase, and the other is a transformation purchase. You could guess which one gets benchmarked and demoed.
From a purely strategic angle, on top, offense is the wedge into the enterprise. Find real vulnerabilities, earn credibility, then expand from there, and this is the wedge benchmarks follow, because labs benchmark what they can sell.
Offense is a solo sport. Defense is an org chart.
Here's what I feel is the deepest reason, and the one I think people underrate the most.
Point an offensive agent at a box, and it just goes. It recons, it finds a foothold, it exploits, it pivots, it goes deeper and deeper. The entire loop closes inside the agent. No approvals, no tickets, no standups, and a very good example of this is the OpenAI & Hugging Face incident. Benchmarks can measure that kind of self-contained autonomy, and to be fair, they should.
Defense never closes the loop inside one agent. Every meaningful defensive action drags dependencies behind it. Touch identity, and the IAM team needs to be involved. Touch infrastructure and platform engineering is in the thread. Legal wants to know why. Access requests, change windows, context that lives in people's heads and Slack channels rather than in any environment the model can see. Which means a defensive benchmark wouldn't really measure the model. It would measure the organization around the model.
Offensive benchmarks measure capability. Defensive benchmarks would measure coordination. A benchmark can only isolate what a single agent can do alone, and only offense is solo work.
Why this matters more than it looks
Benchmarks are not neutral scoreboards. They are the industry's to-do list. What gets measured gets built, and what gets built gets sold.
So we end up with autonomous pentesting agents before autonomous SOC analysts, not because offense matters more. Any CISO will tell you the defense market is bigger, the defense budget is bigger, and the pain is constant. We get offensive agents first because offense is measurable, packageable, and solo work. Defense is none of those things, so it waits.
The danger is a generation of security AI optimized for a scoreboard that only exists on one side of the house. Models get very good at finding holes because we can grade hole-finding. Then those same models get handed defensive work they were never evaluated on, and everyone acts surprised when the results feel thin.
What would change it?
None of this is permanent, and a few things could genuinely shift the balance.
Cyber ranges keep getting cheaper to build. A simulated enterprise that was a research project five years ago is a product demo or a POC today. Realistic defensive environments are becoming feasible in a way they simply weren't before.
Grading is getting easier too. Models are now decent judges of messy, open-ended answers, so you can score a triage note against a rubric without waiting on a human to read it. "Was this any good?" is finally a question you can ask ten thousand times. And the data problem is an engineering problem, not a law of nature. There's nothing fundamental preventing synthetic but realistic SOC telemetry, shared defensive corpora, or benchmark data sitting behind access controls. The CTF world bootstrapped itself out of writeups and culture. Defense needs its own bootstrapping effort, and someone has to start it.
The part worth remembering
A benchmark is a compressed representation of what a field/domain pays attention to. Ours is pointed entirely at the attacker and their mindset, not because defense matters less, but because attack is legible, cheap to grade, easy to demo, safe to publish, and a solo sport. Defense is a ghost. It's the absence of an event, spread across an org chart, owned by no single artifact.
If the point of security benchmarks is to tell us where AI will actually help, the current leaderboard tells us where AI is easy to test, not where it matters most. The first credible blue team benchmark should not just fill a gap. It should move real money, because the defense side of the market has been waiting a long time for a scoreboard of its own.