Dreadnode
Policy at DreadnodeLast updated August 2026

Policy at the intersection of
AI and cybersecurity.

Dreadnode fits squarely at this convergence, building autonomous cyber agents and the rigorous evaluations that measure them. We build autonomous cyber agents and the rigorous evaluations that measure them. That frontline vantage point grounds our policy stance in empirical research. We advance positions only on the narrow set of critical AI and cybersecurity challenges where that research offers direct insight.

The Thesis

The threat, and our response

01 / The Threat

Human-Speed Defenses against Machine-Speed Threats

As AI-driven cyber capabilities accelerate, traditional resilience baselines are rapidly losing efficacy. Yet public and private sector defense standards remain trapped in legacy solutions: manual reviews, reactive compliance, and bureaucratic reporting. In an era of machine-speed threats, strategic advantage belongs to those who actively evaluate, stress-test, and master these capabilities first.

02 / Our Response

Turning Offensive AI into Actionable Resilience

We emulate and research real-world adversarial behavior inside isolated, on-premise environments. By developing model-agnostic evaluations, we make opaque offensive AI capabilities measurable. We translate these findings into custom, red-team attack vectors that position CISOs to build truly adaptive blue-team defenses.

Our Positions

Five positions, grounded in our research

Investment in open research will maintain America's strategic advantage.

Our Position
The United States should treat open weight development and distillation as strategic infrastructure rather than as leakage to be contained. Policy should support domestic open weight release alongside sustained investment in the evaluation and runtime monitoring infrastructure that holds those models accountable after they diffuse.
Why
Geopolitically, the world is using more open weight models than frontier models. The U.S. can either choose to compete in this arena or isolate from this market, which will invariably reduce our competitive advantage. Restricting domestic release does not reduce global adoption; it relocates the ecosystem onto foreign models that American researchers cannot inspect, instrument, or influence. Openness also favors defenders, because open weights can be probed, fine-tuned, benchmarked, and monitored at runtime by anyone, including us. Our own evaluations confirm that the capability is already distributed, which raises the burden on evaluation rather than lowering it.
Related Research
In The Scaffolding is the Red Team, we re-ran 13 AIRTBench tasks that were previously unsolved or solved by only a single frontier model. GLM-5.2 and Kimi-K3 each solved 10 of 13, matching Claude Sonnet 5's aggregate. Given adequate scaffolding, the choice between an open weight and a frontier base model no longer determined the outcome.

Dynamic, model-agnostic evals must target capability diffusion.

Our Position
Security evaluations must be dynamic and model-agnostic. We advocate for evaluation frameworks that test capabilities wherever they land, whether in a proprietary frontier model or an open weight, distilled downstream variant.
Why
Policy attention concentrates on frontier models at release because they are easy to see and control. However, capability rapidly diffuses through distillation, fine-tuning, and open weight releases that bypass standard audits. Static benchmarks go stale the moment a model ships, and models learn to game them once the tasks are known. Evaluation regimes must therefore evaluate the full trajectory of agent behavior in real time.
Related Research
Between DreadIndex, PentestJudge, and ScopeJudge, we use LLMs as dynamic runtime evaluators. DreadIndex measures offensive capability across 76 tasks in 10 categories, and what that capability costs to exercise. PentestJudge and ScopeJudge assess whether agents meet operational requirements and stay in scope during live execution, and benchmark the judges themselves. Across 4,897 offensive tool calls, our best judge entered human range and still missed roughly one scope violation in 10. Every Model Cheats audited 1,518 traces to measure how frequently models cheat on offensive cyber benchmarks, and how far prompt-level mitigations reduce it.

Minimum security baselines must be enforced through federal procurement.

Our Position
Any AI system integrated into defense or critical infrastructure must demonstrate a minimum security baseline before deployment, enforced directly through procurement standards. Deploying unmeasured, unverified AI into critical infrastructure is an unacceptable national security risk.
Why
Today, an evaluation can prove a model is unfit for a sensitive environment, yet it can still be deployed because acquisition rules do not mandate otherwise. Federal procurement, modeled after enforceable frameworks like FedRAMP and CMMC, is the key lever to turn capability measurement into binding deployment decisions.
Related Research
Through our From Compute to Congress policy series, we analyze how federal mechanisms, executive orders, and cybersecurity mandates can bridge the gap between AI capability benchmarks and federal acquisition pipelines. Our most recent installment examines the June 2026 Executive Order and NSPM-11, which introduced a cyber-benchmarking mandate and a clearinghouse seated at Treasury, along with the resourcing gap that will determine whether either holds. DreadIndex supplies the measurement layer that a procurement standard of this kind would need to reference.

Defense must be automated to neutralize machine-speed threats.

Our Position
Automated, continuous defensive evaluation and remediation must become the default posture for critical systems, not a capability reserved exclusively for elite operators or organizations.
Why
Cybersecurity asymmetry heavily favors attackers: an adversary needs only one successful path, while defenders must secure everything. AI widens this gap by allowing adversaries to generate and adapt exploits at machine speed, rendering human-tempo defense structurally obsolete.
Related Research
We developed Ares and DreadGOAD, open-source frameworks for closed-loop threat emulation. While autonomous red-team agents execute complex kill chains in realistic Active Directory environments, blue-team agents analyze the resulting telemetry in real time. Mine the Gap reports what that closed loop measures, and our LLM-powered AMSI provider, paired against a live red-team agent, produced both a detection blueprint and the dataset behind it.

Establish legal safe harbor and telemetry standards for synthetic testing.

Our Position
Policymakers should establish clear safe-harbor standards for air-gapped red-teaming sandboxes alongside standardized agentic telemetry specifications for enterprise and defense deployments.
Why
Understanding how non-deterministic agents behave requires granting them broad action spaces in controlled, synthetic target environments. Overly broad prohibitions on offensive AI research risk stifling defensive innovation while doing nothing to deter malicious actors. Furthermore, closed-box security tools introduce unquantifiable risk without complete trajectory logging.
Related Research
Our work in synthetic simulation environments (Worlds) and agent observability (Agent Lens) demonstrates that security agents are made interpretable not by inspecting raw model weights, but by recording complete execution paths, API tool calls, and decision trajectories in isolated sandboxes. Worlds trained an 8B model to Domain Admin on GOAD entirely on synthetic data, evidence that serious capability research can be conducted inside isolation rather than against live targets.
Work With Us

Closing the security gap requires alignment

Closing the security gap requires alignment across policy, defense, and enterprise operations. Dreadnode bridges frontline offensive research and policy execution through direct collaboration with national security leaders and public-private defense coalitions.

Policy & Regulatory Engagement
We supply empirical research, model-agnostic evaluation standards, and agent observability frameworks to help lawmakers build realistic, enforceable AI security baselines.
National Cyber Defense & Mission Partners
We partner with defense organizations to operationalize advanced threat emulation and simulation capabilities across high-security environments, enabling teams to train, evaluate, and defend at machine speed.
Enterprise & Infrastructure Resilience
We work directly with CISOs and security teams to translate continuous red-team attack trajectories into real-time blue-team detection and remediation.

Email daria [at] dreadnode.io to inquire about collaboration opportunities as we empower policymakers to navigate the AI-cyber threat landscape.