Earlier this month, we hosted a webinar titled Synthetic Data in Cybersecurity: Overcoming Data Bottlenecks in Training Environments and Detection Models,” in collaboration with Carahsoft, Cympire and SRI
The session brought together a diverse audience — public-sector practitioners, civilian and defense agency teams, cybersecurity researchers, cyber-range operators, and industry partners working on AI-driven detection and response.
Despite their different roles, a consistent theme emerged: access to realistic, usable data remains one of the biggest constraints in modern cyber defense.
As organizations are asked to move faster and rely more heavily on analytics and AI, the gap between real-world threats and the data available to train and test against them continues to grow.
The Cybersecurity Data Bottleneck
Cyber teams are not short on data volume — logs, telemetry, and alerts are generated continuously. The challenge is that the data that best reflects real threats is often the hardest to use.
Operational cyber data frequently contains sensitive infrastructure details, personal identifiers, or classified context. For good reason, this data is tightly controlled and difficult to share or reuse for training, testing, or collaboration.
During the webinar, Major General (Ret.) Neil Hersey, former Deputy Commander of U.S. Army Cyber Command, captured this challenge clearly:
“One of our most significant challenges wasn’t people or tools — it was realism.”
Training Without Realism
Most training environments rely on sanitized datasets, static scenarios, or replayed historical incidents. While useful, these approaches struggle to capture complex behaviors, temporal patterns, and rare attack scenarios.
“We were training hard, but in many cases, we were doing it with limited visibility.”
The result is a readiness gap. AI models are trained on partial representations of real environments, increasing the risk that systems perform well in test settings but fall short in production.
Why Anonymization Falls Short
Masking or anonymizing data removes identifiers, but often breaks correlations and sequences that detection systems rely on. What teams need is not just sanitized data — but safe realism.
Synthetic Data: Realism Without Risk
Synthetic data offers a different approach.
Instead of copying or anonymizing operational data, synthetic data is generated by learning the structure and behavior of real systems and producing new data that behaves similarly — without exposing real records.
In cybersecurity, this enables realistic logs, traffic, and adversary activity that can be safely reused, shared, and scaled.
“Synthetic data enables us to generate realistic, dynamic datasets that behave like real-world traffic — without exposing classified information or personal data.”
This makes it possible to:
- Train defenders on realistic scenarios
- Test AI detection models more rigorously
- Simulate rare or future attack patterns
- Run repeatable validation at scale
Enabling Collaboration Without Compromise
Cyber defense depends on collaboration, yet data-sharing constraints often limit what’s possible. Synthetic data helps bridge this gap.
“Synthetic data allows us to collaborate across agencies, with industry, and with partners — without exposing sensitive resources or methods.”
By decoupling data utility from data sensitivity, organizations can conduct joint training, shared research, and cross-organization testing while respecting privacy and governance requirements.
A Foundational Capability for Modern Cyber Defense
A key takeaway from the webinar was that synthetic data is no longer experimental. It is becoming a foundational capability for cyber training, AI validation, and secure collaboration.
As Gen. Hersey summarized:
“Synthetic data isn’t just a technical tool — it’s a strategic enabler.”
At Rockfish, we help organizations operationalize synthetic data for cybersecurity training, AI development, and testing — enabling realism without risk.