Co-authored by Deepti Mande, Director of Product, Rockfish Data, and Sreedhar Rao, Global Telecom CTO, Snowflake.
A sleeping cell is one of the more frustrating failure modes in a mobile network. A tower stops serving traffic. Users can’t connect. But the availability dashboard still reads 100%, because technically nothing about the cell is “down.” Traditional alarms walk right past it. By the time someone catches it, hours of service are gone.
It’s a small example of a bigger problem. The failures that matter most usually show up least in your data. This is true in telecom, in fraud detection, in industrial systems, in healthcare — anywhere teams are training AI on data that captures what already happened, not what could happen next.
If you want AI that catches these failures, you can’t wait for them to occur in production. You can’t copy them from history. You have to generate them.
Why standard ML misses them
Here’s the thing that took me a while to appreciate: this isn’t really a data-quality problem. The early signs of rare events are usually already in your data. They’re just too subtle for standard ML to pick up.
Standard ML learns the alarms once they’ve fired. It doesn’t learn the subtle precursors that would have caught the problem earlier. And as AI takes on more autonomous, closed-loop roles — triaging alerts, adjusting parameters, taking action without a human in the loop — confidence needs to be established before deployment, not discovered after.
Synthetic data is the obvious answer. Teams reach for it and that instinct is right. The catch: most synthetic data tools reproduce the typical really well and the rare almost not at all. For closed-loop use cases, the causal relationships between subtle signals are usually the thing that matters most — and that’s exactly what most synthetic data tools flatten away.
The Rockfish–Snowflake partnership
Rockfish and Snowflake announced an integration in March 2026 to close this gap directly inside the Snowflake AI Data Cloud.
The short version: Rockfish (built on research from Carnegie Mellon University) specializes in the temporal, causal, and cross-signal structure of real-world data — and in amplifying the subtle patterns behind rare events so ML can actually learn them. Snowflake is where a growing share of enterprise data already lives, governed and ready for AI workloads. Putting them together removes the steps that usually sit between “we have data” and “we have a trustworthy AI system trained on it.”
Rockfish runs directly inside the customer’s Snowflake environment. Synthetic data lands in Snowflake tables alongside the real data. The customer’s data stays inside their Snowflake account throughout — no exports to a third-party engine, no copies of the source data, no data movement out of Snowflake.
“Together with Rockfish, we are enabling carriers and vendors to generate test data within Snowflake’s governed environment — so they can move faster with confidence.”
— Sreedhar Rao, Global Telecom CTO, Snowflake
The initial announcement was telecom-focused. But the pattern applies pretty much anywhere edge cases matter more than the average.
Rockfish runs as a Native App inside the customer’s Snowflake environment. Real data stays put; the amplified patterns land alongside it, ready for Snowflake ML.
What the joint solution does
Three things, and they build on each other.
Synthesis inside your governed environment
Rockfish runs as a Snowflake Native App. It reads from your tables, generates synthetic data with the temporal, causal, and cross-signal characteristics of the real data intact, and writes results back to Snowflake. No export step. No external tooling to certify. No data-movement pipeline for security to audit. The synthesis happens where your data is already governed.
generation happens where the governance already lives
Amplify what your models are missing
This is the piece I get most excited about. Rockfish identifies the temporal, causal, and cross-signal precursors of rare events — the spikes, outages, gradual ramps, sustained shifts, correlated failures your production data has hints of but not enough examples to train on — and amplifies them into data your models can actually learn from.
These aren’t random anomalies bolted onto normal data. They’re the causal signatures of real-world failures, faithful to the real system, at a strength ML can detect.
subtle patterns, made learnable
“Rockfish preserves the temporal, causal, and behavioral characteristics that drive real-world network behavior — including rare failure events that may occur once in millions of sessions.”
— Muckai Girish, CEO, Rockfish Data
End-to-end validation, no data movement
Once the synthetic data lands in Snowflake alongside real data, Snowflake ML, the Feature Store, and the Model Registry let you train, version, and validate models where the data lives. Horizon Catalog governs the whole flow, so every test run is traceable and reproducible. Snowflake Cortex adds a natural-language surface for teams that would rather describe what they need than write SQL for it.
The result: faithful data, rare scenarios included, and a full model-validation loop, all inside a single governed environment.
a full validation loop, no data movement
What this looks like in practice
The telecom example is the easiest one to point to, partly because the sleeping cell problem is literal there, and partly because telecom AI is moving fast toward autonomous, closed-loop operation.
The end-to-end flow we’ve demonstrated: a developer requests a rare scenario like a sleeping cell. Snowflake calls Rockfish, which generates synthetic telemetry with the anomaly injected alongside real data. The developer’s model trains and validates in Snowflake against the accuracy targets the team sets. The validated detector goes back to production.
The same flow works for fraud detection, insurance modeling, healthcare, industrial IoT, security — anywhere edge cases are the missing piece of the AI validation story.
Each doing what we do best
The split is deliberate. Rockfish creates the training data: the temporal, causal, cross-signal fidelity of the synthesis, and the amplification of the subtle patterns your models need to learn. Snowflake provides the governed environment, the storage, the ML compute, the Feature Store, the Model Registry, the natural-language surface.
Neither company is in the other’s business. What the partnership does is remove every step between them, so a team gets the full loop without stitching together a data pipeline, a synthesis platform, an ML training environment, and a governance layer.
Getting started
If you’re already on Snowflake, you can add Rockfish through the Snowflake Marketplace listing (Rockfish Synthetic Data Platform) and start generating synthetic scenarios inside your existing Snowflake environment.
For a deeper conversation about the joint solution or specific use cases, visit rockfish.ai/partners/snowflake or reach out to sales@rockfish.ai.
If your AI is heading toward autonomous, closed-loop operation, a safer copy of yesterday’s data isn’t enough. You need synthetic data that moves the way real data moves, that can rehearse the events you’re most worried about, and that gives you the tests to prove your systems are ready. That’s what Rockfish and Snowflake, together inside the Snowflake AI Data Cloud, are built to do.
Rockfish Data, built on foundational research from Carnegie Mellon University, generates high-coverage, labeled datasets purpose-built for evaluating AI models and data analytics agents in production. Unlike generic benchmarks, Rockfish systematically covers edge cases, rare scenarios and domain-specific failure modes — giving teams the evaluation infrastructure to measure accuracy and reliability before deployment. Learn more at rockfish.ai.