Ideas on synthetic data, agent evaluation, and reliable AI
Every AI team tells the same story — the demo works, then the project quietly stalls between "this works" and "this is in production." The gap was never talent or ambition. It was data. Here's where we're headed.
Read more →Across 49 datasets and 11 models, TabQueryBench evaluates synthetic data by the questions people actually ask of it. Joint research with CMU, HKU, Tongji & UMD — showing where even the best generators break.
Read more →Modern analytics and AI systems evolve continuously — schemas change, models are retrained — yet the data used to test them stays static. Here's how Rockfish closes that gap.
Read more →Two conflicting views on the nature of data: is it the "new oil," or something we should share as little of as possible? A look at both camps through the lens of generative AI.
Read more →Data availability has quietly become one of the biggest bottlenecks to AI innovation. Schema-driven generation lets teams prototype, test, and validate before production systems even exist.
Read more →Recap of our webinar with Carahsoft and Cympire — bringing together public-sector, civilian, and defense teams to close the data gap in training and detection.
Read more →From buzzword to business tool. A year of customer conversations shows the shift: teams have stopped asking what synthetic data is and started asking how to put it to work.
Read more →Why should you care about synthetic data quality? Imagine training a new hire from a badly translated manual. A beginner-friendly guide to measuring it right.
Read more →