Are you able to say anything about pricing here?
Are you able to say anything about pricing here?
The primary reason we haven't provided pricing is that we have just launched and wanted to collect more data points before setting making the pricing public.
Our current offering is:
1) Free for diffing datasets < 1M rows 2) $90 / mo / user for diffing datasets of unlimited size 3) Cross-database diff, on-prem (AWS/GCP/data center) deploy, Single-sign-on are in custom-priced enterprise bucket.
We would love to hear your thoughts on this.
Option 1: make it free, up to a certain dataset size. You can harvest interested leads like the gentleman above.
Option 2: (if you don't want to deal with a huge volume) offer it for $50 one-time fee, up to X size, for Y months (e.g. $50, up to 1 GB, valid for 3 months). Nice way to filter qualified leads.
There are variations from the two options above, but I think you can easily get the general idea.
Thoughts?
We're leaning towards Option 1: free diffing for datasets < 1M rows. Option 2 seems a bit tricker since we are in a way creating a new tool category and it can be harder to convince someone to pay before they try and understand the value (unlike, say, a BI tool – everyone knows they need some kind).
I'm also really keen to hear why you built this tool - what use cases you expect. I've used free diffing tools a few times before in the past, but I think every time it was to make sure I hadn't messed up "manual" data migrations (which obviously aren't a good idea).
The main use cases we've seen: 1) You made a change to some code that transforms data (SQL/Python/Spark) and want to make sure the changes in the data output are as expected. 2) Same as (1) but there is also some code review process. In addition to checking someone's source code diff, you can see the data diff. 3) You copy datasets between databases, e.g. PostgreSQL to Redshift and want to validate the correctness of the copy (either ad-hoc or on a regular basis).
We have a signup-free sandbox where you can see the exact views we provide, including schema diff: https://app.datafold.com/hackernews
Is the value prop “you don’t need all they grunt work” as opposed to above direction?
Data testing methods can perhaps be broken down to two main categories:
1. "Unit testing" – validating assumptions about the data that you define explicitly and upfront (e.g. "x <= value < Y", "COUNT(*) = COUNT(DISTINCT X)" etc.) – what dbt and great_expectations helps you do. This is a great approach for testing data against your business expectations. However, it has several problems: (1) You need to define all tests upfront and maintain them going forward. This can be daunting if your table has 50-100+ columns and you likely have 50+ important tables. (2) This testing approach is only as good as the effort you put to define the tests, back to #1. (3) the more tests you have, the more test failures you'll be encountering, as the data is highly dynamic, and the value of such test suites diminishes with alert fatigue.
2. Diff – identifies differences between datasets (e.g. prod vs. dev or source DB vs. destination DB). Specifically for code regression testing, a diff tool shows how the data has changed without requiring manual work from the user. A good diff tool also scales well: it doesn't matter how wide/long the table is – it'll highlight all differences. The downside of this approach is the lack of business context: e.g. is the difference in 0.6% of rows in column X acceptable or not? So it requires triaging.
Ideally, you have both at your disposal: unit tests to check your most important assumptions about the data and use diff to detect anomalies and regressions during code changes.
It's a niche, and a small one at that, but I don't doubt you'll find some enterprises willing to pay what you ask. But outside of enterprise I just can't see anyone paying that, especially when free tools exist (albeit not nearly as polished and features as yours).
I've got to wonder about YC backing for what seems like such a small niche though - very possible I'm not seeing something you have planned for further down the line.
Agree with you about the niche. Diff is our first tool that helps test changes in the ETL code, and the impact is correlated with the size and complexity of the codebase.
Diff also provides us a wedge into the workflow and a technical foundation to build the next set of features to track and alert on changes in data: monitoring both metrics that you explicitly care about and finding anomalies in datasets. We've learned that this is something a much larger number of companies can benefit from.
For mid-sized companies though that could be a bargain. I worked on a data migration project at a company with <100 people and <5 engineers where we had to hack together our own data-diff tools and this would have been a bargain.
For the majority of diffs we see with sampling applied, sample sizes are <1M rows (more is often impractical in terms of information gain for higher compute costs) especially if your goal is to assess the magnitude of the difference as opposed to get every single diverging row.
Hmm. Why not?
Aside from that, it just doesn't seem to fit the spirit of HN - this is just personal opinion, of course.