1,265 karma · joined May 26, 2010
I think it would sell like hotcakes.
Everything that's been added in terms of functionality is just window dressing in IMHO. BigQuery and in particular Snowflake have a superior architecture, with the separation of storage and compute.
Also, someone mentions in on the thread, the Redshift marketing was horrible.
Features, features, features, features, etc.
Compare that to Snowflake, and how their message evolved:
"The Cloud Warehouse" "Data Cloud"
So much more compelling. Also, Snowflake had a killer sales and marketing team.
Many other, little things. The Redshift team was constrained by what the AWS Console would give them. Snowflake could build more, better admin features. Redshift tried to mitigate that by acquiring a client (Datarow), but from what I heard, the acquisition never got integrated.
Having said that, set-up and configured the right way, Redshift was faster and cheaper than any other data warehouse on the market. Except - nobody wanted to spend the time on properly configuring their cluster. People just wanted their warehouse to work, and that's what BigQuery and Snowflake delivered. Even when that meant paying more.
The time you spent tuning Redshift for performance, you spent on tuning Snowflake and BigQuery for cost. Pick your poison. But again, people didn't care about the money - they just wanted things to work.
I didn't really see BigQuery as competition, simply because that meant switching clouds.
What I did hear is that analytics teams preferred GCP overall, and I think that has driven a lot of cloud workloads from AWS and GCP, because in the last few years analytics teams started to be decision market in many companies.
IMHO, GCP has the much better analytics portfolio than AWS. By now, also the better sales team. It's been really smart by GCP to bet on data science and analytics because of data gravity. Once you have the data in your cloud, it attracts workloads.
Snowflake is what eventually killed our performance tuning business. You can find my post-mortem on our company on Medium.
As you can probably tell, I know more about the whole analytics ecosystem than I bargained for.
The much larger problem is industrial manufacturing, in clusters like Ludwigshafen, where you have companies like BASF and Benckiser that rely on Russian gas for their refinery operations.
Fertilizers, chemicals, lubricants, etc.
You turn off Russian gas, you turn off the German industrial complex. Literally over night.
And that’s the reason why the German is so cagey about stopping the import of Russian gas. They can’t, because they don’t have an alternative.
We worked on an infrastructure with 400K+ active resources. It’s just not possible for a human to track that without any help from tooling.
If you don’t do something about those resources - at some point they will become a problem. We had cases where we ran into quota limits, and it would stop the entire operation.
Next to search, we also offer ways for automation. Say you wanted to find all unused load balancers and mark them for clean up. You can persist that search as a Job and let it run on a schedule.
The post describes how we developed a search syntax to solve the problem of finding resources running in a multi-account infrastructure. While the post is about AWS, the search works also for GCP.
The (not so) secret sauce behind the search syntax is an open source project that we developed out of necessity of taming a multi-cloud infrastructure (AWS and GCP). We quickly realized that it's not just about understand WHAT resources are running, but also HOW these resources dependent on each other.
Behind the scenes, Resoto collects your cloud resources and puts their metadata and dependencies into a large indexed graph that can be searched. The search syntax is how.
In a way, parts of our product are an open source alternative to e.g. AWS Config and GCP Config Connector.
Would love to hear any reactions from the community how they're currently solving the problem of keeping inventory of what's running in their cloud accounts.
Disclosure, co-founder here, we're building one of those CLIs. We started as an internal project at D2iQ (my co-founder Lukas commented further up), with tooling to collect an inventory of AWS resources and be able to search it easily.
The result is that you get a lowest common denominator type of dashboard. And hence a whole industry of providing just a prettier dashboard on top of AWS / GCP / Azure metrics.
Datadog started with a prettier dashboard for Cloudwatch data.
Cloudability started with a prettier dashboard for the Cost and Usage Report.
And also works the other way around. The individual product teams buy development environments to circumvent the console restrictions.
For example, a few years ago, the Redshift team purchased "DataRow".
Cloudkeeper is an open-source tool for Site Reliability Engineers (SREs) to perform cleanup of cloud infrastructure “drift,” leaky resources, and services that are triggering quota limits. For example:
- Unattached storage volumes without recent I/O
- Active Amazon Cloudwatch alarms for compute instances that no longer exist
- Load balancers without target groups
Because of cloud-native infrastructure and the ever-expanding menu of services from cloud providers, the potential use cases are limitless.
Cloudkeeper crawls your cloud, indexes and maps resources into a directed graph, and captures resource dependencies. Cloudkeeper ships with both a command-line interface (CLI) and a query language. The CLI makes it easy to search the graph for any asset and build workflows to collect, clean up, and generate metrics. There is also a Prometheus exporter for those metrics.
There is a plethora of existing cloud management tools that promise to perform those jobs. These tools usually fall into one of two categories: (1) individual asset discovery tools and (2) rule-based cleanup tools.
(1) Discovery tools generate long lists of resources—essentially a prettier view of the data from your cloud console. However, they do not perform cleanup. As a result, the vendors of these tools push professional services promising “optimization opportunities” or “actionable recommendations.”
(2) Cleanup tools, on the other hand, enforce rules and policies in your cloud accounts. But they do not facilitate asset discovery, nor do they aid in pinpointing the root cause of resource leaks.
In short, existing tools either provide reporting or automated clean-up, but not both. It is difficult to take insights from the reporting tools and use them for cleanup. As a result, the amount of infrastructure “drift” continues to grow
In our experience, SREs get stuck in a constant cycle of trying to determine what is running, if it should run, who is running it, and if it can be safely pruned. Many SREs maintain collections of scripts scheduled to execute at regular intervals, which quickly become unwieldy as new scripts are constantly added to handle new edge cases. We have spoken with dozens of SREs over the past six months, all of whom experience these problems in their day-to-day work.
Lukas built the first version of Cloudkeeper as an SRE at D2iQ in response to growing sprawl in the D2iQ infrastructure. Here’s an image of what the D2iQ cloud looked like at the time: https://github.com/someengineering/cloudkeeper/raw/main/misc...
As soon as Lukas generated this image, he realized there was simply no way for an individual engineer to manually manage this infrastructure. Lukas and D2iQ open-sourced Cloudkeeper because it helped them dramatically reduce sprawl, maintain control over all assets running, and free up the SRE teams’ time; they knew it could help others as well.
We have spent the last few months adding support for more services on AWS and GCP. We are keeping Cloudkeeper open source, because it helps to address the long tail of cloud services. Closed-source vendors will always have ROI considerations when building out support for a new cloud service. In fact, that is the reason why most SREs we’ve talked to use at least one commercial tool AND maintain a collection of scripts: to address cloud services and/or use cases not supported by the commercial tool(s) at their disposal. Our goal is to develop Cloudkeeper into an extensible product, and for it to be easy to add support for new resources or cloud providers.
Please check out Cloudkeeper on GitHub (https://github.com/someengineering/cloudkeeper). If you need help or support getting started, you can chat with the team on Discord (https://discord.gg/someengineering).
https://github.com/someengineering/cloudkeeper
I’ll reply more in depth since I’m on the run right now, but for now I hope the link is sufficient.
Same playbook - show that you’re better in a key metric that’s easy to understand (performance) to get the attention, but then pitch the paradigm change.
In Snowflake’s case, that was separation of storage and compute.
In Databrick’s case, it’s the Lakehouse Architecture.
I think the reason why Snowflake is so nervous because they know they can’t win this game.
One immediate thought for a feature (you may or may not have already thought of this...). A picklist of options for "how did you hear about us?"
That will help marketing and sales to get more granularity than the usual 80/20 split between "referral" and "direct" in Google Analytics. With all the brand-building efforts like newsletters, videos and podcasts going on - traditional "last touch attribution" tools like Hubspot or even multi-touch tools like Bizible don't capture brand.
In the Mad Scientist BBQ, he suggest to wrap it twice and also put beef tallow into the second wrap. One wrap is to get through the stall, and then a fresh wrap for the rest. Maybe that's worth an additional list item.
For everybody who wants to get going with BBQ, watch these channels:
Mad Scientist BBQ - https://www.youtube.com/channel/UCselvHbb5ah0sEqZrFa-7nA - just fantastic video production
Harry Soo - https://www.youtube.com/channel/UC4dtbTXdvjo272b5x_mxRBw - Harry keeps winning 1st place in all BBQ comps he participates in.
Texicana BBQ - https://www.youtube.com/user/MCglobalvision (he is Aaron Franklin's pitmaster)
In the old approach, you would run the transform BEFORE loading data into the warehouse. The disadvantage of that approach is that you loose all fidelity of the raw data.
In the new approach (Airbyte's approach), you load the raw data into the warehouse, and then run your transform jobs in the warehouse. You can do that because modern warehouses are cheap and scalable. The benefit of that approach is that you keep your raw data with all its fidelity, opening up endless opportunities for exploratory slicing and dicing.
That's why it's called "ELT" (new) these days, to distinguish from "ELT" (old).
My unsolicited $0.02 - I think your approach is spot on.
As a company, you will never have one consistent data set and metrics if you keep building an individual model for each user / use case / etc. And I've seen the explosion of tables and models in real-time. They just keep growing. And how do you even know that the question you're asking in your dashboard is pulling the information from the correct table? I've yet to see a data team that didn't have to deal with drift. Plus, there's a real cost of storing all these stale tables that nobody is looking at anymore.
What your product is doing is what I see companies already trying to accomplish themselves [somewhat]. For the leading companies when it comes to working with data, the warehouse today is already the source of truth, with one dimension table that points back to the SaaS tool / dashboard via an S3 bucket. So the SaaS tool itself is really only the last mile and visualization layer. Run the model, create the table, offload the table to S3, point the tool to the S3 bucket with the table. Update every 4 hours, etc.
dbt wins in that world. (and I assume you're using something like dbt under the hood of narrator.ai?)
That approach is already commoditizing the SaaS tool down to the visualization layer and the opinionated way of displaying data. But that still means there's at least one model per tool, use case, etc. with one table - and you still don't see the entire journey of the user, that's something you either have to create for a single specific use case, or cobble it together ad-hoc. If instead you have one table that has it all - you can move soooo much faster with data, and take out all the friction that comes from having disparate data sets.
narrator.ai wins in that world.
Blinkist is Berlin is following a very similar approach to what you guys have built. This deck is a few years old, but I think the approach described will resonate with you:
https://www.slideshare.net/SebastianSchleicher/tracking-and-...
If I had to look into my Crystal Ball, I think one of your GTM challenges will be to convince existing data teams that everything they've built is somewhat redundant. On the flipside, I can see the same data teams say "OMG, finally!". I'm curious to hear the customer reactions so far.
I'm very excited about this product! I wouldn't be a direct user with my current role, but FWIW, I can share the bruises I got from working in this market.
Would love to hear more!
Too much money gives you a false sense of security.
this is 4 years ago, the market we were going after were the cloud warehouses like Redshift, BigQuery and Snowflake.
In the beginning, we just had an idea for a specific product / service. but we knew that the customer would be the lonely data engineer in charge of building the analytics stack.
I started with cold outreach via my network and linkedin. I would use pretty broad language around data warehouse usage, to cast a wide net with terms like "usage, performance, metadata, etc.", and cover all potential use cases. Engineers are always short on time, so you need to be affirmative, authoritative and present a clear ask that shows what you want and how the engineer will get value out of spending 30 min with you.
Of course I pulled the founder card, and that does help. I made it clear that we don't have a product, but working on building one. Turns out that most people are helpful, and want you to win!
The first signal that you're onto something is when they reply to your message, and are intrigued.
we built our first product around that feedback. Companies like Postmates, WeWork and Udemy bought version 0.5!
I made a point out of keeping in touch with every customer, and do a quarterly check-in. What's changed? What are your plans? For this coming week, month, quarter and year. Where do you want to be in 2 years? What problems are you trying to solve for your company? What are the expectations for you and your team? What tools are you using to solve that problem? What tools have you looked at and decided to not use them, and why? Etc., etc.
I rolled around in THEIR situation, trying to walk in their shoes. We'd talk for sometimes 90 minutes, often in person here in SF, over lunch. We'd often not cover our product until the final 5 minutes.
Of course, that's a huge chunk of time out of your calendar. Huge opportunity cost, in particular for a founder.
So here's actually something I'd do different. We hired our first sales reps after our first 10 customers. That was a mistake. Our product was very technical, it's not like selling email software. So you need a technical rep. Our first rep was a class act, but let me tell you, it was a goat rodeo for him.
Rather, I should have hired a customer success person, and have them do the quarterly check-ins. That way I could have kept selling. Then let the customer success rep also write (technical) blog post, customer stories / case studies, and documentation. You're building out (credible) marketing materials that build the top of your funnel.
Most of the books recommended in this thread assume that you're working for a established firm, with product / market fit, etc.
Clearly that's not the case for a start-up.
Read up on what Pete Kanzanjy publishes https://www.foundingsales.com/ - it covers the the "founder-led sales" phase.
How I did it:
- you interview as many prospects and customers as possible
- you understand what keeps them up at night, what specific pain points they have, the language they use to describe their situation
- you shape your messaging to solve those specific pain points, using their own language
- wrap your messaging into a story - the worst you can do is "problem / solution". people don't buy that way. people buy change, and you use the story to communicate that change.
I wrote a totally too long Medium post on the whole topic:
https://medium.com/@larskamp/the-5-cs-an-operating-framework...
as somebody else pointed out on this thread - be read to deal with objection! You'll likely collect 9 "No's" for each "yes!"
For example, with intermix.io, it was the tracing we had built for other tools like Looker and dbt. The insight was that the result of a DAG involves many different calculations across different tables. The metadata only tells you that the steps happened, but doesn't tell you in what sequence they happened, where the "hiccup", latency, etc.
Redshift is clearly suffering from Snowflake. I wrote about that in my post-mortem. That post also has a few battle scars:
https://medium.com/@larskamp/why-we-sold-intermix-io-to-priv...
Ping me on LinkedIn if you care to hear more :-)
My experience is that monitoring data quality is a still an under-appreciated discipline. I've found that most teams still have an "not invented here" mentality, or don't even know they have the problem! That can lead to a "oh, we can just fix it when it happens" type of mentality. But your timing may be better than ours - we started back in 2016.
I haven't played with your product (yet), only took a look at this thread and your website. Some observations:
- SQL Editor - big plus! I think giving your users a space where they can take action is a super value-add, we didn't have that.
- nice work running the tests inside the customer's warehouse. That has two benefits for you. 1) you're not incurring the cost to crunch the metadata, it can get quite expensive, depending on the number of tables in the warehouse. 2) you're avoiding data access issues, getting access to the warehouse was always a hurdle, even though we only needed access to the system tables.
- pricing model. I think the per-seat model is the way to go. We tried charging by number of rows, and size of the warehouse (number of nodes), but then you run into weird situations with customers who are dealing with huge historic datasets, but really only look at the last 30 of data.
My unsolicited $0.02 is that you think hard about distribution. I think you want to think about hitching your wagon to the cloud marketplaces, and Snowflake's marketplace. For example, attaching themselves to Snowflake is what made all the difference for Fivetran.
I have a bunch of more scars that I can share if you care to know them :-)
Patrick, great job for:
1) trying something different and being persistent and polite, but not annoying
2) reaching your objective, and
2) documenting it in a post.
You've got the result on your side, and you found one possible way to get there. Are there others? Sure! Get an intro via a mutual connection, etc. You chose the one that worked for you given your constraints / circumstances.
When I saw your slide, my immediate reaction was "he must have been at Accenture". Looks like the templates haven't changed since I left in 2013 :-)
there's a bunch of great hosted tools out there, across ETL, workflows, dashboards, etc. Think Fivetran, Segment, Matillion, Periscope, etc. And then of course the warehouses like Snowflake, Redshift, etc.
But I think there are three issues with that stack, roughly like this (I've got to do some more thinking, would appreciate your input):
- Privacy: you have your customer data flying around in all these different tools. it's hard to impossible to track your compliance
- cost: all these vendors charge in some way by data volume, MAUs, etc. - you get taxed multiple times for the same data stream. It all adds up.
- control: your data is subject to pre-determined schemas, proprietary formats, black boxes, etc. - mismatch for the same metric across different tools, less flexibility to manipulate your data and pack up and go elsewhere.
I think there's a valid open source alternative for every layer of the stack:
Segment --> Rudder Labs, Snowplow
Matillion --> Airflow, dbt
Fivetran --> Stitch / Singer
Periscope, Looker, Tableau, etc. --> Metabase, Superset
Warehouses --> just yesterday I learned about materialize.io here on HN
And then add open source products like PostHog, that add additional value for very specific use cases (in this case product analytics).
Not arguing the value of the hosted products. They are amazing to use if you just get started. But there's a great open source "stack" available that long term likely will be more transparent, more flexible, and cheaper.
Would love your thoughts!