Why is Snowflake so Valuable?
freshpaint.io
freshpaint.io
I would be apprehensive in investing in Snowflake long term purely because their product is highly susceptible to being obsoleted in the next 5-10 years.
As far as I can tell, it is a unique product in the database space. Extremely well executed ideas and design.
The killer feature for me was the query profiler - you can see WHY a query is taking a long time and optimise it - BigQuery just felt like Google were brute forcing the performance, and then charging you accordingly.
When the project I was on switched, the micro-clusters (and the ability to recluster a table) as well as the MERGE semantics beat BigQuery hands down - although those features my be out of beta now (but I've moved on to a new gig).
Also, BQ does explain the query plan to some extent: https://cloud.google.com/bigquery/query-plan-explanation. Not quite at the level of a "regular" SQL DB, but it does give you some info to work with when optimizing queries. If you haven't used it in a while I'd give it another try.
What they have now strikes me as an even better solution to the problem of bankrupting someone with a query IMO. Not sure how pricing compares to redshift et al, but pricing is the easiest thing for Google to change.
If you need to read a terabyte of data to answer your query then more slots only gets it done faster.
Wondering if this is a somewhat new feature in BQ since you used it, or if there's still a feature gap here (e.g. see https://cloud.google.com/blog/products/gcp/performing-large-...).
Predicate pushdown filtering enabled by the Snowflake Spark connector seems really promising. Lots of companies are currently running big data analyses on Parquet files in S3. Snowflake has the opportunity to grab a huge slice of the big data market.
Maybe its cynical/paranoid, but in this age of Theranos I must ask: is it possible their algorithm excels at showing you a reasonable looking number, rather than an accurate one?
This course is on the design and implementation of database management systems. Topics include data models (relational, document, key/value), storage models (n-ary, decomposition), query languages (SQL, stored procedures), storage architectures (heaps, log-structured), indexing (order preserving trees, hash tables), transaction processing (ACID, concurrency control), recovery (logging, checkpoints), query processing (joins, sorting, aggregation, optimization), and parallel architectures (multi-core, distributed). Case studies on open-source and commercial database systems are used to illustrate these techniques and trade-offs. The course is appropriate for students that are prepared to flex their strong systems programming skills.
[1] https://www.youtube.com/playlist?list=PLSE8ODhjZXjbohkNBWQs_...
What if Kodak sponsored an imaging class in 1990... what do you think they would have said about film vs. digital photography?
I think any increase in educational content is good, even if ‘bad actors’ are funding it.
I recently saw a criticism of Palantir which went: "The company has largely succeeded, they say, not because of its technological wizardry but because its interface is slicker and more user friendly than the alternatives created by defense contractors."
A lot of the most successful tech firms started post-dot-com are decent interfaces to not-particularly-revolutionary databases. In high-end consulting and investment banking, appearances are hugely important. You can't have trash decks. It's unsurprising to me that the same is true in defense and intelligence. You can get a roof over your head and breakfast at a trashy motel or the Ritz. Everybody knows the Ritz can command a much higher price because "its interface is slicker and more user friendly than the alternatives."
I think the same thing is true here.
Hotel restaurants are the same principle, except replace furnishing with food.
I stand by my claim. The relative differentiation in niceness is swamped by their mass produced boxness.
Ironically, my favorite road chain tends to be Aloft. At least they're upfront about their capsule-esque nature, in a sort of ironic/not-ironic way?
Least favorite: Embassy Suites. shudders It's like every Disney vacationing family's fantasy about what a hotel should be... packed with every Disney vacationing family. Omelette?
Don't get me wrong, there's a benefit to consistency of product (especially when you travel Su-F for consulting).
But that benefit, parent company consolidation, and economies of scale drive a net result of overwhelming homogeneity.
When I'm travel for business or putting my head on a pillow on a roadtrip, consistency makes my life easier and less stressful. I'm a gorilla-sized person :), I would rather stay at higher end hotel that provides an actual bath sheet than a marriott whatever where I have to call for 6 towels. Surprises aren't delightful at 10PM when you've been on the road for 15 hours.
When I start talking about the Ritz and high-end consultants, I'm discussing the interface, which of course includes the "far better beds, cleaner & safer rooms, better food..." and consistency you're trying to contrast with appearance. I would agree that those things are more than superficial and are extremely important to the experience of the user, because that's exactly the point I'm making.
The beds and concierge are nicer at the Ritz, and the interface (note: not appearance) and support are better at Palantir (or, as we're discussing here, at Snowflake).
I wouldn't be that worried about lock-in or being made obsolete. Business logic is going to be pretty easy to port between Redshift, BigQuery, Snowflake, or whatever comes next.
Redshift is backed by worker instances that have their own stores in what's basically an EC2 instance. It's definitely not backed by S3 like Athena.
Bigquery and GCS are both built on top of Colossus, but they have different layers in between them.
Since the object store APIs are almost identical across platforms, it doesn't matter that much which warehouse you actually use for production work. It's something that does massive SQL, imports data from S3, and exports data to S3.
No, most are going to be using SQL IDE's to query and export data.
This isn't even remotely true. Each has unique SQL syntax, and once you have few hundred or thousand queries written using vendor-specific SQL (be it date functions or JSON), it is non-trivial to migrate.
This can be said about most products and companies. What keeps them alive is how robustly they capture (and hold on to) the market, reduce costs through economies of scale, and innovate. This specific market is also very rapidly growing.
Lotus? Delphi?
Dunno if that's changed since Red Hat took them over.
Last year we recruited an attorney from a firm that still uses WordPerfect for all their documents.
Tavis Ormandy (@taviso) Tweeted: @mkolsek Funny you should mention that, I was recently curious if there are any console word processors. I discovered there's a community who still use WordPerfect 5.1 for DOS. They kinda sold me on it, got it working in DOSEMU. https://t.co/t6j0c1G3w1
The legacy players like Teradata and Exadata (from Oracle) really don't scale. Teradata has ~2B in revenue, Exadata is probably in the same range. That's all up for grabs but that's only scratching the surface.
Historically, only transactional data was dumped into the warehouse. Snowflake is selling storage at S3 price (plus you get compression so often ends up cheaper) while they are making money of compute/query. If they can provide all the right query abstractions (SQL, full-text search), in theory all data can be thrown in Snowflake. Yes, tech savy bay area companies can setup their own stack using Presto etc but rest of the world is not like that.
It sounds like a ton of these cloud infra companies have this product strategy (datadog, snowflake, elastic, hashicorp, etc)
My last company was an early adopter of Snowflake. And we tried Presto first, circa 2016 and Presto was sloooow. We were using vertica at the time and it was so much slower. Snowflake on the other hand was able to perform on the same order of latency as Vertica, which was pretty crazy to us.
Their Eon mode product is very similar to Snowflake, with S3 storage and semi-dynamic compute nodes, but they may not be as slick at marketing it or providing a UI.
Snowflake & BigQuery get the ability to have multiple customers on a large cluster.
It’d be cost prohibitive for a single smaller customer to have all that compute sitting idle for a few queries per minute.
Storage also benefits as snowflake/ BQ can shard your data across a much larger array of disk giving you better IO.
Think is it faster to drive a car 100 ft starting at 0mph and flooring it. Or to drive 100ft with a car which starts off doing 120mph
> The legacy players like Teradata and Exadata (from Oracle) really don't scale.
I get why Teradata gets labelled "legacy", but one of Teradata's main differentiators is scale. Teradata engineers have been tackling incredibly interesting scale problems (on many dimensions of "scale") for 40 years. Teradata has many customers who routinely manage and perform analytics on many petabytes of data.
> Historically, only transactional data was dumped into the warehouse.
That was once true, because initially that was all the data that companies had. However, companies have long since used data warehouses for all kinds of data — sensor data, text, behavioral data, product info/BOMs, vendor info, contract info, etc. — whatever's necessary to run the business.
> Snowflake is selling storage at S3 price…
This is important, but not unique. For example, Teradata's current product has native support for S3 and S3-compatible object stores, and you can query them just like any other database table, join that data with data in high-performance native storage, etc.
My experience of TD is > 10 yrs and then the multi-node version was substantially more expensive than the single-node version. Also, storage and compute was coupled which meant I had to pay for nodes even if 99% of my data was cold. That's a problem with RedShift too but not for Snowflake.
De-coupling storage and compute was a brilliant move by Snowflake. BigQuery can completely abstracted compute - you don't provision compute and only pay for data scanned. However, it gives you a sense of insecurity around cost - A single bad cron job running a query every sec can blow up your cost (real-life experience). Snowflake provides the best cost/performance tradeoff I have seen.
The honest answer is "it depends". Because Teradata is a different beast, per-query pricing can be significantly cheaper than Snowflake with high-volume workloads. It's worth trying both to evaluate cost and performance.
> Also, storage and compute was coupled which meant I had to pay for nodes even if 99% of my data was cold.
Yes, it used to be that everything had to into Teradata's high-performance filesystem. These days, Vantage's native object storage support means that you can keep that cold data in S3.
Only if you're using on-demand. Instead reserve some slots and pay flat rate. The minimum quantity is very low, and the minimum time is 1min.
Storage costs for S3 (or any cloud-provider object storage) are only one dimension of the price. The other is interaction costs which can get prohibitively expensive, for example if you accidentally forget to provide a partition key in your query predicate. Snowflake absorbs this cost if you use internal storage (or just copy into tables).
The benefit of native tables for all columnar databases is that it provides an optimized format with metadata for each column, which is then used to eliminate most of the data retrieval during query time. The more selective your query, the faster the results.
It makes sense to centralize, but only at a certain cost. Beyond that cost, it is better to just de-centralize because not every project can spend 4-5 months of meetings to spin up a DB.
The cloud changed this because it became an OpEx discussion and something you could spin up on your own. For non-production workloads, it comes especially obvious to do this.
> This still in no way answers why Snowflake is so valuable, though.
It explains it to a T! You have something you want, but internal company politics and territoriality keep you from getting it the way you want. An outside provider lets everyone get it for a bit of cash. It's basically the same play as Salesforce. It's not some kind of technical moonshot. It has to do with a modicum of technology delivered by a 3rd party who can avoid all of the internal friction.
The next founder who can think of this kind of play, then execute on it, will be the next Salesforce/Snowflake, and will probably have the ear of the same investors!
Edit: why are so many people downvoting this? Is there some other reason for Snowflake's valuation (aside from tech bubble playing a role)?
W/r/t value, the idea is a disproportionate of egress from Oracle, Teradata, etc will end up at Snowflake, hence huge TAM, SAM, and SOM.
But I can probably answer why Snowflake instead of Redshift (sorry, not too familiar with BigQuery)...
First of all it's cloud-provider agnostic so you can set up Snowflake on any or all of the 3 major cloud providers as well as set up replication between them directly or indirectly through their data exchange. Probably the most powerful feature is the way that Snowflake has the ability to scale (up or down) compute (vertically and horizontal) and storage independently of each other. Furthermore, you have the ability to scale compute down to nothing, and spin up "instantly" when the demand arrives. On top of all of this there is an incredible selection of functionality that i could go on and on about.
Edit: seems you’ve already answered this :) https://news.ycombinator.com/item?id=24265856
Marketing. Vast amounts of marketing.
I find that this occurs when an infrastructure team considers itself a "platform". The only supplier of an asset that everyone else demands can set the "price" of the asset as high as they want.
Eventually that blows up too -- re-composing data at the parent level (e.g., for quarterly financial reporting) becomes too exhausting and the company decides to revert to centralized services.
Until you want to join and then all of sudden it's not your problem. Then you end up with a "gang" of cowboy analysts running ad-hoc data-dumps against operational datastores affecting production uptime and stability only so that they can do a lookup between the multiple (source_table_column_count * source_table_row_count) sheets in their uber excel document.
I'm all for decomposing the monolith as long as you have a plan for recomposing when it's necessary.
If everybody cowboys their own storage, you're risking building up a legacy cruft that's very hard to work your way out of later.
The costs if recomposing later can be prohibitive, if a few particularly poor choices were made early on, and the hidden costs of inflexibility can bite too.
There's nothing wrong with outsourcing storage; that's not the issue - the issue is the culture in which it's easier to just not talk to the rest of the company (assuming the company is small enough to have any kind of cohesion in the firs place). If it's too expensive to talk (or even pick from a few common defaults) beforehand, how are you ever going to interop later on? You're getting the downsides of a large organization without the upsides.
Baffled both by how bad your internal infrastructure is and by how easy it is for you to buy stuff.
Probably other things, but many companies exist just to be alternatives to faang. If you’re good enough, you surpass that intention.
I've worked with a large Postgres cluster before (~1PB of data) and have been experimenting with Snowflake recently. I would say there's two clear technical advantages of Snowflake over Redshift. First is there's no maintenance when using Snowflake. You just signup for a Snowflake account, upload a CSV, and you can start querying the data. This is in contrast to Redshift where you have to manually provision a cluster, resize it as you add more data, etc.
The second is their pricing. Storing data in Snowflake costs the same as it would cost to store in S3. The tradeoff is you also have to pay based on how long your queries take. Depending on your workload this can result in a massive cost savings. If you access only small amounts of your data infrequently, it's like you're storing the data in S3 and you only have to pay a bit more when accessing the data. This is in contrast to Redshift where you have to pay for the full cost of the cluster regardless of whether you are actually querying the data or not.
Snowflake also has a ton of quality of life improvements compared to Redshift. One really nice thing is you can change the amount of compute used for any individual query. For example, if you have one specific slow query, you can allocate 4x the compute for that one query, pay 4x as much while the query is running, and get the query to run 4x faster (ultimately costing you the same amount as if you used 1x the compute).
One neat thing is there's ultimately only one "Snowflake instance" in each region. Everyone's tables are in the same instance, but you can only access the tables you have permission to access. This allows you to easily share data between different Snowflake accounts. You can store the data in one account and query it from another.
So the core value proposition is really strong and it also has a bunch of extra features that are all pretty useful at the end of the day.
This post focused on Snowflake solely from a business point of view. I'm considering writing another one that focuses on it from a technical point of view.
I work for VMware and get along well with folks who work on Greenplum. It's still doing massive workloads with massive amounts of data for lots of customers, has the ability to operate over blob stores with predicate pushdowns and recently merged up to parity with the PostgreSQL 12 upstream[0]. It's the fruit of a six year effort to return to the upstream from a heavily modified fork of 8.3. A truly monumental effort.
[0] https://github.com/greenplum-db/gpdb/commit/19cd1cf4b68faff2...
[1] https://www.infoq.com/presentations/snowflake-automatic-clus...
At read time, though, Snowflake's zone map is the same as Redshift's and Vertica's; you'll see similar pruning for many queries.
Redshift however doesn't prune during joins, which is a huge deficiency.
Snowflake looks more flexible about getting the data into its final ordering.
- late 90s: Early DWs like Redbrick
- early 2000s: Oracle, Teradata
- late 2000s: Shared-nothing Data Warehouses (Vertica, Aster Data, Greenplum) - bought up by Teradata, EMC, HP
- early 2010s: Hadoop and Hive
- late 2010s: Redshift and cloud DBs
- early 2020s: Snowflake
- late 2020s: probably something else...
All these technologies felt they were here to stay at the time, but they didn't. Will Snowflake be the exception? Maybe, but the odds are not nearly as great as their valuation implies.
But it does not explain why it is valued at 60B$. Or if that value makes sense.
To put a price on a company, I need to know in projection for the next 5-10 years what income they will generate to shareholders.
The fact that they are a great company does not guarantee they will generate an income to shareholders that value them at 60B$.
This is how bubble starts. Overvalue a great company, then overvalue fair companies just not to miss out and "in comparison" with the great companies, and in the end, overvalue nothing (if Tesla is great, then Nikola was the nothing).
Why is Snowflake so Valuable?
Stock market is not logical or rational.
Thanks for reading.
There are only 3600~ companies on the US markets [1], half of what was there in the 90s. There aren't many places to put lots of money.
The rise of private equity (PE) really has taken lots of the growth out of the public markets but since there are so few, and low interest rates, people are looking for returns. Lots of it is also pump and dump schemes loading up stocks and then short and distort. The problem is the growth is tapped on public markets now, PE drains it before it gets there, so more volatility games are being played. Couple that with less spending purchasing power in the lower/middle and that adds to the games as investment in new consumer focused companies isn't working as well when purchasing power is drained, M2V has hit a precipice [2]
[1] https://www.wsj.com/articles/where-have-all-the-public-compa...
I don't know how can you, as an investor, rationalize the price by any other reason than speculative increase in share price. Which is again driven by some fundamentals but mostly by hype. If they fail to gain almost 100% increase in revenue for the next year, that 60B is looking way too high.
Compare this with something like Splunk. They make high quality product and have over 2B in revenue (with a nice growth curve). Market cap 30B. So is Snowflake as of now worth two Splunks? I don't think so. I could see it maybe being half of that, 15B, if I squint real hard.
I never used snowflake, so hard for me to have a solid opinion of this company. I remember when facebook IPO'ed and people were like 'what? worth 100 billion? OVERVALUED'. and they were wrong in every way. so who knows? Though my gut doesn't tell me this company is the next facebook. coming out of the gate with a 70 bil market cap feels like all the growth is already priced in.
With facebook and on their IPO, I felt just their mobil revenue in 5 years time alone would be worth their valuation. But I had week hands. I think I bought them at 28 and sold at 22. I should have had more conviction because I truly did believe they were worth a lot more.
$60B is still too much.
It's odd that Buffet is in, it's a weird signal, because this is a weird era for markets: all other things equal, we are looking at .com-ish situations here and the timing would be ideal for a true crash.
That the world economy is shrinking by 10% and governments, major industries are going insolvent should be scary.
Perhaps investors think they are preparing for the 'covid future' but this may be a weird kind of inflation whereby everything else (including cash) is crap so they are piling into winners.
There is an emotion to a lot of stocks these days that is probably making every analysts job a nightmare - if the CEO or company is popular, it really messes with valuation.
More like Todd Combs and Ted Weschler...
[1] https://www.ft.com/content/b330e091-2a59-4527-b958-9213731a5...
No. That would require customers to grow their spend 58% every year. Might happen early on, especially as projects like this are usually staged in phases so the initial go-live usage is lower than when you get to full production. But projecting long term revenue growth on the basis of average annual 58% increases in per-customer revenue is simply ridiculous.
1. https://www.sec.gov/Archives/edgar/data/1477333/000119312519... ( see dollar based retention rate)
2. https://medium.com/@sammyabdullah/109-net-dollar-retention-i...
The shoe shine boy told me about it....
How easy/hard is it to fake an NPS score? Is this somehow regulated? Can the company only provide its most satisfied customers (which it knows beforehand) and only have them participate to get a good NPS?
> For every $1 of revenue Snowflake received from their customers a year ago, that same pool of customers are now paying $1.58.
Net Retention is more important, but in this case it also gives credence to the NPS number.
I don’t think it’s that great for software, but it’s very trendy. It can be gamed, like anything else, usually by sending NPS surveys to decision makers who aren’t usually the actual users.
Slightly better metric is Customer Effort Score (CEF) which shows how easy is to do business with a company.
In general the article persuaded me that Snowflake is a great company. But without some maths relating all these numbers to an estimate of future revenue, there’s no way to tell if it’s a great stock — at the current price. It’s the same problem that Tesla bulls have.
But everyone has heard of Snowflake.
And in contrast their Silicon Valley roots means a lot of their tooling/UX/data capabilities are ... undercooked. Their Web IDE feels like a throwback to 2003 Hadoop, their ETL capabilities are a joke, they don't support joins in views ...
And they've also squandered some opportunities to actually offer a differentiating "all in one" data processing experience for ad hoc/exploratory, BI/aggregated, and Big Data/AI/ML model crunching. For example, here's their garbage blog post on Spark SQL - https://www.snowflake.com/blog/snowflake-spark-part-2-pushin...
tl;dr when someone writes a Spark job that includes a filter against data in Snowflake, it's more efficient to let Snowflake filter the data before shipping it off to the (much more performant) Spark engine to do the actual analytical pieces of the query plan, instead of just shipping all the data over and letting Spark do the filtering.
Like ... wow, predicate pushdown is your answer?
Contrast with Azure Synapse providing Spark and SQL Server compute in the same environment; Databricks adding Delta Lake capabilities to be more schema-on-write friendly; Dremio building AI into their caching, and Starburst into their workload management ...
Anyway, I don't see any secret sauce, which means it's still just traditional enterprise sales cycles...
Perhaps you mean they don't support joins in Materialized Views (uppercase M)? We use Snowflake views with joins all over the place. Furthermore if views don't cut it for you, you can always use joins in UDTFs. Or if you really need joins in a materialize view (lowercase M), you can use change streams in combination with joins to maintain your own materialized view (table).
> instead of just shipping all the data over and letting Spark do the filtering
Forgive my ignorance, but in what capacity would this be less efficient? Doesn't it make more sense to reduce you result set before shipping it off to external compute?
And we're in agreement on the second part: my point is not that their solution isn't more efficient - it is - but it's treating predicate pushdown as some kind of deep synergy between Spark and Snowflake.
That is, it's more the marketing aspect - clueless execs see "Snowflake and Spark work great together" and a box is checked.
2) What is their stack?