Snowflake S-1
sec.gov
sec.gov
There's so much they can do from a user experience perspective to make it even better. The integration with Numeracy was a trainwreck, but the fundamentals of the DB are there.
Interesting to see they lose so much money, but I bet their margins have to be so thin running on the cloud. I wonder if they'll ever have to go bare metal to make it work.
Working with it was fraught with issues. Performance was mediocre at best, it was horribly expensive, Python and JS client libs had re-occurring issues with disconnecting and reconnecting. The advice given to us around scaling concurrent connections was bizarre at best. Teammates had numerous issues where it was clear corners had been cut in handling some edge cases around handling certain unicode characters. Their Snowpipe "streaming" implementation was...not good. The idea of having having compute workers that "spun up and down" sounded good in theory, but in practice lead to more bottlenecks and delays than anything else.
The AWS outage last year that prevented you from provisioning new instances essentially crippled our snowflake DB.
I almost go out of my way to recommend people _not_ use it. I keep seeing it pop up, but mostly because it seems they're doing what Mongo DB did in the early days and just throw marketing money to capture mindshare as opposed to being an actually good product.
We changed to ClickHouse and the difference was literally night-and-day. The performance especially was far superior.
They are always going to be less integrated and less infrastructure-cost-efficient than the native options (Redshift and BigQuery), without the R&D budgets and with incremental friction (sales) and risk (data privacy and cybersecurity).
AWS really should get around to buying them, like they should have bought Looker or Tableau or Mode or Fivetran or DBT, etc, ect.
Like, in a sane world I agree with you -- Redshift SHOULD have a crazy competitive advantage. But somehow they've been unable to execute on that goal for half a decade, and I don't see that changing quickly, given Snowflake's mindshare and growth.
Snowflake is better. Redshift has been really slow to execute. AWS is doing the world's worst job of articulating whatever vision they have for analytics. AWS's message is laser-focused on infrastructure folks and machine learning engineers (not analyst, data scientist, not absolutely anything else).
The higher you go up the stack, the slower and less meaningful, AWS's solutions feel. There is a fantastic job opportunity out there for someone to reconcile AWS's data analytics offerings. They have so much upside.
I'm still not betting on Snowflake winning a direct competition with their primary supplier. For the enterprise and the highly regulated: Redshift is good enough, already there, and they don't NEED the efficiencies that Snowflake makes available.
Example: you can play inside ball on storage infrastructure costs to get a 2x cost benefit at the expense of a lot of extra engineering. Better DBMS storage organization, which is available to any implementation, gets you 10x (or greater) improvement. Which would you rather have?
In fact, products like Redshift don't even really game the infrastructure prices. Costs to customers are comparable with Snowflake for equivalent resources as far as I can tell. They both charge what the market will bear.
1. You start out storing it on Amazon gp2 Elastic Block Store, which is fast block storage available on the network. It costs about $0.10 US per month per GB, so that's $102.40 per month.
2. Data (sadly) has a habit of getting destroyed in accidents so we normally replicate to at least one other location. Let's say we just replicate once. You are now up to $204.80 per month.
Now we have a couple of ways of reducing costs.
1. We could make the block storage itself cheaper thanks to inside knowledge of how it works plus clever financial engineering. However, the _most_ that can get us is about 5x savings, because prices for similar classes of storage are not that different. The real discount is more like 2x if we want to make money and be reasonably speedy. You likely have to do engineering work--like implementing blended storage--for this latter approach, so it's not free. So, we're back to $102.40 per month.
2. Or, we could build a better database.
2a.) Let's first build a database that can store data in S3 object storage instead of block storage. Now our storage costs about $0.02 per GB per month. Plus S3 is replicated, so we can maybe just keep a single copy. We're down to $10.28 per month but we had to rewrite the database to get it, because S3 behaves very differently from block storage and we have to build clever caches to work on it.
2b.) But wait! There's more. We could also arrange tabular data in columns rather than rows, which allows us to apply very efficient compression. Let's say the compression reduces size by 90% overall. We're now down to just $1.03 per month. Again, we had to rewrite the database, but we got a huge savings in return, like 100x.
The moral is that clever arrangement of data just about always beats financial shenanigans, usually by a wide margin. The primary reason that Amazon has done well in data services like Redshift and Aurora is partly that they have been extremely smart about data services, not any inherent advantage as platform owners.
Edit: fixed math error
I fail to understand this network effect. Is there any conflation here ? How does data sharing equate to network effect. Something is fundamentally not adding up here. If I share my data with 10 other customers, it should inherently enhance my experience. How does this happen with Snowflake ?
1) Building a common platform to upload datasets by anyone. e.g. weather data, retail data, govt data, other open data, or close data (copyright etc). They gave the example of COVID cases in their S-1 doc.
2) Providing mechanism for others to find data through a marketplace; some data is free, other only via payment (with diff monetisation models, e.g. per consumption, per month). Allow other customers to consume it as & when needed. Note, based on their S-1 doc, data is never copied when shared with others, so cost is limited to share with a wide audience.
3) More data on the platform, more data is shareable in the 'marketplace' and more data used by everyone. This increases the value of the whole platform through network effects.
4) Also opens up alternative revenue streams. e.g. more revenue through storage (more data on platform from different people). and revenue from shared data that is consumed (maybe)
Here is a company that is doing something similar in Australia. https://www.datarepublic.com/solutions/use-cases/data-collab...
If you have 80-100% utilization for a month, perhaps, but the beauty of Snowflake is that you can spin up a 3XL warehouse for a few MINUTES to get answers fast, and then shut it down again and don't pay anything.
Saying "you could run it on self-managed Spark/Oracle/Hive/SQLite" is approximately the same argument as saying "I can run a web server cheaper myself than paying Amazon for an EC2 instance" -- there are cases where that is true, but there are many, many, cases where the "on demand capacity" is the bigger benefit.
Is this why they're making a $350mn annual loss?
A million dollars a day loss would be a pretty big deal to me.
(Small caveat that someone probably gets commission on follow-on expansion revenue, but not that initial subscription amount.)
I had a couple of other minor issues and got very good response.
But Snowflake has really led the way to democratize Data Warehouse the past few years and educating the market. You can start on a $50/month plan, and in our experience, the pricing scale nicely with the value you are getting out of the data. Snowflake (and Bigquery) also made it a lot less scary to get started by having an easy way to ETL data from 3rd parties (google ads, Salesforce, prod DB, etc.) to your warehouse.
Thank you, Snowflake, for paving the path for startups like Census, Fivetran, DBT, Mode to help (data) engineers and analysts do more with their data
General analytics queries for the like of dashboards, CH latencies in the order of < 100ms, Snowflake about a second. Snowflake couldn’t do Geospatial queries when I had to use it, but I was getting responses from CH in like 40ms for a dataset of 10’s millions of points.
This guy has done some really in depth benchmarks: https://tech.marksblogg.com/billion-nyc-taxi-clickhouse.html
And https://tech.marksblogg.com/benchmarks.html
CH is one of the fastest non-GPU databases there.
Scaling redshift up and down was a nightmare. Tracking files on ingestion was a nightmare. Semi-structured data into structured data was a nightmare. i could also go on...
I'm a very early customer, and a big fan of SF if you couldn't tell.
- Performance from tables not cached on the warehouse instance is awful. That the price you pay for shutting a warehouse down.
- I wish it were cheaper. If you run queries against a warehouse 24/7, preventing from auto-shutdown, you better hope it's tiny. And even then, the cost might incentivize you to employ a different strategy entirely.
The other stuff you mention is them building a moat around your data lake. They're pretty good at that. I'd happily get locked into Snowflake for the moment. Redshift really does look like an amateur hour product in comparison to the tools you get with Snowflake.
Even then, as good as Snowflake is, our internal users went from complaining about the performance of Looker and Redshift to complaining about the performance of Tableau and Snowflake. I don't know if you can ever please anyone in this space...
I can't say enough good things about snowflake, and I have plenty of criticism to throw at hadoop, redshift, asterdata & vertica.
Like, is this a indication that a lot of people are trying to exchange their companies for hard cash as quickly as possible? It kind of looks a lot like that. This is what, the 3rd or 4th one to hit the HN front-page lately?
Is... is that a bad sign?
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor...
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor...
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...
I'm wondering which of these S-1s are particularly interesting to discuss in their own right, vs. which are just follow-up/copycat-style threads? Unity is getting a specific discussion (https://news.ycombinator.com/item?id=24261559) but the other ones seem pretty generic. Oh yeah, that's another relevant principle: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor....
(That being said, the end times may ALSO be upon us, but not because of this particular sign.)
http://www.bretswanson.com/index.php/2017/01/reasons-for-opt...
Is that supposed to sound attractive to institutional investors?
They spend a ton of money on Sales & Marketing (293,577k in the last fiscal year). That's what's driving a lot of their growth I assume and it's a lever they can pull back to increase profitability.
We'll also have a lot of elasticity features of snowflake shortly without sacrificing our performance advantages. (https://www.memsql.com/blog/the-future-is-bottomless)
(disclosure: MemSQL CTO).
Very curious about your performance issues?
The issue is we have billions of rows and very varied analytical requirements, so there are quite a few "pathalogical" queries.
Change streams are something we're looking at (as well as MemSQL, Clickhouse etc.)
Data Lake is usually object storage or other large storage pool with raw files. These can be different formats like JSON, AVRO, Parquet containing with strong schemas or unstructured data. Processing can be done by engines like Spark, Presto, Drill, etc that support less advanced SQL but more robust access across data files and storage locations. The point is to serve as a general dumping ground or "lake" of all the data and then manage it afterwards (including cleaning and moving important records to a data warehouse).
SQL Server is a single-node OLTP relational database but most database engines are fast enough now that you can do everything you need up to hundreds of millions of rows. Best SQL and feature support with full update capabilities. Some DBs like SQL Server have also added OLAP features like columnstore tables to further delay or eliminate the need for a data warehouse.
On Data Lakes: I often use an S3 data lake construct as a staging area for my Snowflake data warehouse.
DW is for analytics and reporting.
Data Lake is like many DWs together and other, often "garbage" data, which "might" be useful in future analysis, ML and stuff. It's the unstructured graveyard of data (joking). Schema is defined on read.
Just my experience. Glad to see them reaching for cash. They're effective at what they do.
Because I had the opposite experience - you literally couldn't pay me to use Snowflake again.
> Date Available for Sale in the Public : The 91st day after the date of this prospectus (First Release).
Edit: Seems like I was wrong. This is for current shareholders. I saw somewhere on the internet sometime in October. That's a wild guess though.
Would love feedback! Included some helpful quotes from this thread too on why Snowflake vs Redshift.
Their massive marketing spend is interesting, I suspect that they perceive themselves to be the first (or at least, strongest) mover in a once-in-a generation land grab.