Other than lawyers and economists, does anyone ACTUALLY prefer this raw filing?
EDIT: Adding my preferred link: https://www.cnbc.com/2018/02/23/dropbox-ipo-form-s-1-prospec...
Other than lawyers and economists, does anyone ACTUALLY prefer this raw filing?
EDIT: Adding my preferred link: https://www.cnbc.com/2018/02/23/dropbox-ipo-form-s-1-prospec...
* Reliance and risks of Zynga in the Facebook S-1
* Customer acquisition costs in the Blue Apron S-1
* Growth specifics and positioning of algorithms in the StitchFix S-1
* Infrastructure costs in the Snapchat S-1
Besides, an S-1 filing is not written in legalease, it's written in plain language. One of the target audiences is street investors so it's meant to be accessible. I'm looking forward to digging into this one.
* Our business could be damaged, and we could be subject to liability if there is any unauthorized access to our data or our users’ content, including through privacy and data security breaches.
They have made progress. They managed to get SoC II compliance for all of their offerings. They now offer HIPAA compliant hosting as well.
Not that long ago though (circa 2013) I remember a series of articles that made it clear that DropBox employees had access to customer data.
That spooked me enough to recommend folks pair it with https://www.sookasa.com/ if they were going to use it.
You have to trust the company providing the service, right? Of course in practice, accessing user data should be tightly controlled and require good business reasons and levels of approval.
Zero-knowledge alternatives like Spideroak exist, but this approach makes them sacrifice features. (and doesn't appear very popular based on market share)
(We solve the problem by letting our large customers run their own servers, with their own authentication via single sign on.)
The problem is that the user experience for client-side encryption is awful! Every shared folder will need its own key, and users would need to manage and share their keys outside of our system. That is not sustainable.
But then the major feature set breaks down. Want to access your files in a browser? Not with client-side encryption. Want to email someone a hyperlink to a file? Not with client side encryption.
The major lesson is that the world operates on trust. We can only stay in business if our customers trust us.
But then the major feature set breaks down. Want to access your files in a browser? Not with client-side encryption. Want to email someone a hyperlink to a file? Not with client side encryption.
I do all of this with Boxcryptor. I might be misunderstanding you - do you mean decrypt it without first downloading it from the browser? Because yes, that’s not strictly possible.
But Boxcryptor implements a small wrapper around directories and generates a public/private key pair tied to email addresses. You can client-side encrypt a file with your - or anyone else’s - public key by moving the file into the directory. You can also change the file’s encryption to add or revoke access by multiple users.
If you wrap Boxcryptor around your local Google Drive, Dropbox, Box, etc. directory, it automatically client-side encrypts, then uploads new files. Then you can share a hyperlink to share encrypted files without exchanging keys with anyone. The usability is so great I’ve been able to use this with non-technical clients. You can even use your own key pairs.
Homomorphic encryption doesn't change that.
The only way to compare plaintext is to decrypt the whole thing. So either you must trust a centralized org (like dropbox today), or you must trust a single centralized key (that could be done with homomorphic encryption).
(Also the best homomorphic algorithms still make small programs take days to execute)
Each user generates a symmetric "user key", kU.
The plaintext of each file (or without loss of generality, block of data, etc.), pFile, is encrypted with a randomly generated symmetric key, kFile, producing the ciphertext cFile. pFile is also hashed with a cryptographically strong hash, producing hpFile. kFile is then encrypted with hFile, producing ckFile. The user encrypts pFile with kU, producing chpFile. Finally, the user takes the first N bits of hpFile (for N on the order of, say, 16 or 32), producing hpFileTrunc. The user then submits hpFileTrunc to the server.
The server is, semantically, just a list of 3-tuples: (cFile, ckFile, hpFileTrunc).
The server sees if it knows of the existence of records with the same hpFileTrunc value as the client's submission. If so, it returns them to the client.
The client then tries, for each record returned by the server, decrypting ckFile2 with the client's hFile value, potentially producing kFile. If this is successful, the client then decrypts cFile with kFile, producing pFile. Finally, it compares this pFile to the original. If it matches, a match has been found, and the client exits the loop. If not, (or if either of the two decryption steps failed), it continues to the next record the server returned. If there are no more records, the client instead submits the tuple (cFile, ckFile, hpFileTrunc) to the server, which stores it.
Finally (whether or not a match was found), the client stores chpFile locally, to be used when retrieving the file.
To retrieve the file, the user decrypts chpFile with kU, producing hpFile. They truncate hpFile, producing hpFileTrunc, and submit it to the server. They perform the same process described earlier to retrieve the matching pFile.
(Note: truncation may also be replaced by, or combined with, a second round of hashing.)
With this scheme, assuming secure primitives (authenticated encryption and hashing), I don't believe it's possible to learn any information about a file unless you already have its contents.
So the server can tell if you're accessing (storing or retrieving) a particular file if and only if the server knows what it's looking for.
TL;DR: you can totally construct a scheme that allows meaningful comparison of plaintexts!
But... this is probably a bad thing. Comparison of plaintexts is a vulnerability: the server being able to see who's storing a particular "bad" file has a real impact on privacy. And likely more subtle impacts, too...
> The server sees if it knows of the existence of records with the same hpFileTrunc value as the client's submission. If so, it returns them to the client.
And by doing this, provides a way for clients to verify if any user on the file storage server has this file. So if I wanted to know if your mozilla thunderbird has a mail I have the source to, I simply try to store this and get these duplicate records.
Most people would consider this extremely unacceptable.
> The client then tries, for each record returned by the server, decrypting ckFile2 with the client's hFile value, potentially producing kFile. If this is successful, the client then decrypts cFile with kFile, producing pFile. Finally, it compares this pFile to the original. If it matches, a match has been found, and the client exits the loop. If not, (or if either of the two decryption steps failed), it continues to the next record the server returned. If there are no more records, the client instead submits the tuple (cFile, ckFile, hpFileTrunc) to the server, which stores it.
Why would the client have the keys to files stored by other users ?
Unless you mean that you can only deduplicate within a single client, in which case that's of much more limited use (and I might add, your encryption scheme is way more complex than it needs to be).
Yes. This is the reason you don't want this property (being able to deduplicate encrypted files)!
But you can provide it, while still providing meaningful security against other attacks.
The client has the keys to files stored by other users because the keys are the hashes of the plaintext, and the client can hash its own plaintext when it has the file.
(Note a trivial modification to this scheme, solely client-side, allows for certain files to be totally secure, with the cost of them being exempt from deduplication)
Personally I find only people explicitly authorized have the key to be the whole point of security. And you're suggesting this as a solution to the problem that organizations providing file storage could see what files you're storing.
Under this scheme, it wouldn't just be that organization, but everybody who is a client, that could see what files you're storing (or at least verify if you're storing a particular file or not)
So I find your assessment:
> But you can provide it, while still providing meaningful security against other attacks.
Very dubious indeed, especially given the context of securing centralized file storage, where the whole point would be to deny others access.
I mean it's a true statement, because you don't specify what "other attacks" are.
I posit that given that this system leaks the plaintext of your files I find it strictly worse than just giving Dropbox or Microsoft access to my files.
You can do this today, with Dropbox or whatever else- anything that does deduplication, if it saves bandwidth by not asking for files it already has.
You can't tell who is storing a particular file- only if anybody is. Does this leak information and impact privacy? Yes! But it still provides other useful properties.
If you have a copy of a file, you can see if anybody else does- a boolean value. (And if the server is malicious, it can tell who does (if it logs).) If you don't have a copy of a file, you can learn absolutely nothing about it.
So, for example, if a user uploads a, uh, personal image to the service- with Dropbox, in theory (they likely have strong organizational and technical controls against this sort of thing, mind you) if the server is malicious they can view that image.
With this scheme, the server can't.
On the other hand, if you, say, save a file containing only your social security number- or a similar low-entropy value- the server can crack the hash and decrypt that file. That's the price you pay for being able to deduplicate.
(Perhaps one could only deduplicate large files- thus handling the case of movies, music, Ubuntu ISOs, large system files, etc. To implement selective deduplication- if you want a file to not be deduped, replace all uses of its hash with, instead, a unique random value to identify the file. Server requires no modification.)
https://www.dropbox.com/help/syncing-uploads/upload-entire-f...
So, when quoted out of context it sounds really extreme.
The way Stitch Fix talks about it in their S-1 makes it seem like the latter is the priority. I'm not yet convinced that the practical value driven by algorithms at Stitch Fix is up to par with how much they talk about it.
I was interested in growth to understand both their growth rate but also to get a feel if it was driven by increasing user acquisition costs like Groupon, Blue Apron, etc or if it was organic.
I never tried myself but it seems quite feasible to build a decent profile of someones taste in fashion from a bit of data.
Fast fashion gave us the logistics (no more 9mo from concept to store). But Zara and friends still supply only the major trends. We're still missing for someone to reliably market the "long tail".
I've thought about this as well. How would you personally try to build this profile for users, then market to them based on it?
That could give you an starting point. But I believe the main issue is that fashion products have, by definition, short shelf life, so you can't run the algos on SKU data. Then you can use deep learning on product images + user categorical data to try to predict preferences, maybe simple binary classification?
I guess using images as input should give better features than textual description.
You would have to tag the hell out of these photos, right? Disambiguating preferences is the challenge -- the user may have liked images #3 and #7, but why? Specific items, or the color palette, or the silhouette, or just the model? A post by Chicisimo on the front page addresses these hurdles [1].
I've also seen some of these image recognition apps in action -- they pick up on patterns and color very well, but struggled with silhouette.
[1] https://hackernoon.com/how-we-grew-from-0-to-4-million-women...
That's part of what makes this industry hard.
(which even if you disregard everything else is a really good line to take if you're after recruiting good machine learning engineers!)
(Note: I'm biased and tweetstormed about Stitch Fix's most recent blog post earlier today. https://twitter.com/achompas/status/967085860763193345)
It seems similar to talking to people at Google in the early days. The thing that they cared about from an engineering point of view seemed weird and they language was alien. 3 years later the rest of the world hits the same problem and I remember the conversations and think "oh so this is what they were talking about".
The standouts to me were -
* They cut costs on an absolute and relative basis for the last two years. This is fantastic and I hope the trend continues.
* I don't understand how the $112 ARPU number foots with their pricing. They are telling a story that "teams" is driving growth but on the surface that's not reflected in the ARPU number. There's some sort of promotional discounting that's happening that isn't exposed here. I hope that they are aggressively managing their promo strategy internally because unchecked it can tank a whole company (see GAP).
* No idea where they expect new users to come from given they have 500M accounts. Presumably people have 2+ dropbox accounts (personal + work). I'm interested to know how many unique active users dropbox has.
I'd imagine there's quite a few with dozens, as they offer more space (up to a bit over 30 referrals) if you refer new users. I remember "referring" myself with a new email address whenever I'd need an extra 250 MB.
500M accounts, but only 11M paying users.
To be honest I'm not super familiar with the comps to know how good or bad their numbers are relative to others. But just from their own reporting they look like they are on track. That's not by accident. It's very likely that constructing this trend has been a major focus for the company in the last few years.
Facebook was at $1/user/year in their S-1, Twitter was < $1.
From the SNAP S-1:
"We rely on Google Cloud for the vast majority of our computing, storage, bandwidth, and other services. Any disruption of or interference with our use of the Google Cloud operation would negatively affect our operations and seriously harm our business."
"We have committed to spend $2 billion with Google Cloud over the next five years and have built our software and computer systems to use computing, storage capabilities, bandwidth, and other services provided by Google, some of which do not have an alternative in the market."
It usually goes POC->Cloud provider->Your own gear
Github: https://githubengineering.com/evolution-of-our-data-centers/
Backblaze: https://www.backblaze.com/blog/our-secret-data-center/
Twitter: https://blog.twitter.com/engineering/en_us/topics/infrastruc...
LinkedIn: http://www.datacenterdynamics.com/content-tracks/design-buil...
FastMail: https://www.fastmail.com/help/ourservice/security.html
Stack Overflow: http://highscalability.com/blog/2014/7/21/stackoverflow-upda...
Wikipedia: https://meta.wikimedia.org/wiki/Wikimedia_servers
OpenStreetMap: https://blog.openstreetmap.org/tag/infrastructure/
The Internet Archive: https://www.theregister.co.uk/2017/11/16/head_like_a_memory_...
Gitlab tried, but didn't have the necessary in-house experience before they made the attempt: https://about.gitlab.com/2017/03/02/why-we-are-not-leaving-t...
Instagram was migrated from AWS onto Facebook's infrastructure: https://www.wired.com/2014/06/facebook-instagram/
WhatsApp was migrated from IBM to Facebook infrastructure: https://www.cnbc.com/2017/06/07/facebook-planning-to-move-wh...
Hacker News and Pinboard (acq. Delicious) run on a single server.
It's not hard, but you do need to know what you're doing and have resources to do it (most orgs rent colo space in someone else's datacenter, they don't build their own). There's a reason AWS margins are so high (which leaves a lot of cost savings to be had when your workload isn't highly variable). Any questions, email is in my profile. I spent ~16 years building data centers, hosting environments, infrastructure, etc.
But in any case, the examples given are B2C free to use products which are generally going to provide... not a ton of revenue per user.
Their paying users have increased, but the revenue per user has decreased (based on the S1...which I assume is because of enterprise deals).
This has been explicity stated by the saas business ‘gurus’
So even at 3$ per user that goes lower pretty quickly.
Plus they get alliance w/goog. Aws is a force but if i had a choice id have goog as a best friend over amazon.
Preferably id build my own data center and keep those assets.
If by modern, you mean 2015. Things definitely changed by 2017.
"In its disclosure, Snap has said that it is contractually obligated to “spend $2 billion with Google Cloud over the next five years and have built our software and computer systems to use computing, storage capabilities, bandwidth, and other services provided by Google.” Of the current losses at Snap, more than 80 percent of those funds go straight into Google’s pockets."
https://www.recode.net/2017/2/7/14526832/snap-ipo-snapchat-s...
https://www.forbes.com/sites/quora/2017/02/22/could-snapchat...
https://www.cnbc.com/2017/02/09/snap-cloud-bill-aws-google-c...
EDIT: @dfee How that revenue is recognized is usually based on when the services are delivered, but I am not an accountant.
When talking about deals that are the enterprise valuations of Fortune 500 companies, do you value the sale as if it were an acquisition? Or, some other way?
The Company’s cryptographic systems depend in part on the application of certain mathematical principles. The security afforded by the Company’s encryption products is based on the assumption that the “factoring” of the composite of large prime numbers is difficult. If an “easy factoring method” were developed, then the security of the Company’s encryption products would be reduced or eliminated. Even if no breakthroughs in factoring are discovered, factoring problems can theoretically be solved by a computer system significantly faster and more powerful than those currently available. If these improved techniques for attacking cryptographic systems are ever developed, the Company’s business or results of operations could be adversely impacted.
[1] https://www.sec.gov/Archives/edgar/data/932064/0000950135000...
https://www.reuters.com/article/us-usa-security-nsa-rsa/excl...
Investors. The prospectus filing is informative and you don't have to read it all to get good idea of the company.
The summary from CNBC omits all the details and you can't even trust it to have the numbers correctly because if they screw it up, it has no legal consequences for them.
i am not an economist or lawyer and prefer it
i don't have to worry about a writer/editor injecting personal opinions that i could care less about
In a mere 160 pages. Unless you know what to look for it most certainly does not get to the point rather quickly.
Feel free to link to "news" (read: opinion) posts about the filing.
I was really excited to see they cut their infra costs while supporting more revenue (bottom of page 71), that bodes very well for their future prospects.
Apparently they do, since it's at the top of HN. If they didn't, it wouldn't be at the top. Alternatively, my guess is that many HN people saw this document, upvoted it, and then went to Google for more context.
it contains the "source" facts with limited embellishment. there are all kinds of regulations that essentially result in SEC filings being fairly standardized, limited-BS documents. you dont have to worry about distilling commentary and opinion with fact, generally
for commentary and context, equity research reports are generally helpful as long as you're aware of the inherent bias. articles online can also be helpful, but generally the quality varies, and you have to spend commensurately more time fact checking
once you learn how to ctrl-f the right terms and understand the general skeleton of each type of report, it is actually pretty easy to find the relevant information