Cloud Filestore, high-performance file storage for GCP users
cloudplatform.googleblog.com
cloudplatform.googleblog.com
Thanks
Further minor versions add other things that could be useful for cloud usage, e.g. in NFSv4.1 (https://tools.ietf.org/html/rfc5661 ) parallel data access, sessions, improved delegations, and in NFSv4.2 (https://tools.ietf.org/html/rfc7862 ) server-side clone/copy, sparse file support.
Does it support extended attributes? You'd imagine so, but you'd have also expected NFSv4...
Will it ever support SMB?
Can you set per-UID quota limits, both in terms of bytes and IOPS/disk time?
I work for Google and am the product manager for Cloud Filestore. I would have responded sooner, but I was busy with with announcement related events :)
If you (or anyone reading this) want to discuss any aspects of Filestore, My email is my Hacker news login with at google.com appended.
(1) Filestore is a zonal product and provides high availability (HA) within the zone. We are considering adding regional regional HA, but it's not entirely clear what the use case is, as there is a cost & performance tradeoff. I'd be happy to chat with you in more depth to understand what you would like to see here.
(2) It's hard to be precise about what latencies _you_ will see, as the set of benchmarks and workloads run against NFS is so varied. Anything I say here, will be true for some workload, but undoubtably is bound to find someone who can find a workload where it's not true :). So, TL;DR: YMMV, best to test your workload when the beta launches soon (signup to be notified when it launches in a few weeks at https://goo.gl/forms/Hx6XkobcwNo5DoA33)
(3) We support close-to-open consistency, but it's really up to the client. See this Linux NFS FAQ for details: http://nfs.sourceforge.net/#faq_a8. TL;DR: If you're running a Linux version ≥ 2.4.20, and haven't mounted with the 'nocto' attribute, then yes, you'll see CTO.
(4) We don't have any plans for pub/sub integration on the roadmap, but I'd love to talk to you about the use case (see info about my email addr above).
(5) Yes, we have NFSv4 support on the roadmap. We launched with NFSv3, because it's still widely used, and in many cases customers won't see any appreciable performance delta from NFSv4. That said, we agree that it is very important, and NFSv4 can often help wih some metadata heavy workloads, and has a more extensive authentication and authorization model which some workloads require. Ultimately we made a time-to-market tradeoff.
(6) For backups, we support any of the standard commercial backup software that's certified against GCP and can backup NFSv3 shares. We don't have a native backup solution planned, but we do have snapshots on the roadmap, which in some cases are sufficient.
(7) As to implementation, sorry no, I cannot.
And to answer a few more questions from the nested comments so this is all in one place:
* Snapshots are on the near term roadmap, and are very high priority for us to get supported.
* SMB, extended attribute, and quota support are all on the roadmap, and like NFSv4 are high priority
Unfortunately I can't be more precise about when to expect these features.
-Tad
So unless you run GCP at a very high % of provisioned capacity, it's more expensive.
I'm pretty confident EFS can't match GCP's 700 MB/s and 30,000 IOPS though.
For small block sizes, I got hundreds of IOPS.
Latency was pretty terrible in almost all cases.
I wrote a great deal on this topic for internal consumption. Those benchmarks weren't really meant for the general public. My biggest conclusion is that EFS isn't useful for any workload where performance is a concern. Unfortunately, it's priced so high, that I'd never consider using it otherwise.
That's still a drawback as your app needs to manage dynamic resizing. While aws may not be more expensive, it's more convenient
Also I haven't had a single issue with performance. It's not particularly busy though (~800 iops?).
So large files work well, but if you have many smaller files it's killer.
We built our app to do caching locally into NVMe drives though, so we async pull data from EFS and push back as necessary.
Definitely going to give Filestore a spin and see if it is a good fit for our use cases.
You get an account manager rep and upon occasion, interactions with AWS engineers, depending on what you're doing and how exciting it is.
One might say a lot about AWS - but I don't fault them on the customer service front for a serious business spending serious money.
The only constant with tech support is that it's highly variable and depends on many factors. Enterprise MSDN/SQL support is not bad though and you should've gotten a quick reply if you're using the proper channels with an appropriate support plan.
I'm sure there is some good use for it; I just haven't found one yet.
Will also be interesting to find out how they are solving/pricing the immense bandwidth/storage requirements needed to make such work practical.
The data all lives in the cloud, only the models and specs are sent back and forth. Zync is their rendering tech: https://www.zyncrender.com/
There’s nothing to prevent any application from doing full error checks on every disk i/o call, dealing with timeouts, etc. Except that nobody wrote that stuff into software designed for local disk.
Such as?
What's the corresponding story re security for cloud filestore?
I do not know if Cloud Filestore offers encryption on the NFS application layer.
[1] https://cloud.google.com/security/encryption-in-transit/#vir...
The application had grown and ran in production over a number of years using NFS for storage before it was moved into AWS, so this probably naturally steered devs away from trying to use the shared file storage as a high performance cache (eg they might try it, and find it to be slow, and figure out a different way of doing whatever they needed to do, probably using the db).
Our usage pattern for NFS did not generally involve needing to read or write many small tiny files with any degree of performance. E.g. we might need to do reads/writes of dozens of files, each of a few mb in a process, with the overall running time of the process being largely governed by our DB performance / patches of poorly written algorithmic code running in a slow scripting language. Amdahl's law - for us making the NFS access go infinitely fast or 5x slower hardly changes the overall total process execution time.
I can certainly appreciate that EFS is an abysmally slow for dealing with large numbers of reads/ writes of tiny files, in another part of the system one of my colleagues set up a read only cache of a large number of tiny files in EFS, due to the per-file write latency it took around a week to load in < 100gb of data.
EFS is not even in the same ballpark in terms of speed :/
Using persistent storage with Google sets me back a little over $175 per TB of SSD, per month, without networking factored in.
At $0.20, time a thousand for a TB, Cloud Filestore comes to $200. Let's see how the performance goes.
While it doesn't make sense from an Individuals perspective, for enterprises it certainly is pretty affordable.
I believe it is only multiple read only mounts — there can be only one writer.
Cloud FIREstore is a document-store database for Firebase, which is the GCP suite of services primarily for mobile apps.
Cloud FILEstore is a NFS file system mountable across multiple compute engine VMs.
(disclosure: I work in GCP, but not on this product. Just happy to see it go public for more users.)
Any one group trying to build a solution to serve all needs is likely in line for complication, which is why great partners and options continue to be a wise choice.
(Gonna duck out of this thread, etc. now since this isn't my release and I assume people who work in storage can go a lot deeper.)
It’s like choosing which car to buy. They all get you from A to B, but there are still many options to choose from. ;)
[1] https://cloudplatform.googleblog.com/2014/01/easier-faster-l...
Example use case, screenshots of websites where instances both request and write / overwrite them.
If you don't actually need a disk-based file system and just want to read/write individual files as objects, then object storage like Cloud Storage is your best option.
If that latency is low and your workload can handle disk concurrency well then it works fine. It helps if you use (or configure) a database with more sequential access and buffering for large updates rather than lots of random small writes, as well as spreading the data over several disks.
AWS EFS has latency problems which make it problematic but this product seems to have better performance profile which could work well.
I did a quick lookup in the MySQL docs (https://dev.mysql.com/doc/refman/8.0/en/disk-issues.html) and was surprised that this isn't really an issue.
Learned something new, thanks!
You shouldn't use multiple servers writing to the same volume for database drives, but other than that it's no different than any other disk that might lose connection. Most VMs "local" disks are still attached over the network anyway, emulating a PCIe bus interface instead of NFS.
Much of the EDA world though uses in-house tools, or a mixture of commercial and in-house tooling.
"Google Cloud Storage" could be a product but also the encompassing category of all the things you mentioned?
Also on the Firebase pricing page (https://firebase.google.com/pricing) they have a "Realtime Database", is that related to datastore or cloud storage?
They also have another item there just labeled "Storage". Is that one of the above?
And at the bottom you get "Google Cloud Platform" on the "Blaze Plan". Is that the products you mentioned that all start with "Cloud"?
No approach is "the best". They're all very different services and if you're going to use them then you would've read the overview anyway.
Thanks for the feedback. As the person that named both products, I can say we spent a ton of time debating this but we felt that the fact one is an enterprise file share and the other a document database service focused on mobile and web would mean very little conflict for customers. We will keep an eye on any customer confusion it might cause.
It explains why this is linguistically bad. Basically, a billion people on this planet don't distinguish between the l and r sounds, so for 1/6 of the planet, these names are identical.
Perhaps also worth looking at the screenshot in the blogpost.
You have in there:
---
Datastore
Storage
Filestore
---
So, datastore is not storage, nor is filestore. What is it storage storing if not data or files? Why are files not data? I have no idea what should go where.
> We know folks need to create, read and write large files with low latency. We also know that film studios and production shops are always looking to render movies and create CGI images faster and more efficiently. So alongside our LA region launch,
So I couldn't create, read or write before without low latency? I thought this was already a feature of your other products
> we’re pleased to enable these creative projects by bringing file storage capabilities to GCP for the first time with Cloud Filestore.
For the first time? I couldn't store files before?
I'm not trying to be an arse, but I really don't get from this what the key difference is from everything else you offer.
Objects. Cloud Storage is the S3 competitor.
> Why are files not data?
“Data” as in rows in a database. Like Dynamo.
Everything on a computer is data. The thing you’ve got to understand is that the terms we use, “objects”, “files”, “data” — these don’t refer to types of data, but rather to access paradigms for data. The semantics of their storage, indexing, mutability, etc.
An object is a blob of data named by a key, that you can retrieve entirely, or overwrite entirely, and where usually you automatically get a version history of old versions that have been overwritten that you can retrieve, with a cutoff for automatic GC.
“Data” is a structured tuple that a database knows how to index into, and sort by the columns of. You insert rows, update columns of rows by a key, or delete rows by a key.
“Files” are seekable streams where you can index anywhere into a file by position and then read(2) or write(2) data at that position, and where other clients can see those updates as soon as you sync(2), without needing to close(2) the file first.
All could be used to implement the other (S3 is implemented in terms of Dynamo rows holding chunks of object data, for example.) But each access semantics has use-cases for which it is an impedance match or mismatch.
> An object is a blob of data named by a key, that you can retrieve entirely, or overwrite entirely, and where usually you automatically get a version history of old versions that have been overwritten that you can retrieve, with a cutoff for automatic GC.
And yet they refer to the objects inside as "Files" and support seeking
https://cloud.google.com/appengine/docs/standard/python/goog...
https://stackoverflow.com/questions/14248333/google-cloud-st...
I know this is just bikeshedding about names and terms but it feels confused.
I think some of the confusion in the list is because of the mix of generic and product naming.
Data can be stored in datastore. But also in "spanner" or "bigtable", which are not parts of "datastore", or in "SQL" which is a language. Object can be stored in the object store called "storage" which is also within an entire category itself called "storage". So there's "Storage" which is a group of all these kinds of stores, and "Storage" which is a very specific type of store.
The reason native Japanese speakers struggle with "R" and "L" sounds is because they just have one phoneme to work with, which sounds (to a native English speaker) like a combination of "R", "L", and "D". If you aren't exposed to phonemes at a young age, it is difficult to expand your set later in life.
An analogous difficulty might exist for English speakers if a Chinese company came up with two product names which used the exact same sequence of syllables, but had "tonal" differences in pronunciation.
Hopefully they won't launch Cloud Pyrestore any time soon...
Floating plastic would go in the GyreStore.
If you had a bunch muck it would go in MireStore.
Thanks for highlighting. We were aware of this and working with the local sales teams to make it as easy as possible.
But it does make GCP's storage product naming even more confusing overall (after "Cloud Storage" vs "Cloud Datastore" mentioned above).
Please, use useful names that do tell what is it about.
How about naming one as Cloud FileStore and other as Cloud DbStore?
Wow. Don't mean to be rude but go for a walk outside and speak to 3 people outside the bubble.
Thanks for the feedback. We spoke to number of existing GCP customers for feedback, but it's fair to say we can always talk to more non-customers.
It's a clear blunt suggestion to get out of their bubble.
No one outside Google would hear that explanation and say "Yeah totally makes sense one of them is an enterprise file share and the other a document database service focused on mobile and web, crystal clear and very little confusion.".
What more do you want out of my comment for it to not be low effort? Write a 3 page essay about it carefully making a case based on peer reviewed scientific evidence?
I was confused when they introduced Cloud Firestore to compete with their Realtime Database (https://firebase.google.com/docs/database/rtdb-vs-firestore). Now it seems like they're doing it again with Storage vs Filestore, not to mention the horrendous choice of names.
Cloud Storage is an API-level object store (e.g. S3) that requires specific application support.
If you can discern between "now" and "not" then you can deal with "Firestore" and "Filestore"...
The difference here is between /faɪl/ and /ˈfaɪəɹ/, which is much more subtle. It comes down to the difference between /l/ and /əɹ/. The [ə] is an uncommon vowel in languages, unstressed, and mostly subsumed by nearby sounds. And worse, more than a billion people on the planet grew up speaking a language which doesn't distinguish the [l] and [ɹ] sounds (they're both approximants with only slight differences in articulation). So when you say "file" or "fire" these people can't distinguish which one you're saying, and when they say it they use something like the tap [ɾ] or retroflex [ɻ] instead, both of which sound ambiguous to native English speakers. Or some non-native speakers will use [l] exclusively, for both /l/ and /ɹ/.
Sure, they are one-letter away from the other, it's a fact. But to turn this into a problem, well.. no...
In the case of Google Cloud Platform products, many of them [1] are subject to the deprecation policy [2]. Basically it states that they'll give you one year advance notice of any intent to deprecate those products. This is functionally the exact same policy as that offered by AWS [3].
[1] https://cloud.google.com/terms/deprecation
[2] https://cloud.google.com/terms/ (see Section 7)
Rather, they seem to double down on all of them.
I guess you could say this is the replacement for Cloud Storage FUSE [1], but it's not and if it were I don't see the problem with that.