Cloud Firestore: A New Document Database for Apps
firebase.googleblog.com
firebase.googleblog.com
We built it because we know it can be challenging to build complex apps with our original database -- Firebase Realtime Database -- where we optimized for ease-of-use & real-time sync over querying functionality. For more info, see this comparison between the two[2].
This is a open beta launch, so the product has limitations[3] you should be aware of. We’ll be working to remove/raise these before General Availability.
A note on naming: Firebase launched in 2012[4] with a single product, the original database. As we added products (Firebase Auth / Firebase Hosting / etc), we began calling the original product ‘Firebase Realtime Database’, or RTDB for short. If you haven’t followed us lately, we’ve grown to become Google’s app platform, and now have 16 products to help you build/grow apps.
We’re grateful for the support HN has given us over the years and we hope you enjoy Cloud Firestore!
[1] https://cloud.google.com/datastore/ [2] https://firebase.google.com/docs/firestore/rtdb-vs-firestore [3] https://firebase.google.com/docs/firestore/quotas [4] https://news.ycombinator.com/item?id=3832877
Thank you!
Transactions: If I want to insert 2 documents, but zero if either fails, what does that look like with Firestore?
And if I store a document and run a query milliseconds later, will the query include document I just stored? Or are queries eventually consistent?
So yes, if any of the operations in the transaction fails (and cannot be retried) the whole transaction will fail. It's atomic as you'd expect. If you've ever used transactions in Firebase Realtime Database you'll be happy to know that the transactions in Cloud Firestore are much easier to use since they don't require selecting a common data parent.
Queries are always strongly consistent. If you do a write-read sequence you'll see the value of the latest write.
PS: Hi Chris! :-)
Providing a one-size-fits-all solution here is probably impossible, but it seems like it would be nice to provide some mechanism to be notified that you're making edits based on stale information. If such a mechanism existed, it would be easy to add a bunch of canned merge strategies. In doing so you can probably teach people a little bit about the pitfalls they're likely to run into (these sorts of bugs are insanely difficult to track down), while not really making them do much work.
The approach we've taken in Eve is that we can't solve all these problems for you, but we can at least let you know that things can go sideways and prompt you to make a deliberate decision about what should happen. It's amazing how helpful that ends up being.
Can you link to somewhere where this layering is explained?
The www.firepad.io site has documentation on how to use the editor, but I'm interested in how "OT on top of last-write wins" is achieved.
But basically, sync is split into two halves: writes and listens. Clients store pending writes locally until they're flushed to the backend (which could be a long time if the app is running offline). While online, listen results are streamed from the backend and persisted in a local client cache so that the results will also be visible while offline (and any pending writes are merged into this offline view). When a client comes back online, it flushes its pending writes to the backend which are executed in a last-write-wins manner (see my answer above to ibdknox for more details on this). To resume listens, the client can use a "resume token" which allows the backend to quickly get the client back up-to-date without needing to re-send already retrieved results (there are some nuances here depending on how old the resume token is, etc.).
Very timely ;)
Thinking about https://news.ycombinator.com/item?id=14359801, what improvements have you made in the last 139 days that would make us want to build on this?
What sets Amazon apart from many other companies is their reputation for relentless customer service. Customers pay close attention to the quality of answer or lack of answers to questions like this one.
Every cloud company that asks us to invest our time in their proprietary API is asking us to trust them. If you violate one vocal customer's trust, we're going to notice. If you don't make things right, we're going to notice. Even before I read the article, I first read the comments to gauge the reception because I care more about the perception of trust before I even consider using something like this.
Customers don't always stop trusting companies for fair or valid reasons, but that doesn't matter. Trust is much more of a visceral thing than a cerebral thing.
So here's my question to other HN readers. What would Google need to do to improve its reputation for customer service? What would they need to do to make that commitment resonate with you on a visceral level?
I hear this about AWS all the time, and I've even experienced it myself. One client some time ago had an AWS novice who was confident they could handle setting up the auto-scale groups. They made a small mistake, which lead to the scaling trigger being permanently on, and it auto scaling to 1024 ec2 nodes within the first half hour. Immediately after deploying a fix for the scaling trigger, I phone AWS to see if there was anything that could be done about the cost of this accident, and I had a small monthly credit to compensate, and a respectable discount on the next appropriate AWS training course so that the novice could learn more and hopefully avoid future mistakes.
THAT is the kind of story I want to start hearing about Google Cloud. "I used it and did not have problems", and "It works and was cheaper than AWS" is just not enough.
There is very little time for developers of a system to use on support, so how do you go about building the baby padding around them? I would imagine at least 97.5% of support requests are a problem on customers end, so just how much support staff would you need for a global service like this?
They didn't go back retroactively and ask for back payments, but just started charging them accurately going forward.
Of course, that bill caused shock for some people, as they suddenly realised, "Oh s*it, I'm being charged accurately now, my bills are huge!!!"
In this customer's case, apparently they were also doing something bad with TLS tickets and not setting keep-alives - basically causing them to spam Firebase with connections:
https://news.ycombinator.com/item?id=14357270
Of course, there's the other issue of a communication breakdown which they said they're working on.
Bug 1: we under-reported bandwidth (in particular SSL overhead)
Bug 2: we were not enforcing quotas for all accounts.
For most users, the fixes had little-to-no impact. For a few users who were using the Realtime Database with large volumes of small reads and writes, the impact was large. You could mitigate this impact by updating your client code, but unfortunately a user who had shipped code to their IoT devices couldn’t. This user was also simultaneously forced to upgrade to the Blaze pay-as-you-go plan due to quota enforcement on our $25/mo Flame plan. These combined resulted into a large billing increase for this user. We weren’t quick enough to provide this user with credits due to poor internal communication).
To address these we have (1) worked to make billing more transparent on Realtime Database and (2) are working on improving support.
1a. We rewrote our documentation to add more detail on billing mechanics and how to optimize bandwidth (https://firebase.google.com/docs/database/usage/billing).
1b. We rewrote the documentation for our profiler tool which was confusing to many developers (https://firebase.google.com/docs/database/usage/profile).
1c. We now have better alerting for if/when we find errors in our codebase that can impact their bill (up or down).
1d. We will soon be releasing (spoiler alert) a new monitoring API to let developers directly analyze their database billing and performance data.
2a. We raised the quota on free technical questions from 5 => 10. Questions on accounts/billing/bug reports are still unlimited
2b. We worked to increase Support CSAT. It is up by 15% since the billing issue in May.
Finally, the new database we’re launching today, Cloud Firestore, has daily budgets. You can use these to set exactly how much you’re willing to spend per day (more here: https://firebase.google.com/docs/firestore/usage#limits) We’ve also got extensive pricing docs: https://firebase.google.com/docs/firestore/pricing
I hope this answers your question!
Where can I read about this?
Look for the optimize billing section
This is a totally new product that builds on what we learned from those two products, sharing some of the best features of each and bringing in some totally new things as well.
All three are (in my opinion) great choices for building a new app today so make sure you evaluate all options.
Cloud Datastore has been around in some form (started as App Engine Datastore) since 2008.
Cloud Firestore was a massive multi-year joint effort between Firebase and Google Cloud. Google is investing heavily in both, and this is a big deal for our teams.
After all, if they are truly confident, there should be little cost in doing so.
I'm not sure what other guarantees would even make sense to offer. If anything, I'd look at this announcement as ongoing proof in the magnitude of investment Google is making in Firebase and Cloud.
Maybe publish a long term (5+ year) plan / roadmap? idk.
Roadmaps are subject to change and even more subject to be delayed, publishing them tends to disappoint more than reassure. If we gave a forward commitment the questions would just be "why not longer?" or "what happens in X + 1 years?".
All we can do is say what I'm saying now: we stand behind this product 100%, we think it solves real problems for developers, and we really hope people will try it out and find it useful.
I get the doubt, truly I do. But Google's incentives are clearly aligned with Cloud Firestore's success: if you folks use it and grow your app to be successful, we make money. If you use it and really like it, you're more likely to use Firebase and Cloud's other products, which will make us even more money.
the issue is the inverse is also true: If you folks don't use it, we don't make money, and we deploy these resources elsewhere. See Parse
I'm struggling to see what this is giving me that I didn't have with RTDB though
The queries seem to be doing what orderByChild equalTo startAt endAt limitToFirst and limitToLast were already allowing.
Is there mostly a performance gain, or am I missing something that I can do now that I couldn't do before?
Cheers
---
We can now apply a filter and sort in the same command
We can chain filters, though they can only apply to one field if they specify a range
We can make a query that will only return documents, not the subfield data (i.e. what a lot of us were doing with the deprecated REST interface shallow=1)
Queries in large data sets will be faster I guess?
Things it'd still be nice to see -
References that allow you to pull down related documents with a single query
OR queries
---
It's been a bit of rollercoaster seeing this announcement..
"Oh wow we can do proper queries now! Oh wait no, the queries don't seem to allow more than what we had before.. oh wait there are some improvements, it's a bit better now.. oh the shallow querying will be super handy! and they do geopoints.. but don't appear to have any way of searching for radius.. but they say they will soon"
Anyway, good work all the same :)
It would be nice to develop a way in Firestore for users to be able to pay for the number of shards assigned to their workload to reduce relatively artificial limits such as write limits within collections and index update rate.
That way, developers wouldn’t have to plan to leave the platform if their app is successful.
One of the things we liked about Firestore is that it takes the best practices of Realtime Database and makes them more explicit. Before, your RD database structure would look like `collection/{id}` and `collection_sub_collection/{id}/{sub_id}` in order to avoid loading sub-collections in top-level queries. With Firestore, this collections pattern is now part of the API itself, and sub-collections aren't fetched when the parent is fetched.
Another feature we liked is that transactions are no longer limited to a subtree of your database. Before, you would have to structure all of your transactional data under a single path. This would sometimes lead to having to pile-in unrelated data into a single object, such as adding payment data under a users object instead of a separate collection, so that you could atomically modify both user and payment data. With Firestore, transactions are global, so this isn't a concern anymore - we are free to structure our data in any way that makes sense for our app.
Overall, we had a great experience with Firestore during the alpha, and we'll definitely be keeping it as part of our technology stack. Congrats on the launch!
(disclaimer: this post isn't sponsored by Firebase)
While I like self hosting (was previously a RethinkDB user), from a business perspective, it doesn’t make sense to spend time on operations if it doesn’t give you a competitive advantage. It’s going to be very difficult to outpace a business that only has to focus on development versus one that has to do development and operations.
Your data is your most valuable asset, and by using this you're locking it inside Google servers. If they decide five years from now to discontinue it, or to raise the pricing 10x, you're screwed.
Are most developers only working on short term projects? Why would you put yourself in such a situation instead of using open source technologies that can be deployed anywhere?
You'd be able to move your data off of firestore. And there's legal business contracts around pricing. Google can't just raise pricing 10x overnight.
It's a very expensive move that I don't think people consider when choosing this kind of solution.
If you're concerned that a service might shut down then you need to architect your application with that in mind, in which case how much of a rewrite is necessary is essentially up to you. Usually there's a tradeoff between going fast and engineering solutions that will work in the long term. Most startups never get to the stage where they need to swap out a service, so closely tying your application to a service is probably OK at the start.
If the service that shuts down is reasonably popular though it's likely there'll be very little code to change. API-compatible competitors will pop up to replace it. It happened when Parse closed.
[1] https://appbase.io [2] https://store.docker.com/images/appbaseio [3] https://appbaseio-confidential.github.io/streams
It seems a lot of developers use the HN stack: eg things thatvwent thru YC or that they hear other devs talking about a lot on HN.
That said, I really hope there are plans for some full text search ability beyond the current suggestions[1]. I would very much like to ditch Elasticsearch in favor of db engine provided search. Even a small subset of the Elasticsearch/Solr feature set (similar to the full text search capability now available for Postgres[2]) would be a very welcome addition.
[1] https://firebase.google.com/docs/firestore/solutions/search [2] https://www.postgresql.org/docs/9.5/static/textsearch.html
I have a client who has invested significant amount time and money in Firebase Realtime Database. Now with this move, I am not sure if Google will support Firebase Realtime Database for next 5 years. So a full rewrite might be needed.
Once more it seems that going Cloud Native on one of Google's proprietary tools is very risky. I know many of the readers will say that Firebase Real-time Database is still supported. But the main question is: will it stay supported for years to come?
Google please please make an announcement and make a commitment to keep Firebase Realtime Database alive for "X" years to come. Otherwise, you are just making us developers lose faith in you.
Regarding deprecation: you can be comfortable continuing to build on the Realtime Database. We don't intend to deprecate either database, since both are useful in different situations, depending on what you're building. We recommend using the Realtime Database for a number of usecases[1]
We're not posting a "Realtime Database will be supported for X years" statement because many may interpret this as "the Realtime Database is deprecating in X years", which isn't the case.
[1] https://firebase.googleblog.com/2017/10/cloud-firestore-for-...
Can't you just pretend this new feature you just learned about 2 hours ago didn't exist if you don't want to use it?
Isn't releasing this very feature improving Firebase, like you're demanding?
From linked page:
> designed to easily store and sync app data
Maybe I'm misunderstanding, but that's not what I understand a "document" to be.
For me a document
- is a file that can be stored on a file system
- can be send via mail
- is using a standard document representation so can be used by different applications, such as markup (XML, HTML, SGML), or maybe PDF
Just answer.
Besides, a "document database", like many term in computing, can be very confusing. Come on, we all had to be explained what the difference is between a software server vs hardware. This is not different.
Smells like ring voting: 4 downvotes in < 0.5 min? @dang?
As the site has grown, HN has become a bit of a mouthpiece for large organizations through these de facto voting rings.
Best idea I have is for HN to add a profile field like: "Organizations: [google]" which would prevent voting on any Google-related submissions. It could also add a disclaimer in each comment of these submissions, so users wouldn't have to remember to do that.
In don't mind (in fact, appreciate) this aspect of HN as long as it's done openly and civilized.
> Best idea I have is for HN to add a profile field like: "Organizations: [google]" which would prevent voting on any Google-related submissions. It could also add a disclaimer in each comment these submissions, so users wouldn't have to remember to do that.
That's a brilliant idea.
I love that suggestion, but how would you validate it? Registering company domains would be an exhaustive process.
What's ring voting ?
Coordinated up- or downvoting from multiple accounts
> Please don't accuse others of astroturfing or shillage. Email us instead and we'll look into it.
> Please don't comment about the voting on comments. It never does any good, and it makes boring reading.
I just noticed it's much easier to understand a new datastore by reading its limitations (usually carefully omitted from pr articles or documentation).
For example, I suspect that Firestore must be built on top of Spanner infrastructure, as it's the only way to get usable cross-datacenter many-row transactions. And Spanner's limitation is it's high price. And if it's not Spanner, but more like Cloud Datastore or the old Megastore, then there should be limitations on transactions.
Sounds amazing anyway, more datastorage choices is always better for the world.
For example Cloud Firestore can't do native full text search, but we made sure it's easy to integrate with a search provider like Algolia.
As for your guess about Spanner, you're right that Cloud Firestore uses the same technology as Cloud Spanner to ensure consistency at scale.
Would love to know what you think about the product once you've tried it!
I believe this custom auth layer is best handled by allowing code with restricted api access and time constraints. The technical approach cloudflare used to implement its edge workers by embedding v8/js [3] would be perfect here.
[1] https://firebase.google.com/docs/firestore/security/get-star...
example:
service cloud.firestore {
match /databases/{database}/documents {
match /{document=**} {
allow read, write: if false;
}
}
}
[2] https://firebase.google.com/docs/database/security/[3] https://blog.cloudflare.com/introducing-cloudflare-workers/
We're working on opening up the tools to work with the rules language so you can test and iterate more easily.
Someone tell me they’ve found the holy grail?!
A SQL-ish API might even be better, although I can see some pitfalls. Maybe base it on GraphQL as a compromise?
Either way, I'd love to have something like this.
It is exactly what you're asking about.
Disclaimer: I work at Hasura. Feel free to ask me anything you want to know about.
https://docs.hasura.io/0.14/ref/cli/hasuractl.html
:(
Please login to https://dashboard.hasura.io and create a free trial project.
"exactly what you're asking about"
"deprecated local feature"
Eh?
the following 3 have some potential:
https://feathersjs.com/and postgraphile https://postgrest.com/en/v4.3/ https://www.npmjs.com/package/postgraphile
Some cool differences:
- MIT/ZLIB/Apache2 licensed.
- Graph data structures (document, key value, tables, relations, and more!)
- Decentralized like IPFS!
- Alpha prototype for end-to-end encryption and user authentication / permissions security.
And a bunch more :).
It provides the same name document/collection concepts, but search abilities are still under development.
Side note, Im one of the cofounders :D
Generally this move by Firebase makes a lot of sense! Building large real-time apps usually requires smaller documents or else things get super odd at scale.
Edit: Minus rule based auth.
Do you think $0.18 per 100000 writes solves this problem?
If you're concerned about price per GB stored, Cloud Firestore will be much cheaper than Realtime Database (0.18/GiB/month). We evaluated many use cases when deciding on prices and we believe developers will be happy with the new model.
However since Cloud Firestore charges by operation, it's important to evaluate your use case when thinking about pricing. For example if you're running a fleet of IoT devices checking in a few times per second with very small payloads, you'd be doing a lot of write operations with very little storage and Cloud Firestore could be more expensive in that case.
firebase <-- presence, real-time editing, chat, in-memory things
firestore <-- things that have a save/submit button, transactional database
quick question: So the real-time features in firestore are not ideal for real-time text editors and chats. But saving documents from firebase after to user has left is perhaps a good middle ground.
Would firestore charge read for each reader that's listening in real-time when a document is updated?
In my experience both databases are totally appropriate for a chat app. Even though in Cloud Firestore you will pay a document write for each new chat message, that's only $1.80 for a million chat messages and you get all the rich querying from the Cloud Firestore API.
Regarding your pricing question: https://firebase.google.com/docs/firestore/pricing#operation...
> When you listen to the results of a query, you are charged for a read each time a document in the result set is added or updated. You are also charged for a read when a document is removed from the result set because the document has changed. (In constrast, when a document is deleted, you are not charged for a read.)
They have invested heavily over the lifetime of the company in avoiding providing support. That reputation can never be revived.
Google does not have direct customer support in its DNA, it has the opposite, whatever that is.
They have not demonstrated relentless commitment to being available to resolve issues, and nothing matters more than this if you've bet your company on their platform.
If something goes really wrong. Amazon has your back and you'll find someone who'll listen who has the power to resolve it. Google, you're stuffed. If you've built you're business around that thing that went wrong, well time for regret.
What's the underlying synchronisation mechanism?
Is it based on CRDTs?
How does this product differ from Gun.js and realm.io?
It uses Paxos as the consensus algorithm, along with internal systems like the TrueTime API to enable us synchronously replicate across multiple data centers.
NoSQL is cool, but ultimately GraphQL/Apollo serves many of the same issues but has the capability of a much richer, standardized and potentially lower cost backend.
did Kreatank do that logo?
graphql/rest -> openresty ( + custom code ) -> postgrest (custom) -> postgresql
yes he did :)
https://firebase.google.com/docs/firestore/query-data/querie...
It seems as though there is just the capability for AND logic between the WHERE queries, is that correct. Is the any plans to add AND/OR and perhaps a concept for representing grouping like parentheses?
And for the data structures:
https://firebase.google.com/docs/firestore/manage-data/struc...
Is there any capability to perform joins or have the data populated at query time?
And for the last thing:
It is mentioned multiple times in the documentation that nested collections will not be deleted if the parent is deleted. I'm just curious why there isn't the capability to insert a document with options that would allow something like a cascade delete. Since you allow indexes on collections, there is obviously some sort of metadata maintained, why not just add an additional flag that could be set so that whenever a record is deleted it can optionally have its subcollections removed as well?
The biggest difference is the integration with Firebase, so you have access to Mobile (iOS/Android) and Web SDKs along with a native offline mode. This comes along with the real-time synchronization feature that makes serverless app development a breeze.
Cloud Datastore is great for large scale server-side development where you manage your own connection to your app, such as running your own website on App Engine or via Compute/Container Engine.
> multi-region replicated database [..] once data is committed, it's durable [...]
> strongly consistent on the server-side
Do you mind elaborating a bit more here? Around perhaps what happens underneath the hood when failing over, etc. Do you have a single "co-ordinator" of sorts ensuring strict serializability, if so what do you do when failing this over? Or is it a quorum based approach like Paxos/Raft?
I'm curious, what's the react native support like for Firestore / firebase in general?
The Cloud Firestore integration is fresh out of the box, but I'd certainly recommend trying it out!
"The Blaze billing plan (pay-as-you-go) for Firebase is required for Google projects with billing enabled.
To use the free tier, you must first turn off billing in your Google project."
I don't want 15 "projects". I want to add this to my existing one. Why on earth is this not possible?
For this reason we are not going to invest in making a GeoFire library for Cloud Firestore and spend that effort getting the native functionality ready.