Convex vs. Firebase
docs.convex.dev
docs.convex.dev
The CTO worked with Turing-award winner Barbara Liskov on Viewstamped Replication Revisited [1], the revision of the pioneering consensus protocol, and later the founding team pulled off the migration of Dropbox from S3 to their own custom storage stack called Magic Pocket [2].
The deterministic simulation testing techniques [3] they used to develop Dropbox's revised sync algorithm are still state of the art today (few systems are designed to be tested like this), and the online verification techniques [4] they used to verify their production systems are vital to building large-scale systems safely.
I can't think of a stronger technical team to do something like Convex, and couldn't imagine a better devops team to be running the backend.
[1] https://pmg.csail.mit.edu/papers/vr-revisited.pdf
[2] https://www.wired.com/2016/03/epic-story-dropboxs-exodus-ama...
[3] https://dropbox.tech/infrastructure/rewriting-the-heart-of-o...
[4] https://www.oreilly.com/library/view/velocity-conference-new...
Hi Shawn. I don't remember you bringing this up when we spoke in person recently, but it's a great question.
In our opinion, the very best way to evaluate yourself as a backend developer is how directly you solve problems for frontend developers. We believe in the merit of customer obsession, and the customers are not buying queues. They're buying the product as they see it: its surfaces, workflows and experience. And that's what the frontend developers, PMs, and designers are creating.
Historically, all these backend technologies that only interoperate with each other are only useful so long as they make product creation and improvement easier, more reliable, etc. We strongly believe as soon as you don't need them anymore, you should toss them out. They're complex and not proprietary to your product.
Convex (and serverless in general) is just the next step in providing more powerful abstractions that allow companies to double down on frontend engineering (work that adds product value) instead of reimplementing the same backend/devops plumbing the users never see (work that, at best, merely sustain product value).
So, given that we recognize this need, I respectfully disagree that we're not well equipped to solve these problems for frontend developers! Most of our team's recent our work has been designing synchronization and storage platforms to enable product development, including work on web, desktop, and mobile libraries/SDKs. We feel like we have both a lot of empathy and experience for this space, and we're very proud of our early product and the enthusiasm from the web dev community.
great answer :) if you can solve DX for frontend you're solving it for everyone.
I was the one of the original TLs of Google Photos, specifically on Android and eventually managed all of the frontend teams for Google Photos. So I hope it's not a stretch to say that I have deep empathy for frontend problems. We also have great frontend oriented folks on the team.
I also worked with jamwt at a previous company (Bump) doing frontend work. When we worked together, I remember the acceleration of our ability to execute when the backend people were closely involved in our data syncing problems. Heck jamwt and team even wrote most of the initial client syncing code. When backend and frontend folks work closely together to solve data problems you can make magical things happen. In the end we're just product oriented engineers trying to ship delightful experiences.
const querySnapshot = await getDocs(collection(db, "messages"));
const userSnapshots = await Promise.all(
querySnapshot.docs().map(async messageSnapshot => {
return await getDoc(docSnapshot.data().creator);
})
);
Phew!Thanks, but no. Never. I will keep doing it server side:
$messages = DB::select(
'SELECT * FROM messages JOIN users ON users.id=messages.user_id'
);
It is amazing with how much cruft developers are willing to deal with these days. And how much CPU cycles get burned for nothing, as the Firebase example fires one query per message to get the user. This would be bad enough on the server. But with the Firebase example, it would also create a client-server http roundtrip for each message. Mind-boggling.1) You are severely limited in how you can query. The list of limitations is too long to recount here, but querying is nearly worthless.
2) The database design strongly pushes you towards nesting collections, making side effects cleanup a disaster, especially as a database grows in complexity.
3) You cannot sort on a field without creating an index first. I get why creating an index is a good idea, but I can't even write a simple analytics script without indexing the fields first.
4) It gears itself towards frontend developers who don't know how to write a backend, and encourages bad practices for them. One example: Firebase lets you manually edit production data very easily from within their dashboard. Like it's treated almost like it's a CSV.
5) The recommended development approach is to directly query the DB from the frontend, with no server in between. This means any data security has to be implemented in a separate DB rules document. The syntax and structure of this doc is super limited and often results in a giant, unmaintainable file.
6) Firestore cannot count. As in, you literally cannot query the number of records in a collection. If you want that value, you have to store it as a separate field in the collection, and then update that value each time you add and remove a doc. MADNESS.
I could go on, and on, and on.
If the goal is to create a prototyping DB, ideally there’d be some nice off ramping or migration tools for when the app needs to become production ready.
If you use it as a system component and not a do everything backend it is the most production ready clientside technologies on the planet.
[1] https://docs.confluent.io/cloud/current/clusters/cluster-typ...
This is truly spot on. I'm one of those devs who couldn't write a good backend when I first began coding out apps (as a hobby). Started out with firebase because that was the tool used by most YouTube videos and Medium blog posts introducing people to app development. In the end, I eventually had to learn other technologies such as using Elasticsearch, DynamoDB (which I would also stay the hell away from) and PostGRES.
Thankfully Supabase decided to stay away from the NoSQL format, so learning how to get started with PostGRES was made smoother.
Looking back I don't even know why I jumped into using Firebase, since I do have a pretty good footing in SQL querying. I don't know why Google isn't bothering with an Elasticsearch-like NoSQL solution for Firebase. And quite frankly, I would only use Firebase for Auth and RTDB for basic database stuff. If I had known this earlier, I might have saved months of learning Firebase Cloud Firestore crap online.
that is probably a business decision. Google has api limites, so figures.
It does provide a lot of serverless scalability for what it offers, but it's the classic case of optimizing for a situation that won't happen for 99% of their apps.
I want records where record.score >= 10 and record.date <= 2022-01-01
The first form (single field) is a simple range query on a single index - that's what Firestore is optimized for. The second form (with different fields) potentially requires walking the near-entirety of both indexes looking for matches, and therefore has unbounded time and computational requirements.
Firestore is designed so that you can't do things that don't scale. Sometimes that sucks, especially when you know that the data volume for that query will always be "reasonable". But the limitations are not arbitrary.
* Firestore is truly a "fire and forget" datatabase that scales without effort or maintenance. If your app works for 500 users, it will work for 500 million. Without a devops staff.
Yes, Firestore (aka Cloud Datastore) feels crippled compared to running aggregations and joins on an RDBMS. If your data and load fit on a single node Postgres, by all means use it, that's a great solution! When your requirements exceed that, you're in a different world. You can look at Spanner ($$$) or its clones (operational load, maturity). Or you can do what I did, and run the firestore/datastore as a master database and replicate data to other stores (eg BigQuery) for analytics.
Firestore's sweet spots are very small (eg, you have many microservices and want simple cheap zero-maintenance persistence) or very large (where scaling and availability would give you headaches anyway). In the middle, traditional RDBMSes are great.
another area where it truly shines is zero downtime for schema updates.
#3 is not really correct. You can sort on any single field but if you have multiple fields, then yes, you must create an index.
#4 I partially agree. The web dashboard makes things I don't want to do (accidentally edit or delete a field/document) dangerously easy, and things I do want to do (copy the contents of a document, save the results of a query, copy the text of a field) exceedingly difficult. The truth is that firestore is geared toward people who want an easy way to get near real-time data synchronization. It really sacrifices almost everything else.
The number one most annoying thing to do with firestore is work with their security rules.
1) That’s a JOIN and Firestore is not unique among NoSQL databases for being bad at it.
2) That’s the code you’d use in a web browser to load Firestore data. That code also handles all of the API and Authentication code you’d normally put in front of another DB. It’s serverless. So that is a fully functional snippet, your SQL example needs an API layer and frontend deserialization code.
3) It does not make a new HTTP connection per request. It uses a long-lived gRPC connection.
3: It is still a "a client-server http roundtrip" for every query. It does not open and close the http connection every time. But it sends a "GET / HTTP ..." request for every query. With hostname, accepted formats, encodings etc.
The entire point of the Fire* style of database is that these trade-offs are not worth it in the long run, that few development teams have the skill and time to implement this themselves, and that databases can and should solve this for you.
I have little love for Cloud Firestore, it's a trash fire riddled with poor decisions, but if you don't even understand what the problem is with that SQL query, you don't understand the expectations users have nowadays of front-end applications.
> "that few development teams have the skill"
This is the fundamental problem. There's no magic answer to a lack of skill.
You can do the same thing on a simple by selecting by a `chatid` in a table that'll get the latest messages/inserts. Again, the only thing needed is the pub/sub layer, not an entirely different database.
It would be a bit more time consuming and challenging to set this up across iOS, android, and web. A lot more time would be spent figuring out the system to make this work well and without issues at scale. On the other hand… It takes like 20 min to set up a real time feature with firestore that works across all platforms. It scales well and works great for what it is.
If people try to use it in place of a rdbms or complex use cases, they’re going to have a bad time. But for those who can utilize their tool belt effectively, it’s awesome.
I love firebase products.
const userSnapshots = await Promise.all(
getDocs(collection(db, "messages")).then(
docs => docs().map(snapshot => getDoc(snapshot).data().creator)
)
)
Ultimately you're kinda just comparing preference for SQL syntax over JS in that case.The JS example is also async (always adds visual complexity to code calls), whereas the PHP example is blocking. That SQL query doesn't look particularly efficient either...
There's plenty of good solid criticisms of Firebase & NoSQL but these aren't it. e.g. your HTTP roundtrip argument has been debunked by sibling commenters, but client-side DB calls is still riddled with adjacent problems & ultimately adding more complexity to the server-side avoid the need for client-side DB access is often worthwhile.
The fact that Firebase isn't using SQL is obvious and comes with up and downsides.
> In Cloud Firestore, the data on the client are loaded from the database at different points in time. Even if you listen for realtime updates, results from separate queries will not remain in sync. This creates consistency anomalies and bugs in your app.
Here is a link to the protocol documentation that the clients use to support it: https://github.com/googleapis/googleapis/blob/d0b394f188e8c3...
I'd link to the client implementation but it's quite involved.
There are ways to opt of of this and get stale cached data if you're offline, but you're explicitly opting in at that point
I have a lengthy list of complaints about both firestore and the firabase flavor of cloud functions, however, I will say that the ease of getting started with the firebase suite is unmatched, in my experience. Compared to any product on AWS or raw GCP, it feels like an actual product with people thinking about their users.
There is also a large community around the firebase products, the main example being Invertase.io which provides amazing open source native clients for firebase.
Regarding Convex specifically, the approach of writing queries server side seems great. The docs aren't clear (to me) about whether the queries only send incremental state changes. I would assume and hope that is the case.
In this bit[1] it seems like the function needs to execute the entire query again, which could become a significant performance issue.
[1] https://docs.convex.dev/understanding/convex-fundamentals/fu...
This is by no means a panacea, but it does result in different performance characteristics than one might expect.
> Later on, if any mutation inserts, updates, or deletes a record that overlaps with the read set, Convex knows it needs to recompute the listMessages query. If the result of listMessages changes, the new value is synced to the client and the component rerenders.
FWIW Firestore is able to send incremental updates and not just rerun the query when data changes. There is a complex system for broadcasting individual documents as the commit happens. I posted a bit about this a long time ago: https://news.ycombinator.com/item?id=26910411
Another key distinction is that "query" here doesn't refer to a database read, it could be a complex function containing multiple reads, relatively-arbitrary compute, etc.
I talk more about how we perform this comparison in this (now rather outdated) talk: https://youtu.be/iizcidmSwJ4?t=1218
Note that there are false-positives here, i.e., it's possible for the inputs to a query function to change without the outputs actually changing, but this is not a significant issue in practice.
as james mentions, we entirely re-run the javascript function whenever we detect any of its inputs change. incrementality at this layer would be very difficult, since we're dealing with a general purpose programming language. also, since we fully sandbox and determinize these javascript "queries," the majority of the cost is in accessing the database.
eventually, I'd like to explore "reverse query execution" on the boundary between javascript and the underlying data using an approach like differential dataflow [1]. the materialize folks [2] have made a lot of progress applying it for OLAP and readyset [3] is using similar techniques for OLTP.
I wish y'all the best I do think this is a super cool product!
The idea that n queries instead of a join is slow is not as true as you would think. Firestore supports streaming and pipelines at its core, and can reuse cache across operations. At the end of the day, the data goes over a narrow network channel. If you can saturate the channel, and don't leave any gaps, what's the performance difference if the data comes from a single query or many that are back-to-back. The data is transferred to the client either way. Both Firebase databases are pipelined, so this "many round trip" argument is not a decent argument if the client can issue the queries without waiting for responses (such as the code in this article).
The other is consistency levels and correctness. I constantly see devs call Firebase an eventually consistent database which is wrong, its causally consistent [1], and this makes a huge difference when trying to do OLTP. The offline capabilities are built on the consistency primitives, and it's the only way it can work. So while this convex article is banging on about "End-to-End Correctness Philosophy", they miss the most important quality of correctness, and if they are not careful, will miss the required engineering, and then be unable to deliver an offline cache over real-time streams. I see this playing out with Supabase, I warned them personally before they got into YCombinator that what they were building was not causally consistent. Since then, they have had to rearchitect their real-time features after shipping them. (I have not reviewed their latest design yet so I have no idea whether they have it right yet).
Many things sucked about Firebase. The bespoke security rules and the lack of views. So Convex is on the money shipping functions on the backend. I think Supabase is shipping competitors' mistakes with row-level security language. Personally, I think Firebase's mistakes can be fixed with the addition of an open-source Firebase server [1], as the clients are already open source and the mistakes are all to do with just the server. The real tech was always in the clients anyway (offline cache, connection management, operation queues).
It will be interesting to see if building expressly for React is a good idea. Firebase shipped many adapters, like https://github.com/FirebaseExtended/reactfire, using the "thin-waist" principle of not over-fitting. But Javascript technology moved from callbacks to async while Firebase was in the field, so the current API is not now idiomatic. But convex is setting itself for even more ecosystem fragility, what if React changes API or falls out of favor? This is a big risk! I hope they can roll with whatever happens!
The specific concern with regards to multiple round trips is scenarios where the client needs to send multiple serial requests to render a web view, e.g., having to fetch a list of posts first before knowing which post ids to fetch comments for. These request waterfalls can lead to high page load times for many web apps.
Consistency-wise Convex supports full serializability. The main point this article is making with regards to consistency however is providing tools to ensure not only that the database is consistent but that the data rendered to the end user is also a consistent view. I talk a little more about that here: https://youtu.be/B9aeddqwVas?t=446
Btw we are huge fans of Firebase and think it did a ton to advance the industry forward. Thanks for your work on it!
> Consistency-wise Convex supports full serializability.
As discussed elsewhere both Firebase DBs too and also provide clients with a causally consistent snapshots too, so you are not actually adding anything over the Firebase offering. Firebase has optimistic updates too and can use clientside persistent storage (I think convex has just an in-memory cache at the moment).
> e.g., having to fetch a list of posts first before knowing which post ids to fetch comments for.
Yes this does perform badly. Though this type of join is also an IO hog on a relational DB too, just it's not so visible.
RethinkDB was such a great project.
On RethinkDB: The post-mortem of RethinkDB is a good read https://www.defmacro.org/2017/01/18/why-rethinkdb-failed.htm...
If anyone has questions regarding Thin I'm happy to answer.
There's a lot of comparisons in here about querying in here those are fair-ish / true-ish, I do think that's also missing what Firebase is "about", at least that's not what I've found it useful for.
If you're thinking about weighty sever side business logic and SQL queries ... yeah Firebase isn't built around SQL, it's not built to optimize big business logic and SQL like queries (granted, there are still ways to do these things).
Firebase isn't likely replacement for your enterprise CRUD app that carries with it a ton of logic for each and every possible CRUD action. And you do have to stop and think about how you are going to structure your collections, docs and so on.
Having said that I think NOT being the solution for complex CRUD apps is what makes Firebase pretty great.
I think the idea that business logic on the server being kinda wonky on Firebase is true generally. At the same time:
"Once again, with Cloud Firestore you could put this code into Cloud Functions and use them as a business logic layer, but it's messy and requires giving up some of the platform's other features like reactivity and optimistic updates."
I think that statement is deceptively broad. Just using cloud functions doesn't eliminate optimistic updates and so on, the effect is limited to when or how you use them...
Except for maybe in the sense that convex is built on functions from day one so the docs and platform will lead you there by default rather than loading the documents which is non-optimal.
But that's a different argument than:
> With Cloud Firestore, the client interacts with its data by loading documents straight from the database.
Part of the difficulty of writing about Firebase is that it's actually a whole collection of tools with a host of complex tradeoffs. But that's also a lot of the difficulty in using it as well! Understanding when to use Cloud Firestore vs Realtime Database and when to load data directly vs via Cloud Functions are all complex questions with unclear answers.
Firebase Functions are asynchronous, and can be reordered relative to DB operations. They would be much more powerful if synchronous and on the wrtie path (so you could put auth logic in them instead of being forced to use the bespoke security language).
Firebase functions are at-least-once though, so they are pretty good but could be better.
I use Flutter for my front end, and I simply haven't had any of the trouble they say people are having with Firebase and React. Perhaps the problem is not the Firebase half of that combo.
I do use Typescript for Cloud Functions in Firebase and it IS slower than manipulating documents directly. I'd love having a more responsive mutation path with server-side business logic. But I'd love it more if I didn't have to write it in JS/TS.
Seems that Google has been flooded with Auth0 tutorials, but I'm wondering what alternatives mobile devs would recommend.