HNHacker News
TopNewBestAskShowJobs

quinthar

163 karma · joined October 29, 2008

submissionscomments
quinthar··on Expensify achieves extreme concurrency with NUMA balancing
I had some free time this holiday season and wrote up a blog post about how we use NUMA balancing on our 384-CPU, 6TB monster servers to keep the site humming reliably (and some of the challenges we've had to overcome). It's pretty low level, but I found there wasn't a lot of good info on how to optimize for high core-count servers, so I thought I'd share some experience. Enjoy!
quinthar··on SQLite – The “server-process-edition” branch
we have been using this page locking technology for quite a while now and it works incredibly well. Having both read and write concurrency is super powerful and scales fantastically.
quinthar··on SQLite – The “server-process-edition” branch
It's scales so perfectly we are deploying it across 3 datacenters, each with four 384 core machines with 3 terabytes of RAM. I can't speak highly enough about sqlite and the team behind it.
quinthar··on SQLite – The “server-process-edition” branch
Expensify is powered entirely using sqlite (inside bedrockdb.com) and it is freaking amazing.
quinthar··on Linus Torvalds on Wireguard
At Expensify we have been astonished by how easy it was to set up -- as simple as ssh -- and its incredible performance relative to OpenVPN. It's night and day due to it's multi threaded design, meaning it doesn't have the same single cpu bottleneck as OpenVPN. It's clearly the future.
quinthar··on Steven Pinker’s case for optimism
I think a big challenge (and perhaps the root problem) is weighing the relative importance of different relative changes. Car deaths are down 97%, but suicides are up 30%. On balance, are things getting better or worse? (I'd say that car deaths are by far more common so the reduction there is more significant, but suicides are probably more directly correlated to happiness that car accidents...)

I'm more persuaded by the optimistic perspective, personally, but these are great contrasting pieces to highlight that it's not a slam dunk.

quinthar··on Steven Pinker’s case for optimism
This is a much, much better rebuttal than the other one. Thanks!
quinthar··on Steven Pinker’s case for optimism
Hmm how is this a good rebuttal? I haven't read the book, but the original article is based primarily on clear assertions of fact and statistics. A "good rebuttal" world presumably show that those stats are wrong, or at best misleading. But John's rebuttal is just a bunch of ponderous, abstract philosophy, with the most compelling part (and it's not very compelling) being:

"Much of its more than 500 pages consists of figures aiming to show the progress that has been made under the aegis of Enlightenment ideals. Of course, these figures settle nothing. [Uh, of course? Care to elaborate?] Like Pinker’s celebrated assertion that the world is becoming ever more peaceful – the statistical basis of which has been demolished by Nassim Nicholas Taleb [Oh snap! It was demolished by this dude I've never heard of, using reasoning that isn't provided? Dammnnnn, sick burn] everything depends on what is included in them and how they are interpreted. [Alright, I'm ready for you truth bomb]

Are the millions incarcerated in the vast American prison system and the millions more who live under parole included in the calculus that says human freedom is increasing? [Um, are they not? Or are you asking me to do the homework to see if your point has merit?] If we are to congratulate ourselves on being less cruel to animals, how much weight should be given to the uncounted numbers that suffer in factory farming and hideous medical experiments – neither of which were practised on any comparable scale in the past? [Good question... that you decline to answer]"

Basically: "X is misleading, because what about Y?" I don't know man, can you just tell me rather than leaving me guessing? It would be more compelling if he did his own homework rather than leaving his sucker punch as an exercise to the reader.

I think the original article is far more compelling than this so called rebuttal.

quinthar··on SQLite Query Language: upsert
Fyi, Expensify uses sqlite as it's primary database in a HA clustered way, see bedrockdb.com for more details.
quinthar··on Self-Awareness for Introverts [pdf]
I think salary notifications in general are a sign of the company compensating in a fundamentally unfair way. Unless you are being hired to negotiate, paying you based on your negotiation skills seems unfair. It's harder -- a lot harder -- but I'd encourage the CEOs reading this to create a compensation scheme that is "fair by design" and pays correctly without the need to negotiate. Here's what we do, though I imagine there are likely even better ways available: https://blog.expensify.com/2016/06/17/expensifys-comp-review...
quinthar··on Consus: Fast Geo-Replicated Transactions
Cool! Also, you might take a look at http://bedrockdb.com for a geo-replicated, ACID, SQL database (basically a replication layer atop SQLite).
quinthar··on Single-file C/C++ public-domain/open source libraries with minimal dependencies
Well if the rules are:

- Libraries must be usable from C or C++, ideally both

- Libraries should be usable from more than one platform (ideally, all major desktops and/or all major mobile)

- Libraries should compile and work on both 32-bit and 64-bit platforms

- Libraries should use at most two files

Then ya, I think it'd fit. I'm not sure why size matters here: the goal is to have libraries that are easily embedded. Anyway, I was mainly curious for the reason. If size is the reason, then that's that. Thanks!

quinthar··on Single-file C/C++ public-domain/open source libraries with minimal dependencies
In the FAQ it just says "Come on" for SQLite included in the list. Can you provide a bit more detail? This feels like a pretty great option to include here.
quinthar··on Bedrock – Rock-solid distributed data
We're using GitHub Issues: https://github.com/Expensify/Bedrock/issues However, honestly those specific issues aren't on the list. In general we focus less on what could happen, and more on what actually does happen. Those specific issues haven't ever occurred, and thus never got "fixed" because they never became real problems. But PRs welcome!!
quinthar··on Bedrock – Rock-solid distributed data
And a bit more here: http://bedrockdb.com/synchronization.html
quinthar··on Bedrock – Rock-solid distributed data
Cool. It looks like they have 500-1000 Android installs. As context, Expensify has over 1M Android installs, and over 4M total users. So I think Bedrock is probably operating 4-5 orders of magnitude more activity, at least in this case.
quinthar··on Bedrock – Rock-solid distributed data
Ah, sorry for the confusion. Every transaction is given an incrementing ID by the leader, and every follower commits the transactions in ID order.

Furthermore, every commit has a running SHA hash of all prior commits (and every node keeps a history of the last few million commits). This way any two nodes can compare their journals to make sure they agree -- and if there is any split, then the cluster kicks that node out.

Basically, there is no scenario in which a node that commits a different transaction (or a transaction in a different order) is allowed to remain in the cluster.

quinthar··on Bedrock – Rock-solid distributed data
> What do you do during index creation: Do you drop write requests on the floor? Do you store them in a persistent queue somewhere? Or do you remove a node from the cluster, create the index, re-add the node to the cluster, and repeat this on every other nodes?

The last of those -- for small indexes, we just replicate them out like normal queries. For large indexes, we take that node down and add offline.

> So, basically, you're doing a rolling upgrade? Have you evaluated compiling your stored procedures to dynamically loaded libraries, or using an interpreted language like Lua/Python/JavaScript?

Correct, rolling upgrades. It's worked well to date, but the idea of putting the plugin into a dynamically loaded library is really interesting. Our plugin system is relatively new (only the past few months) so it hasn't really been considered. Great idea!

quinthar··on Bedrock – Rock-solid distributed data
To be clear, we're talking about functionality that is not implemented or fully designed. Today all transactions are committed on all nodes in the same order, which is a much simpler world. I agree, the multi-threaded replication case is a much more complex and interesting world, with much greater performance opportunities. Lots of exciting problems to solve when we get there!
quinthar··on Bedrock – Rock-solid distributed data
Correct!
quinthar··on Bedrock – Rock-solid distributed data
Lots of comments! I'll try to address the key points:

Re: "brand new technology not yet in trunk" -- I'm not sure what you mean by that. Bedrock has been used continuously for 8 years.

Re: "struggles over high-latency, low-reliability WAN connections" -- Ah, I mean MySQL's replication is designed for active/passive deployments with manual failover connected via fast/reliable networks. Not saying it can't support slow/unreliable WAN connections, only that it requires a lot of glue code and manual recovery when things go more wrong thn it's designed to handle.

Re: "the engineering required to correctly implement this distributed system is genuinely challenging" -- Agreed! This is precisely what Bedrock provides.

Re: "SQL fanout" -- To my knowledge, no RDBMS does this automatically. Furthermore, very few real world applications actually require this. Don't get me wrong: this is cool stuff. But this isn't a common requirement. Most businesses will never exceed the capacity of a single server -- the number of businesses that truly require this level of scalability is very small.

Re: "I just don't understand what niche Bedrock is trying to fill." -- The niche of businesses that want a simple SQL database that has built in automatic failover.

Re: "multiple logical SQL instances" -- I'm not sure what you mean by this. Bedrock maintains a single contiguous SQL database across all nodes.

Re: "you only get as much read throughput for as much data as you are willing to duplicate" -- Yes, I'm saying there are very few reasons not to duplicate it all for the vast majority of real world use cases.

Thanks!

quinthar··on Bedrock – Rock-solid distributed data
Unfortunately you can't add new nodes without reconfiguring the cluster, and currently that requires restarting each server (though so long as you don't do it all at once, server restarts are normal maintenance that causes no downtime to the end user). This could likely be added without too much effort, but also the real-world use case of this is uncertain.

The consensus algorithm is only used to elect a master -- once the master is identified, it coordinates all the distributed transactions. If the master dies, everyone who remains elects a new master and re-escalates all unprocessed write transactions to the new master.

Read transactions are processed locally by each node, so only writes need to be escalated.

quinthar··on Bedrock – Rock-solid distributed data
It's not called out specifically (actually, when writing it I didn't even know what Paxos was, and only realized I had implemented it years later). However, the logic is here: https://github.com/Expensify/Bedrock/blob/master/sqliteclust...
quinthar··on Bedrock – Rock-solid distributed data
We do most of the development in MacOS, but deploy it to Linux. I've never compiled it for Windows, but I imagine it would be straightforward.
quinthar··on Bedrock – Rock-solid distributed data
Every host has all data. (This sounds crazy, until you do the math on how cheap storage is and how comparatively "little" data people actually need in the real world.) So we have 6x redundancy in normal operation, across three datacenters (and three different power grids, three different providers). There are also nightly backups.
quinthar··on Bedrock – Rock-solid distributed data
Great questions! Yes, indexing is actually a significant challenge at our scale, as it can take a long time to add and currently, it is a blocking operation. However, the sqlite team is amazing and has all sorts of write-concurrency tricks up their sleeve to allow for creating indexes in a parallel thread. We haven't used it yet, but we plan to.

As for the stored procedures, yes they are written in C++, and thus deploying them requires compiling and upgrading the server itself. However, we have 6 of them (and honestly, any one of which has enough read capacity to satisfy about all our traffic) so a minute of downtime to upgrade for each independently isn't a problem.

The result is zero downtime as perceived by the user, even though each server has occasional downtime for maintenance and upgrades. In general I'd say we upgrade the database about weekly, or more frequent depending on how active we are in the stored procedures.

quinthar··on Bedrock – Rock-solid distributed data
Ya, I'm not sure I agree with that. I understand that argument in theory, but in practice the Jobs and Cache plugins have like 1% as much code as the Bedrock core -- and use all of it. So the idea that it makes sense to have a separate database, job queue, and cache -- despite all doing 99% of the same work -- strikes me as rather redundant and needlessly complex.

In practice, a job queue, a cache, and a database all do largely the same thing: maintain some internal state on disk, respond to some networked API calls, and synchronize across multiple servers. The actual logic of a job queue or cache is pretty trivial compared to what's underneath.

Indeed, I think the lesson of Redis is exactly why you should build these systems on top of a real database: they are struggling to hack on replication and reliable storage, even though these are "free" when built atop a database.

quinthar··on Bedrock – Rock-solid distributed data
Every write transaction ever committed to the database is assigned a unique ID, and the most recent few million transactions are kept in a "journal" table. This is what allows nodes to figure out what happened while they were offline and re-execute the missing transactions. A side-effect of that is we know exactly how many write transactions have ever been committed.
quinthar··on Bedrock – Rock-solid distributed data
To be clear, my sense is that every replica would receive every transaction, and would be free to commit the transactions inside each batch in any order.

Now I do agree that it's tricky to avoid "gaps" in the failure cases. However, every replica keeps a record of the past several million transactions (we aim for 3 days), and every transaction is assigned a unique ID. When a replica starts up, it "synchronizes" down every missing transaction it has, and at this point would "repair" any gaps it somehow obtained when it went down.

Admittedly, the exact details of that part are TBD, but it doesn't strike me as an unresolvable problem on the surface.

quinthar··on Bedrock – Rock-solid distributed data
Well, perhaps "new" is relative. But yes, I think strong application design moves the queries into stored procedures that execute inside the database, exposing a higher-level interface to the application layer. This not only maintains better layering from a software engineering perspective, but also provides better isolation from a security perspective:

Your webserver is the first thing to be hacked by an attacker because it's the server that sits on the internet. If all your security logic is built on your webserver, an attacker can easily bypass it all. However, if your security logic is built into your database, then it dramatically limits the damage a hacker can do from the webserver.

In particular, I would recommend you create an "Authenticate" stored procedure that accepts a username and password, returning an "authToken". Then make all your other functions into similar stored procedures, each of which accepts and verifies an "authToken" before returning the results.

(And in the case of Bedrock, you might also disable the Bedrock::DB plugin entirely to prevent direct access to your database from the webserver, forcing everybody to go through your stored procedures.)

This design means attackers who root your webserver can't randomly access your database without knowing the username/password (or valid authToken) -- all of which is neatly and securely contained inside the database itself.

To be clear, if you do keep your SQL on your webserver (as most applications honestly do, though I believe it's a mistake), then the Bedrock::DB plugin is perfect for you. But I think Bedrock::DB (or really, any direct access to a database outside of stored procedures) should be largely viewed as a programming and maintenance convenience, at the expense of security.

Page 1 of 3Next →