HNHacker News
TopNewBestAskShowJobs

otoolep

5,064 karma · joined April 3, 2014

Software Engineer based in Pittsburgh, PA. Engineering Manager at Google, building large-scale data systems. Creator of rqlite[1], the lightweight, distributed database built on SQLite.

https://www.philipotoole.com

[1] https://www.rqlite.io

submissionscomments
otoolep··on How rqlite is tested
rqlite supports a few levels security, which you can use separately or together.

You can enable BasicAuth on the HTTP endpoints. You can also require that any client presenting BasicAuth credentials has the right permission for the operation the client is trying to perform. Finally you can also enable Mutual TLS, requiring that any client that connects must first present a cert that is signed by an acceptable CA.

For full details on security check out https://rqlite.io/docs/guides/security/

otoolep··on How rqlite is tested
It's not done by Aphyr himself, but someone ran Jepsen-like testing a couple of years back.

https://github.com/wildarch/jepsen.rqlite/blob/main/doc/blog...

otoolep··on How rqlite is tested
Yes, Replicated for example: https://www.replicated.com/blog/app-manager-with-rqlite

It's still in active use by them. https://www.textgroove.com/ is another company that uses it, though I know less about their use case.

otoolep··on How rqlite is tested
rqlite has been in development for about a decade too!

https://github.com/rqlite/rqlite/blob/master/CHANGELOG.md#10...

otoolep··on How rqlite is tested
Yes, that is right. So here's my disclaimer: I'm the author of rqlite.

I really admire the testing that the SQLite team does[1], and the way it allows them to stand behind the statements they make about quality. It's inspiring.

IMO there have only been two really big improvements to software development relative to when I started programming professionally 25 years ago: 1) code reviews becoming mainstream, and 2) unit testing. (Perhaps Gen AI will be the third). I believe extensive testing is the only reason that rqlite continues to be developed to this day. It's not just that it helps keep the quality high, it's a key design guide. If a new module cannot be unit tested during development, in a straightforward manner, it's a strong sign one's decomposition of the problem is wrong.

[1] https://www.sqlite.org/testing.html

otoolep··on How rqlite is tested
Thanks for flagging -- fixed.
otoolep··on Show HN: Embed an SQLite database in your PostgreSQL table
rqlite[1] creator here, happy to answer any questions about it.

[1] https://rqlite.io

otoolep··on Building Observability with ClickHouse
There is at least one basic factual error in this blog post, which makes me discount the whole thing.

"But if you will use it, keep in mind that [InfluxDB] uses Bolt as its data backend."

Simply not true. The author seems to have confused the storage that Raft consensus uses for metadata with that used for the time series data. InfluxDB has its own custom data storage layer for time series data, and has had so for many years. A simple glance at the InfluxDB docs would make this clear.

(I was once part of the core database team at InfluxDB and have edited my comment for clarity.)

otoolep··on Rearchitecting: Redis to SQLite
Yes, that is right, it would require a new service running on the host machines.

That said, I do think it depends on what you consider important, and what your experience has been in the past. I used to value simplicity above all, so reducing the number of moving pieces was important to my designs. For the purpose of this discussion let's count a service as a single moving piece.

But over time I've decided that I also value reliability. Operators don't necessarily want simplicity. What they want is reliability and ease-of-use. Simplicity sometimes helps you get there, but not always.

So, yes, rqlite means another service. But I put a lot of emphasis on reliability when it comes to rqlite, and ease-of-operation. Because often when folks want something "simple" what they really want is "something that just works, works really well, and which I don't have to think about". SQLite certainly meets that requirement, that is true.

otoolep··on Rearchitecting: Redis to SQLite
rqlite[1] could basically do this, if you use read-only nodes[2]. But it's not quite a drop-in replacement for SQLite at the write-side. But from point of view of a clients at the edge, they see a SQLite database being updated which they can directly read[3].

That said, it may not be practical to have hundreds of read-only nodes, but for moderate-size needs, should work fine.

Disclaimer: I'm the creator of rqlite.

[1] https://rqlite.io/

[2] https://rqlite.io/docs/clustering/read-only-nodes/

[3] https://rqlite.io/docs/guides/direct-access/

otoolep··on rqlite: A lightweight, user-friendly, distributed relational db built on SQLite
Not particularly anything to do with Go. A C compiler is only needed for the SQLite source code.

I originally provided these musl-based builds so I could provide rqlite Docker images based on Alpine[1]. But now the Docker release process simply builds rqlite from the source during the image-creation process[2].

[1] https://hub.docker.com/_/alpine

[2] https://github.com/rqlite/rqlite/blob/master/Dockerfile

otoolep··on rqlite: A lightweight, user-friendly, distributed relational db built on SQLite
rqlite creator here, happy to answer any questions.
otoolep··on Postgres as a Search Engine
I wrote about this too[1], and since rqlite[2] puts a HTTP API in front of SQLite, you've got a SQLite-backed search engine available over the network.

[1] https://www.philipotoole.com/building-a-highly-available-sea...

[2] https://rqlite.io

Disclaimer: I'm the creator of rqlite, and it's not the only piece of software to make SQLite available over the network.

otoolep··on Making database systems usable
Or, I should say, I don't add the feature until I can figure out how it can made be easy and intuitive to use. That's assuming the feature is even coherent with the existing feature set of the database.

Of course, it's easy for me to do this. I am not developing the database for commercial reasons, so can just say "no" to an idea if I want. That said, I've found that many ideas which didn't seem interesting to me when an end-user first proposed them become compelling once I think more about the operational pain (and it's almost always operational) they are experiencing.

Automatic backups to S3[1] was such a feature. I was sceptical -- "just run a script, call the backup API, and upload yourself" was my attitude. But the built-in support has become popular.

[1] https://www.philipotoole.com/adding-automatic-s3-backups-to-...

otoolep··on Making database systems usable
>They care less about impressive benchmarks or clever algorithms, and more about whether they can operate and use a database efficiently to query, update, analyze, and persist their data with minimal headache.

Hugely important, and I would add "backup-and-restore" to that list. At risk of sounding conceited, ease of use is a primary goal of rqlite[1] -- because in the real world databases must be operated[2]. I never add a feature if it's going to measurably decrease how easy it is to operate the database.

[1] https://www.rqlite.io

[2] https://docs.google.com/presentation/d/1Q8lQgCaODlecHa2hS-Oe...

Disclaimer: I'm the creator of rqlite.

otoolep··on Building a highly-available web service without a database
>I'm curious whether you've had concerns about write amplification though?

I mean, yes, the more disk IO rqlite has to make to more write performance will be affected. However the advantages of running with an on-disk SQLite database are worth it I believe. In addition rqlite supports storing the SQLite database file on a memory-backed filed system if users really want that[1]. That can help squeeze more write throughput out of rqlite.

>My understanding is that rqlite Raft entries are mostly SQL statements (is that right?).

That's right, rqlite does statement-based replication, though I'm currently looking into extending it so it also does changeset[2] replication where it makes sense.

[1] https://rqlite.io/docs/guides/performance/#use-a-memory-back...

[2] https://www.sqlite.org/sessionintro.html

otoolep··on Building a highly-available web service without a database
>You could just use sqlite with :memory: for the Raft FSM

That's the basic design that rqlite[1] had for its first ~7 years. :-) But rqlite moved to on-disk SQLite, since with WAL mode, and with 'PRAGMA synchronous=OFF' [2], it is about as fast as writing to RAM. Or at least close enough, and I avoid all the limitations that come with :memory: SQLite databases (max size of 2GB being one). I should have just used on-disk mode from the start, but only now know better.

(I'm guessing you may know some of this because rqlite uses the same Raft library [3] as Nomad.)

As for the upgrade issue you mention, yes, it's real. Do you find it in the field much with Nomad? I've managed to introduce new Raft Entry types very infrequently during rqlite's 10-years of development, only once did someone hit it in the field with rqlite. Of course, one way to deal with it is to release a version of one's software first that understands the new types but doesn't ever write the new types. And once that version is fully deployed, upgrade to the version that actually writes new types too. I've never bothered to do this in practise however, and it requires discipline on the part of the end-users too.

[1] https://www.rqlite.io

[2] This might sound dangerous but in the current design of rqlite, the underlying SQLite database is completely rebuilt from the Raft log on startup (which is fsync'ed on every write). So any corruption of the SQLite database due power loss, etc is moot since the SQLite database is not the authoritative store of data in rqlite.

[3] https://github.com/hashicorp/raft

otoolep··on Building a highly-available web service without a database
As for your second question, I don't think you'd benefit much from than that, for two reasons: - rqlite is a Raft based system, with quorum requirements. Running 2-node systems don't make much sense. [1] - Secondly, all writes go to the Raft leader (rqlite makes sure this happens transparently if you don't initially contact the Leader node [2]). A load balancer, in this case, isn't going to allow you to "spread load". What is load balancer is useful for when it comes to rqlite is making life simpler for clients -- they just hit the load balancer, and it will find some rqlite node to handle the request (redirecting to the Leader if needed).

[1] https://rqlite.io/docs/clustering/general-guidelines/#cluste...

[2] https://rqlite.io/docs/faq/#can-any-node-execute-a-write-req...

otoolep··on Building a highly-available web service without a database
It depends on what kind of transaction support you want. If your transactions need to span rqlite API requests then no, rqlite doesn't support that (due to the stateless nature of HTTP requests). That sort of thing could be developed, but it's substantial work. I have some design ideas, it may arrive in the future.

If you need to ensure that a given API request (which can contain multiple SQL statements) is atomically processed (all SQL statements succeed or none do) that is supported however [1]. That's why I think of rqlite as closer to the kind of use cases that etcd and Consul support, rather than something like Postgres -- though some people have replaced their use of Postgres with rqlite! [2]

[1] https://rqlite.io/docs/api/api/#transactions

[2] https://www.replicated.com/blog/app-manager-with-rqlite

otoolep··on Building a highly-available web service without a database
rqlite author here, happy to answer any questions.
otoolep··on Building rqlite 9.0: Cutting disk usage by half
Transactions -- or the lack thereof -- have nothing to do with the consistency guarantees offered by rqlite.

You may wish to read this:

https://github.com/wildarch/jepsen.rqlite/blob/main/doc/blog...

rqlite -- to the best of my knowledge and as a result of extensive testing -- offers strict linearizability due to its use of the Raft protocol. Each write request to rqlite is atomic because it's encapsulated in a single Raft log entry -- this is distinct from the other form of transactions offered by rqlite[1], but that second form of transaction functionality has zero effect on the guarantees offered by Raft and rqlite (they are completely different things, operating at different levels in the design). If you know otherwise I'd very much like to know precisely why and how.

[1] https://rqlite.io/docs/api/api/#transactions

otoolep··on Building rqlite 9.0: Cutting disk usage by half
Can you be more specific?

"Evidence Dump Fallacy." This fallacy occurs when a person claims that a certain proposition is true but, instead of providing clear and specific evidence to support the claim, directs the questioner to a large amount of information, asserting that the evidence is contained within.

otoolep··on Building rqlite 9.0: Cutting disk usage by half
Wow, a lot there. Thanks for your comments.

>One of the problems is if you're working with developers, the log replication contents is the queries, instead of the sqlite WAL like in dqlite.

I think you mean rqlite does "statement-based replication"? Yes, that is correct, it has its drawbacks, and is clearly called out in the docs[1].

>Another issue is if I want to architect a system around rqlite, it wont be "consistent" with rqlite alone. The client must operate the transaction and get feedback from the system, which you can not do with an HTTP API the way you've implemented it.

I don't understand this statement. rqlite docs are quite clear about the types of transactions it supports. It doesn't support traditional transactions because of the nature of the HTTP API (though that could be addressed).

>Furthermore to this point, you can't even design a consistent system around rqlite alone because you can't use it as a locking service. If I want locks, I end up deploying etcd, consul, or zookeeper anyways.

rqlite is not about allowing developers build consistent systems on top of it. That's not its use case. It's highly-available, fault-tolerant store, the aims for ease-of-use and ease-of-operation -- and aims to do what it does do very well.

>If I had to choose a distributed database with schema support right now for a small scale operation, it would probably be yugabyte or cockroachdb. They're simply better at doing what rqlite is trying to do.

https://rqlite.io/docs/faq/#why-would-i-use-this-versus-some...

Of course, you should always pick the database that meets your needs.

>If building a reliable database was as easy as integrating sqlite with a raft library, I would have shipped nearly 10 years ago.

Who said it was easy? It's taken almost 10 years of programming to get to the level of maturity it's at today.

>They need a more robust design and better safety guarantees than rqlite can offer today.

That is an assertion without any evidence. What are the safety issues with rqlite within the context of its design goals and scope? I would very much like to know so I can address them. Quality is very important to me.

[1] https://rqlite.io/docs/api/non-deterministic/

otoolep··on Building rqlite 9.0: Cutting disk usage by half
rqlite creator here.

That's clearly a mistaken attitude because both Consul and etcd also use a single "Raft group" and they are production-grade software.

Ruling out a piece of software simply because it doesn't "scale horizontally" (and only writes don't scale horizontally in practice) is a naive attitude.

otoolep··on Building rqlite 9.0: Cutting disk usage by half
Ah yes. If you are accessing the SQLite files directly I cannot be sure what you will see during the snapshotting. I haven't tested that, since accessing the SQLite files underneath rqlite is not officially supported.
otoolep··on Building rqlite 9.0: Cutting disk usage by half
Thanks for the question, but I don't follow it -- I don't see any data staleness if you query rqlite during snapshotting. Granted the blog post doesn't go into every single detail, so this might be hard to follow.

Can you expand a bit more on your concern? What scenario do you have in mind?

otoolep··on Building rqlite 9.0: Cutting disk usage by half
Agreed, in the sense that while rqlite has a lot in common with etcd (and Consul too -- Consul and rqlite share the same Raft implementation[1]) rqlite's primary use case is not about making it easy to build other distributed systems on top of it.

[1] https://github.com/hashicorp/raft

otoolep··on Building rqlite 9.0: Cutting disk usage by half
Just to be clear, rqlite is not a library. It's a complete RDBMS. rqlite has everything you need to read and write data, and backup, maintain, and monitor the database itself. It's not just a library (unlike, say, dqlite).
otoolep··on Building rqlite 9.0: Cutting disk usage by half
I usually bump the major version number anytime I introduce important new functionality, major performance improvements, or a major new design change. While the API hasn't changed in years, the underlying implementation and file layout can change a lot between major versions. I want to communicate that.

Also rqlite doesn't support seamless downgrades between major versions, only seamless upgrades. I want to communicate that too (I've put a lot of work into the backup-and-restore system[1] so users can protect themselves if they are concerned about the seamless upgrade failing on them).

So by bumping the major version it helps people understand that they are upgrading to a substantially different version of rqlite, even if their client code doesn't have to change at all.

[1] https://rqlite.io/docs/guides/backup/

otoolep··on Building rqlite 9.0: Cutting disk usage by half
rqlite creator here. That's a mischaracterization.

rqlite has been in development for 10 years[1], it's a long-running project and its design goals have never changed. The API hasn't changed in a breaking fashion since 2016 and rqlite has supported seamless upgrades for years now.

In other words rqlite users have been upgrading from version to version for over 8 years, without having to change a single line of their code.

[1] https://rqlite.io/docs/design/

← PreviousPage 2 of 11Next →