SQLite 2022 Recap
sqlite.org
sqlite.org
The JSON -> and ->> operators introduced in 3.38 are my personal favorites, as these really help in implementing my "ship early, then maybe iterate as you're starting to understand what you're doing" philosophy:
1. In the initial app.sqlite database, have a few tables, all with just a `TEXT` field containing serialized JSON, mapping 1:1 to in-app objects;
2. Once a semblance of a schema materializes, use -> and ->> to create `VIEW`s with actual field names and corresponding data types, and update in-app SELECT queries to use those. At this point, it's also safe to start communicating database details to other developers requiring read access to the app data;
3. As needed, convert those `VIEW`s to actual `TABLE`s, so INSERT/UPDATE queries can be converted as well, and developers-that-are-not-me can start updating app data.
The interesting part here is that step (3) is actually not required for, like, 60% of successful apps, and (of course) for 100% of failed apps, saving hundreds of hours of upfront schema/database development time. Basically, you get 'NoSQL/YOLO' benefits for initial development, while still being able to communicate actual database details once things get serious...
Then one of the devs left and I inherited his code, and started to realize exactly how much the rigid schema helps make an app more robust and more maintainable. I will never, ever take that shortcut again. The production outages/downtime caused by database issues that never should have been an issue in the first place (like a corner case row that missing data? Spreading rampant defensive nil checks out across every function that checks any field of a record since any at any time might be inconsistent or nil) contributed to killing the company by pissing off customers.
Elixir/Phoenix does help a lot though because I still treat most rows like JSON objects, but under the hood it's a normal rigid postgres schema. Best of both worlds IMHO.
I've been hacking together one-off python scripts to parse out bits of saved API responses, flatten the data to csv's, and load it to SQLite tables. Looks like I can skip all of this and go straight to querying the raw JSON text.
These days though it's _so_much_simpler_and_cleaner_ to just wrap the call in a decorator that caches the request to sqlite3 and only makes the call if the cache is stale.
I don't worry about parsing the results or doing any of the heavy lifting right away - just cache the json.
Sqlite is so good at querying those blobs and is so fast it's just not worth munging them. Nice.
And using something like datasette to prototype queries for some of the more complicated tree structures is a breeze.
It's good for single-page applications. Many datasets are relatively small -- pushing them to the client is reasonable. In exchange, you get zero-latency querying and can build very responsive UIs that can use SQL, versus writing REST APIs or GraphQL APIs.
Taken to an extreme, it permits publishing datasets that can be queried with no ongoing server-side expenses.
A wild example: Datasette is a Python service that makes SQLite databases queryable via the web. It turns out that since you can compile Python and SQLite to WASM, you can run Datasette entirely in the user's browser [1]. The startup time is brutal, because it's literally simulating the `pip install`, but a purpose-built SPA wouldn't have this problem.
[1]: https://lite.datasette.io/?url=https%3A%2F%2Fcongress-legisl...
Plenty of interesting databases fit into less than a MB even.
I've been using SQLite in WebAssembly for my Datasette Lite project - a Python server-side web app running entirely in the browser: https://simonwillison.net/2022/May/4/datasette-lite/ - here's an article showing how that can be useful: https://simonwillison.net/2022/Aug/21/scotrail/
It's also available in Observable notebooks, which is really handy. Here's a project I built on top of that: https://simonwillison.net/2022/Nov/20/tracking-mastodon/ - notebook here: https://observablehq.com/@simonw/mastodon-users-and-statuses...
The SQLite team have been working with browser developers to ensure the new API is sufficient to enable all this.
Honestly, and I keep going on about it, SQLite in the browser via WASM is the missing piece to make "offline first" PWAs a serious contender when deciding an architecture for an app.
2023 is going to be the year of SQLite in the browser.
https://webkit.org/blog/12257/the-file-system-access-api-wit...
It's "offline first", not "offline only".
At the moment, as I understand it, the OPFS virtual disk is completely isolated from the users disk.
This means you cannot just lightly query a 1GB file without first copying the 1GB from the users filesystem to the OPFS.
Any writes mean you must then copy the 1GB SQLite db file from OPFS to the users local filesystem too.
Is this correct?
Yeah, it's going to be confusing for users when they want to actually want to use one of these files outside of the application that created it. But it's better than nothing!
Adding those API's to read/write parts of the file is the next logical step. But it looks like it is limited to the isolated OPFS only.
I kinda wish sqlite has more functions though.
Or Rust - https://ricardoanderegg.com/posts/extending-sqlite-with-rust...
But I would prefer using something Sqlite officially maintains and not to maintain anything myself. I also don't prefer using libraries from random people unfortunately. Working at a fintech makes me paranoid...