We Don't Need No Stinkin' Databases
btorpey.github.io
btorpey.github.io
Better: "With unix's join command you can join files somewhat similar to database joins"
Working with text based datasets is a skill I don't see enough of even though it is quite prevalent. (I have spent more time than I've wanted using tools like Monarch to make text databases relational through an automated workflow)
The reality of development is working with an existing, or multiple datasets homogeneously. Cool use of the JOIN command, a title like "Join text files using the command line" probably would receive a positive response here.
It'd be sweet if everything was a nice and clean API data call with a bow on it, but part of being a developer is enabling that where it need to be.
If you read my blog, you'll see that I try to keep things light. (No cat pictures, though -- at least not yet ;-)
Having said that, the description in the header is "Data manipulation with plain text files". Maybe I'll see if I can find a way to show that on GitHub ...
Basic language containers (arrays, lists, etc.) tend to be much more convenient for lightweight data. No server to install configure and secure - everything can just live in code and the occasional serialized file. On the more heavyweight side of things, SQL isn't really amenable to GPUs - problematic for some heavyweight data crunching, or even basic computer graphics.
It starts to make sense in the context of a server farm, but leaving it publicly exposed to the internet is a security fail - so the first thing you do with this networked appliance is lock it down to only be accessible from your server farm, or even just your local server via loopback if you're feeling particularly paranoid - writing your own layer that exposes a severely restricted API that only allows the exact operations you want to permit against your database. A lot of people will get this far and then still contribute to the CVE statistics on SQLI vulnerabilities, despite parameterized queries solving this problem. Or store blobs and defeat most of the point of using SQL in the first place. Or denormalize and shard their database for scaling purposes, losing many of the benefits of SQL enforced invariants in the process.
It's not so much that I keep reinventing SQL on my projects, so much as I keep failing to come up with a good excuse to reinvent my existing approaches in SQL.
- Command-line tools can be faster than your Hadoop cluster [https://news.ycombinator.com/item?id=8908462]
- Going Deep [https://news.ycombinator.com/item?id=8902739]
It also reminds me of one of my favorite aphorisms: "Perfect is the enemy of good enough".
What some other commenters appear to have missed is the use case I'm discussing here: scraping log files to generate text files to feed to gnuplot.
And while the approach presented here may not work at Google scale, it works just fine for my situation. It takes me about 10 minutes on my MacBook to churn through around 50GB of logs to produce a few dozen charts. It took about a week to put the scripts together, and now anybody on my team with a shell prompt can do the same thing (including the QA folks).
Sometimes you need to dot all the i's and cross all the t's, but sometimes it's just about getting the job done quickly, with minimal dependencies.
Like figuring out a reasonable access path between tables when you're dealing with billions of rows
Like guaranteeing consistency between your relations. While the join command sounds useful (I didn't know it and I'll look into it) it's not really a replacement for all the constraints, which a database provides to ensure that you don't fubar your data
Backup, recovery? Just hope to you really grab everything. No only the data, but also your carefully crafted queries. Sounds like a bummer to figure them out anew
Implementing ACID also sounds like quite a challenge
A database is not just a shipping container, where you cram in your data and provides a few tools to organize it. There's a lot more under the hood, which is just not supported in plain old files.
If you think about it, a modern disk filesystem is really a form of database (but not the relational kind): various kinds of data (filename, date, access time, etc., and finally the actual data) are stored in a schema on the disk (which itself is abstracted by the disk controller) in a way that it can be looked up and retrieved quickly.
Most modern databases I've heard of use regular filesystems to store their data, so it's a database using another database to store data.
Wouldn't it be faster to eliminate the filesystem altogether, and just have the SQL database go directly to the disk, just like Linux swap partitions do?
After a little googling, it does seem that Oracle used to do this (but no more), and MySQL's InnoDB still offers this. Others claim that there's no advantage to this because the kernel handles dealing with raw devices better than the DB does, but I think it's more likely that the DBs just don't have enough development resources to try doing it better; filesystems are necessarily going to incur a certain amount of overhead, so it only makes sense that theoretically you could go directly to the raw device and cut out some of that overhead, though you would be replicating some of what the kernel's filesystem driver is doing, but in a way that's more optimal for the DB.
Hmmm...that might be a good topic for another article.
In the meantime, here are a few packages you can install using HomeBrew (brew install <package>) that help bridge the gap between Mac and Linux for other tools:
bash binutils coreutils gawk gnu-sed gnuplot
> Knuth wrote his program in WEB, a literate programming system of his own devising that used Pascal as its programming language. His program used a clever, purpose-built data structure for keeping track of the words and frequency counts; and the article interleaved with it presented the program lucidly. McIlroy’s review started with an appreciation of Knuth’s presentation and the literate programming technique in general. He discussed the cleverness of the data structure and Knuth’s implementation, pointed out a bug or two, and made suggestions as to how the article could be improved.
> And then he calmly and clearly eviscerated the very foundation of Knuth’s program.