How to Design Programs
htdp.org
htdp.org
1: Start with the data structure
The tools to operate on the data can be rewritten or improved without much pain. Changing the data structure is what hurts. If the data structure is elegant, the code will fall into place. If the data structure is messy, the code will grow into a giant hairy ball of madness.
2: Don't use a DB. Use human-readable text files
It makes everything so much easier. You can see if your data structure is elegant. Because then the files will be easy to read and navigate. You will be able to start up your software 20 years from now without any installation and without having to go through the hell that is setting up the environment and updating dependencies. The website you are reading this on, Hacker News, stores all data in text files.
3: Build small tools to operate on the data
Like Linus started git by writing small tools "git-add", "git-commit", "git-log" etc. On the web you can have one url to do one thing (/urser, /item, /edit here on HN) and map them to one code file on your disk. If you want to use a framework like Django, Laravel or Express, those will unfortunately make it a bit harder to do that. But you can still stick somewhat closely to this approach.
4: Let people use it right away
If you are building it for yourself, you already have a user. Great. If not, let your target group use it right away. Otherwise, you will waste many years of your life building software features nobody understands or even wants.
This is interesting, can you elaborate? Do you mean like CSV files, or some JSON structure, etc?
I do use CSV and JSON for some types of data, yes.
Sometimes I use YAML. Sometimes just free text. Sometimes I roll my own format.
So: Hacker News, tiny team (one person?), very simple application, doesn't change. An accounting system: rather complex, frequently changing requirements, often developed by larger teams, needs to be accessible to non-developers, has data integrity requirements. You've really got no choice except SQL.
You should definitely consider every application's needs individually, but there's a line for everything where it makes more sense to use a DB, and not understanding where the line is will bring you a lot of pain. As for longevity: the SQLite3 format has been around since 2004 and is utterly ubiquitous. I think it's pretty safe.
Larger programs? Look at Linux and its "everything is a file" philosophy. Linux is one of the largest codebases out there.
SQLite will be around for a while, but not as long as text files. Who knows how long you can easily access present day sqlite files without fiddling with SQLite version numbers. And the format does not offer the easy access to the data that text files do.
I'm not saying nobody should ever use a DB. But files are often the superior approach.
Git is actively maintained by a number of people: https://github.com/git/git/graphs/contributors
Most commercial software is developed by teams.
> Larger programs? Look at Linux and its "everything is a file" philosophy. Linux is one of the largest codebases out there.
"Everything is a file" is an abstraction that Unix presents to the user. It's certainly not the case that that's how everything works underneath — the file system itself isn't a file, for example.
> I'm not saying nobody should ever use a DB. But files are often the superior approach.
And I'm not saying nobody should ever use text files. But a database is often the superior approach.
Otherwise I would counter "Even SQLite stores its data in files".
In one of my Racket applications with high integrity demands I've resorted to writing the temp file, swapping it with the original, then opening it read-only and verifying all of its contents. That is very reliable but obviously slow. But in another application written in Go I've switched entirely to Sqlite even for simple settings files and had no problems since.
When interacting with a typical web application, the user cannot terminate it in the middle of a file write application.
A related thing I discovered recently is the potential non-atomicity of writes using O_APPEND on Linux[1], although I couldn't get the attached test program to fail on my machine. I would love to find some kind of confirmation that this behaviour has been changed and that appends can be relied upon to be atomic.
In the application I'm working on at the moment I've managed to get add more certainty around this stuff since I'm only ever appending to the file in 4096 byte chunks, so can perform a simple file size check to see if data was written completely or if the file needs to truncated to the nearest multiple of 4096. I'm fortunate that it prevents having to do something like open it read only to verify contents as you mentioned.
Of course, it should be perfectly possible to get this right in languages like Go or Racket. At least Go has low-level enough interfaces to the filesystem API. My point was merely that using Sqlite turned out to be less error-prone in my experience for end-user applications, especially if you use WAL and set some cautious Pragma options. It's really good at dealing with unusual filesystem conditions.
I agree that reliably writing files is a sneakier problem than it first appears and that the likes of SQLite have already solved it.
Despite that my preference these days is still for a DB, like almost all dependencies, to have to justify its existence in my projects rather than defaulting to it as many seem to favour. It's straightforward to add SQLite but I'd typically rather take a little extra pain around the filesystem upfront to not have to deal with all the extra complexity of an SQL database if I can avoid it.
Naturally, this kind of simplicity is just one factor among many such as the size and shape of the data, expected access patterns and a host of others to weigh up during the trade-off decision.
One of the linked papers[1] in that article is highly recommended.
1. https://www.usenix.org/system/files/conference/osdi14/osdi14...
What about when the data grows a little bit and you need an index? Or you need to do a simple join? Implement your own BTree? Write your own join? Reimplenting a worse version of SQLite is not my definition of everything being easier.
I agree with the rest of the points btw.
This migration is easy to go from a flat file into a database tbh. And you can do this migration in pieces - copy the flat file and perform the ETL as a separate thing, all while your existing app runs fine unchanged. You can then implement the needed joins or indexing on this copy in the DB, and then schedule some sort of job to re-import the DB from the flat file.
Sure it's not real time, but it's enough to implement most features, with little effort from migrating the read/write code in your app from flat file to the DB. And nothing stops you from migrating the read/write api from flat files into the DB, after you find that it is really needed.
So, almost always.
Honestly, the software I wrote 10-20 years ago, if it doesn't run now it's not because I used a database system that's fallen out of fashion. If anything, systems built back then most likely would have leaned into storing everything into some cronenbergian XML-amoeba. It's unpleasant to work with, but it's not like it's somehow no longer viable.
I can see the appeal: text files are forever! read and write as simple strings!. But what do you do when you want to search ? whip up a regex or a for loop ? can you see how this will get unimaginably nasty with any remotely complex query, you will end up implementing a buggy, slow and informal subset of SQL without realizing it. Or, what if the data has constraints in it, like uniqueness of a particular column value across all rows, or other complex constraints, or types. Text files can't help you. Hell, if you aren't careful with charsets text files produced on 1 OS turn out to be garbage on another.
Maybe use text files as a conceptual\prototyping tool to imagine the persistent form of your data structures, or as a dump of the database for backup. That's reasonable.
It is also not too difficult to switch from plain text files to a database, if one does this step in time, before implementing ones own persistence, joins and whatever on basis of files.
Do you store in plain text or do you simply use CSVs and treat them like a database?
Do you find that the repository pattern helps mitigate the issue with long-term DB deprecation? You could backup your databases as CSV files, and then load those into any DB server later. Your repository layer should allow you to use whatever DB tooling is fancy at the time and the CSV backups would be around for longevity?
For even faster prototyping -- maybe better called experiments -- skip the text files and do everything in memory. Just use the built-in collection types of the programming language. In Python for example, that would be lists, dictionaries, tuples etc, then you can simulate database queries with set comprehension expressions.
https://www.cs.utah.edu/~mflatt/htdp-lite/
https://my.eng.utah.edu/~cs5510/htdp-videos.html
Here's another course which I have used to train interns. It's faster paced and useful for people who learn and progress at a faster pace.
https://web.cs.wpi.edu/~cs1102/a16/index.html
Here's another version of the same/similar where the assignments are challenging enough. You can do them in either racket or pyret.
How to Design Programs (2014) - https://news.ycombinator.com/item?id=26493990 - March 2021 (136 comments)
How to Design Programs, second edition - https://news.ycombinator.com/item?id=16561815 - March 2018 (30 comments)
How to Design Programs, Second Edition - https://news.ycombinator.com/item?id=14932552 - Aug 2017 (40 comments)
How to Design Programs - https://news.ycombinator.com/item?id=12768134 - Oct 2016 (3 comments)
How to Design Programs, Second Edition - https://news.ycombinator.com/item?id=8778569 - Dec 2014 (39 comments)
How to Design Programs, Second Edition - https://news.ycombinator.com/item?id=6150967 - Aug 2013 (26 comments)
How to Design Programs, Second Edition - https://news.ycombinator.com/item?id=2958108 - Sept 2011 (18 comments)
How to Design Programs: An Introduction to Computing and Programming - https://news.ycombinator.com/item?id=2049477 - Dec 2010 (17 comments)
How to Design Programs - https://news.ycombinator.com/item?id=637152 - June 2009 (4 comments)
I learned how to build a vocabulary for the problem domain first, and then use the vocabulary to solve the actual problem through this course. Norvig has great examples on bottom-up problem solving approach.
If I were to do it consistently for an hour a day 4ish days a week I think I could have finished it all in 2-3 months.
could you please explain how this book helped in tackling a new class of problems?
I like it so far, but was initially surprised that there's a systematic way of designing programs. I've just been winging it throughout my self-taught career.
[1] https://www.youtube.com/channel/UC7dEjIUwSxSNcW4PqNRQW8w/fea...
I think it is better structured than HtDP and the recipe pages in the wiki are priceless. Do they have a course website?
https://edge.edx.org/courses/course-v1:UBC+CPSC110+2021W2/77...
https://mitpress.mit.edu/books/structure-and-interpretation-...
Does the assertion still hold?
https://parentheticallyspeaking.org/articles/how-not-to-teac...
Scheme is not what makes SICP challenging. The sheer density of concepts is.
HtDP is a bit repetitive and gets a bit boring over time.
While some things overlap, I would recommend to first work through HtDP (or better through this online course https://www.youtube.com/channel/UC7dEjIUwSxSNcW4PqNRQW8w/abo...) and then work through SICP - especially if you try to learn programming by yourself.
Everyone leaving a CS education should work through SICP.
https://www.reddit.com/r/compsci/comments/vyxqqt/comment/ig5...