I feel like I get a lot more out of messing with / hacking on code than I do from reading it. I'm sure people vary, but I've got loads more out of open source contributions to sometimes small projects, and not very much out of trying to do something like read the code for the Glasgow Haskel Compiler or something.
So for code to read, I think having an in is crucial, at least for me. So I'd say, find a cool, maybe small open source project, look at the issue tracker if it has one, and try to implement something. You'll only really know if it's good code (or why) after you start trying to change it.
Either way, have a goal in mind when reading code. If nothing else, take notes on what you learn.
I typically find the recommendations people make like "here's an example of particularly good code for this particular language / style / architecture" to be more enlightening.
Nowadays if I want to learn some new thing I'll find the 'right' code to read. Someone who has done somthing similar to what I am trying to do. But when I was younger and still trying to understand more basic concepts I just wanted to read 'good' code. I wanted to know how to structure programming logic.
Now I don't want to know how to structure logic unless I'm really interested in learning some new fangled concept. Typically something having to do with concurrency since there always seem to be new ways to express concurrency in programming.
Most of the code I look to read these days is because I want to start using a new library and I want to see how someone else has done it.
I like this and it is true. In addition, I have found that good code also tends to survive refractors despite having the quality of being easy to change.
The Ask HN post specifically mentioned good code but what follows are some (of my subjective) thoughts about the benefits of reading bad code. Bad code is haphazard and varied and good code is “samey”. You are a pattern matcher and this is part of your training. Good code will make more of an impact on your understanding if you have read a lot of bad code beforehand. Being able to read and understand bad code is a far more lucrative skill to have in the workplace than being able to quickly understand good code. In a big team you will spend a lot (most if you’re senior) of time reading other people’s code. Best to get good at it.
I have yet to be presented with a steaming mess of code to read and explain to an interviewer; yet it is the very first thing most new joiners face at any seniority level.
https://cs.opensource.google/go/go/+/refs/tags/go1.21.0:src/...
https://cs.opensource.google/go/go/+/refs/tags/go1.21.0:src/...
The difference in a well structured codebase is that some of the code prevents you from having to read huge amounts of other code. All code is bad, it starts out bad just by existing, it's only redeeming quality is preventing you from having to deal with more bad code.
Everyone thinks they write good "clean" code, and it's never true. Good programmers are good because of the architecture of their code, not because a single excerpt of code in isolation looks a certain way.
What you really want to read about are good designs. Read APIs, models, concepts, schemas, etc.
Another comment mentioned the Go standard library, and I totally agree. But stop at the APIs, if you look inside, you'll see that it's also mostly garbage. It's good because the APIs are good, and you don't have to read the rest.
I'm not sure what you mean by "understand good architecture." One thing that makes an architecture good is its simplicity and clarity. If it's hard to understand, it can't be that good. I can say for certain: if you understand the problem being solved but not the architecture, then it is for sure a bad architecture.
(Unfortunately many people are chasing the Holy Grail of "functional programming" and never finding it because "functional programming" is a pale shadow of what's possible when you understand how compilers work: this is how Common LISP and scheme are so much more profound then, say, Haskell)
It's a little out of date but I was lately thinking up about the Scott Adams adventure games of the early 1980s that were written with a specialized interpreter which could be implemented in BASIC but was also implemented in assembly language for better performance. See
https://ahopeful.wordpress.com/2020/09/13/digging-up-adventu...
If you tried to implement a game like that directly in a language like BASIC you would be driven nuts because that kind of game is fundamentally "object oriented" in that there are a number of things like rooms and items that are all mostly the same except they are different in some ways and trying to code that with IF, THEN and ELSE is bad enough even before GOTO gets added to the mix. The thing is that GOTO becomes quite benign and even useful when it is used to implement interpreters.
So back in the day you would study systems like that to stretch your skills, today I would look at compilers and related technology, like the Jena rules engine.
Huh? What does this even mean? Both languages are very different from each other and the only thing they really have in common is having functions — how is one more “profound” than the other or any less worthy of being “functional programming”, for whatever that means to you?
As for something profound I'd point you to
https://github.com/cerner/clara-rules
which many people will struggle with because like many other production rules engines in LISP (and many other examples of simple compilers), there is hardly any code! Contrast that to the orders of magnitude larger rules engine Drools
https://github.com/kiegroup/drools
which is so crazy-complicated primarily because the Drools language is Java-based so you need all sorts of things that Clara or CLIPS don't need. Note that both of these systems use variations of this algorithm
https://en.wikipedia.org/wiki/Rete_algorithm
where a set of rules can be compiled to a network of transitions that can happen when facts are added to the knowledge base so it is very much an example of compiler technology. (That said, Drools supports a lot of features that Clara doesn't and also supports a very advanced RETE-like algorithm that can exploit parallelism that early versions couldn't.)
Safety critical code seems to still be usually in C/C++, for reasons having nothing to do with the language being safe. Why are these people not studying Erlang, MISRA C, Rust, some kind of formal proof solver language, or OpenCL?
What do people aspire to one day work on that would make the Haskell or LISP more appealing than a more locked down, static checking, kind of language?
Programming culture is more confusing than the code itself!
functional >> OO
and "OO sux" and all that. I think most of these people hadn't worked in the industry enough to have a clear picture of what better productivity and lower defects would look like but instead they were looking for a movement to join.It is interesting how functional programming has to consider data organization completely differently, but sometimes it is just that, a tool or methodology, not a philosophy to subscribe to. Ideally, the good (or future good) engineers realize this and try to understand the less talked about soft skills needed to succeed.
The stuff I was really excited about was mostly "tech in the physical world" related.
Although with the AI boom, and how I'm too clumsy for jobs involving driving... I might be making a lot more if I'd at least tried to focus on the pure programming stuff instead of embedded, even though I never had any real talent for advanced algorithms...
from drools website
"Drools is a Business Rules Management System (BRMS) solution. It provides a core Business Rules Engine (BRE), a web authoring and rules management application (Drools Workbench), full runtime support for Decision Model and Notation (DMN) models at Conformance level 3 and an Eclipse IDE plugin for core development."
No expert here, but it doesn't come across as a fair apples-to-apples comparison.
Drools even has some projects like OptaPlanner and jBPM that are in different source code repositories.
The core in Drools really is dramatically more complex than the core of Clara and some of that really is the higher level of functionality (Drools supports both hash-based and ordered indexes for data because you need that for complex event processing, Clara has only hash indexes) but a lot of it is that the syntax of Drools is much more complex because it mixes Java expressions and statements with the rules languages thus it needs a real parser.
I was working on a project with Drools and had extreme difficulty because Drools error messages didn't make a lot of sense to me so I did a lot of looking at the Drools source code, running the compiler in the debugger and such and still didn't get very far. I switched over to the Jena Rules engine
https://jena.apache.org/documentation/inference/
which still emits lousy error messages but I quickly was able to understand everything about the Jena Rules Engine, get really good at writing extensions, and use it for things the developers said were unsupported such as using it as the control plane of a data processing engine that bridged stream and batch processing: I'd use Jena rules to control the process of setting up and tearing down reactive streams that would do a data processing job. That is, I got good at the Jena Rules Engine that I kept discovering new things I could do with it.
Clara is very similar to the legendary CLIPS rule engine from the golden age of ai
and if you find Clara is too small to understand, CLIPS is even smaller. The massive reduction in code size comes from being able to lean on the affordances of LISP to develop a DSL for writing rules, if you have to write a parser and all the stuff that comes with defining an external DSL rules language your code starts to get much more complex.
Jena Rules Engine is written in Java and I think it is very nice code and you'll learn a lot about RETE engines and other advanced programming concepts by studying it, you might even get more out of it than you get out of Clara. It was common at the tail end of the golden age of A.I. for people to write a first draft of a system in Common Lisp and then end up re-writing in it C++ for performance. Many of the ideas you get out of Lisp programming apply just well to other programming languages but require a huge amount of elbow grease and in-depth understanding of compiler technology whereas you can frequently use macros and similar affordances in Lisp to implement very sophisticated ideas with postage stamp size coded (that granted many people find challenging to understand.)
Also, having asked about Haskell, I'm not sure why you brought Java into this!
My area so I'm extremely interested - please elaborate
That makes me think of parmesan cheese -- "params" would be a better fit.
Good code is working code, code that pays the bills. Focus instead on writing code you can throw away easily, code that you are wholly unattached to and is isolated enough that rewriting it won’t cost absurd hours.
If that’s what you need to do to solve the problem, do that. The point is to stop focusing on the code as the work product, and instead focus on the solution as the work product, of which code is one part.
Good code should always solve the solution, to be considered good code. Bad code might solve the solution, it might not.
The problem with believing that "good code" is good enough to deliver business value is short-sighted.
Good code is highly maintainable so that you can continue to meet business objectives in a timely manner without regressions. Often this means (ironically) taking a little bit of extra time early on to think about how to make your code readable and "simple enough" for someone else to be able to jump in and maintain it.
> Good code is highly maintainable so that you can continue to meet business objectives in a timely manner without regressions.
I think you essentially agree on what people should do, regardless of whether you call this 'good code' or just maintainable code.
Write tests. Give things good names. Use comments to explain why you're doing something, not how to program. Don't copy and paste the same implementation x times because that way there's only one place to fix it.
But also sometimes it makes more sense to copy and paste over trying to fit an abstraction where it shouldn't be. :)
I think the purpose of questions like the ones by OP is not to figure out "rules" (which are useful only for beginners) but to figure out where and why rules were broken. Sometimes (often) the answer is time, but that in and of itself is a useful example.
Good intermediate (I suppose Sr. in our industry) level code is notoriously difficult to find examples of and mentor toward.
depends on goals. diverging modules should be copy-pasted, two modules that rely on the same functionality should be consolidated (not necessarily abstracted, but synchronized somehow). Those are both two common cases, so there's no general advice on which is better.
>Good intermediate (I suppose Sr. in our industry) level code is notoriously difficult to find examples of and mentor toward.
so much industry code and knowledge is proprietary, so I imagine that is by design. even intermediate code has a bunch of value to a company, even if the company lets go of that engineer to make their earnings report 0.1% higher.
I find that's a smell of limited languages: Maybe a language has poor error handling semantics, maybe it's not expressive enough to make a parameter generic.
It can also be a smell of not understanding the language well enough, too. Maybe there is no need to copy and paste, but the programmer didn't understand the language well enough to make a generic abstraction.
But you can only do what your tools allow you to do.
Seems not ideal for end users, unless you're working with microservices or something with well defined specifications that people are actually paying attention to.
Beside, "rewrite" here doesn't mean "new repo, new project, new everything" it means reimplementation, usually based on the lessons learned from the previous implementation, and that does include edge case handling, as well as expanded functionality to "underwrite" or justify the effort spent on the rewrite.
If you design your app as a collection of subsystems, it's sometimes faster to destroy and rebuild one of them than it is to refactor.
A good interface limits the blast radius of my refactor.
and that paradigm is exactly why my domain is the start opposite of front end web development. Code I write may be used by thousands of other engineers over decades. I can't take into account every edge case, but I do try to write with the assumption that one day some archeologists will uncover that code as some Rosetta Stone. Of course, modern demands never let me reach that ideal, but it teaches care and good documentation.
But hey, they do justify it in responses. So maybe they are indeed playing to their philosophy of "work fast"
But in a way you're right; if there is no to ensuring your code follows more conventions, go for it. That's exceedingly rare, however, as a situation to be in.
Is it because it's perfectly abstracted and encapsulated and SOLID and DRY?
Or is it because it's written in a manner, easy to understand, easy to extend, and easy to throw away and rewrite if necessary?
I'm guessing solid and dry are a big part, but lack of cleverness, language choice that doesn't require cleverness, and heavy reuse with libraries and frameworks, and choice of feature complete libraries is probably a lot of the speed.
Something in C is going to be more work than JS or Python or Dart and work takes time. Code you write takes more time than code that already exists, unless the code that exists disappears one day.
This. But IME when you write for this purpose you tend to end up with code that is in fact confusing, not documented, and not test covered. So you need to slow down even if your goal is to one day completely re-write that module.
Perfect abstractions are never perfect unless you completely architect out the product, down to every single edge case. That's virtually impossible with a large codebase, be it due to an API or even compiler level bug.
This is not relevant to the vast majority of developers, realistically (and is an important caveat to your original comment).
The first is the business problem, most important.
The second is the problem of maintaining and iterating on the solution to 1…
Without solving the first problem, the business dies and you don’t even need to care about the second problem… however the second problem can also kill the business if not solved eventually…
As I said, it depends on the project. I'm not going to approach a leaky faucet the same way I would an industrial sewage pipe. Fortunately the industry is big and I can choose to work on larger problems where longevity is valued over throughput. You're fine whipping together a React App in a weekend While I'd be more the end of maintaining the React repository. To each their own.
> Your “craft” is not to write code, it’s to solve problems.
Writing code is part of solving problems.
The quality of code has an impact on the various aspects of how a group of people solve problems.
As a small example - consider onboarding time for a codebase/system. Longer onboarding time means - lower profits end of the day for the organization (I'd also argue that longer onboarding times correlates higher talent attrition). And code quality has a strong influence on onboarding time.
Code quality has an impact on the "debuggability" of your systems. How quickly can you fix stuff when things go wrong?
Code quality has an impact on the "deployability" of your systems. If your code is well-done it is easy to deploy, redeploy, etc.
I can cite maybe 10 more properties crucial to org health, which are influenced at least partly by code quality.
So, "the problem" is not as simple as it may seem at first glance.
This is why the pursuit of “high quality code” is pointless; you are not an artist, you are aligned to solve a problem. If you do that, code or not, you are doing good work as an engineer. If you are not, you are not. Whether the code fits some arbitrary definition of “good” separate from your ability to solve a problem doesn’t play into it.
The second issue is that you think it is impossible to define an "abstract good" in code quality given a specific context (team, product, market). The "abstract good" stems on its own for the given internal culture and market situation. Through some common sense examples, it is easy to see how "abstract good" wields influence on "practical parameters" critical to business survival/thriving.
As an analogy, I can say the "body is healthy". I am aggregating a bunch of metrics to say - "this is healthy". It doesn't mean the term "healthy" is meaningless. The term "healthy" has useful meaning although not at a mathematically precise level. One could even argue that the term "healthy" captures something even precise mathematics cannot capture (it's abstracted at a higher level). Apply similar argument to the term "code quality".
Edit: Maybe it is better to explore the idea of code quality "via-negativa". Find what's actively harming beneficial outcomes. And remove it. If you cannot find many harmful things, then it has high code quality.
You can call that good code if you want, but my argument is to stop caring about the code’s “quality”, as a value it carries independent of the problem.
The physician - operates "via negativa". He tries to find faults with the given body, tries his best, and when can't find - he calls it "healthy".
The engineer/businessman can look at the code from an empirical point of view.
If onboarding is bad -> code is bad
If understandability is bad -> code is bad
If deployment is bad -> code is bad
And so on. As you eliminate these issues, your "code quality" increases (just like as you eliminate disease, the body becomes more healthy).
Look into say, Taiichi Ohno's Toyota Production Management - one associates "zero defect" ~= "quality". So, the term quality alludes to a continuous elimination of faults and shortcomings.
The aggregate placeholder/banner term 'code quality' stems from very firm practical sources, that can be inspected, amended and improved.
> make it work, then make it pretty, then make it performant
Solve the problem first, then improve the solution
On the other hand, those aren't the only two choices, by a long shot.
Code that works is, usually, acceptable. But code that is acceptable while making no unnecessary maintenance trade-offs is much better. Good code is code that uses standard techniques in standard ways to achieve a result without being verbose or inefficient. But that is a much higher bar than code that is literally good enough.
Writing bad code that just "gets it done" is the proverbial broken window. It's how you end up with shitty code-bases that get shittier with every change, until it all collapses under its own issues.
Doing this on established code-bases is basically what the classical duct tape programmer does. Sure, you deliver "business value" in the short term (normally to claim credit and gain favor) but at the expense of everything and everyone else.
Not saying one is better than the other. Sometimes being first across the line is make or break for a company. But I wish companies could be more honest about what they want.
Congratulations on figuring out how to create a massive churn rate among your actual good engineers.
I mean, juniors can do this. Once you're a senior, there will inevitably be come coupling you need to make in order to "pay the bills". Or you may make the first part of a system that will be a pain to re-write, even if it's the most elegant, readable code ever.
Nothing wrong with preparation.
A painter might study and practice enough to be as good as Michealangelo, but the attempt to be 'more good' or 'better than Michaelangelo' stops as soon he starts a serious painting. He has to finish the painting with the skill he has at that point.
I agree code should be unattached and you can throw it away. As soon as you start coding a program for a client, your attempts to become 'more good' as a coder and to write more ideal code, have stopped and it's time to make some code you can throw away or sell to the client.
But all the 'training' paintings you made on the journey to becoming 'good' have to be kept, because you trashed thousands of attempts along the way in pursuit of an ideal painting. The good painting is framed and hung on my wall. The commercial painting is sold to the client or trashed.
Yes, give up on 'good code' at work. Keep the ideal of good code as a direction to improve towards, not an end result in commercial works.
As soon as 'good' became explictly defined/bike-shedded it died anyway..... it's an infinite direction, not a limited thing that can be defined and boxed up.
The first is Jones Forth (https://github.com/nornagon/jonesforth), start with jonesforth.S and move into jonesforth.f. I really enjoyed following along with it and trying my hand at making my own stack based language.
The other is Xv6, a teaching operating system from MIT (https://pdos.csail.mit.edu/6.828/2021/xv6.html), not all the code or implementations are top notch but it shows you non-optimized versions (just because they're simple and more readable) of different concepts used in OS design.
If you're interested in the embedded world, there is a really neat project I've been following that feels a more structured and safe (as in fault-tolerant) while still staying pretty simple (both conceptually and in the code itself): Hubris and Humility (https://hubris.oxide.computer/).
Sqlite is supposedly high quality C code: https://github.com/smparkes/sqlite
For videos of someone (Casey Muratori) writing video game code and debugging it, Handmade Hero: https://handmadehero.org/
A blog post about how to write code by the same author: https://caseymuratori.com/blog_0015
For how to implement a fairly advanced type system, Typing Haskell in Haskell: https://gist.github.com/chrisdone/0075a16b32bfd4f62b7b
But, honestly, you're probably better off writing code yourself and learning by doing.
Actually most of the big Square OSS libraries are great to read - okio, okhttp, picasso.
Eventually, after working on, say, half a dozen code bases, you'll start to understand intuitively what good code is, providing you get lucky enough to find a good code base, or a code base with a significant amount of good code.
It's a long old journey, but once you have the skill, it never goes away. It's like learning a musical instrument or a foreign language. (By which I mean you can read as many books as you like about it, but without application, you haven't yet begun. Nevertheless, read the books.)
Warning: most developers never attain this skill, but almost all of them believe – truly believe – that they write good code; just as everyone thinks they are a good – nay excellent – driver.
Warning: no one writes good code. Good code becomes good through iteration, just as good writing becomes good by iteration/editing. The reason for this is obvious; but if you don't know why, then you haven't done enough yet.
Warning: everyone has biases. Learn to recognise yours and when you are applying them. Learn to ignore them and see things through a different lens. Explore with an open mind.
Iterating to good code is one of the most satisfying things you can do with software development.
It also links off to this which looks like a good read https://third-bit.com/sdxjs/
Someone can tell you "This is good code." but good for what? Why is it good?
It is fast code? Is it highly maintainable? Is it well documented and kept up to date? Is it a code that is highly reliable? Is it code that solves and important problem?
My rule of thumb is: Ugly code usually comes from ugly problems. Ugly code can often be some of the most valuable code, because... it does the ugly things! It does what we want 99.9% of the time, using heuristics, and other nasty stuff.
So don't judge code on if it is "good" or not. Judge it on if it does what the author intended, and if it doesn't suck too badly to read with no reason.
Code bases I've worked in and have opinions on:
Samba: Good code base, but you MUST understand the idioms of the codebase, or it is absolutely horrible. It also alas, has the wisdom of 20+ years of existence in it... so it isn't always pretty.
Illumos/OpenSolaris: Nice codebase. Get the SmartOS distribution and you can literally type a few commands and build an entire OS and userland.
FreeBSD: See above. Great codebase, ans also, it can build userland + kernel, though it takes a few more commands. I'll admit I haven't read this one in 20 years. But I always found it a good codebase to work in back when :).
Grab the source for a library you use all the time, you know the useful one but the API feels a bit off... Download it and look at why the API is the way it is.
When looking at code, do NOT neglect looking at the history of a given file or piece of code, it often can teach you quite a bit. :)
To folks starting out with Elixir, I suggest reading its standard library. From my experience, there was this aha moment when I started reading `Enum` module.
Also, Elixir's documentation is one of the best out there.
I would look for places the code seems to be really liked by its users. Maybe its very reliable, or extensible, or fast, or something else. How do they achieve that? Why do the users say these things? How do they measure / focus / make tradeoffs to focus on those attributes?
Then for painful to use software & common painpoints, why does this happen? Is it a fundamental design decision? Is it just sloppy code? Is in just intentionally slow to be more user-friendly? Or hard to read code because the focus is on speed?
It's all about the tradeoffs and intentional choices...
Ultimately I think the single most important rule for clean code is: skinny controller, fat model. If you are doing batch data, then this applies still I think. You should have all of the logic you can in the model, avoid data objects. And the code paths that alter things should be as thin as possible. I honestly think it is better to have a 5k line model if it avoids more.
The most unlcean code I have seen usually falls into the abstracted out processes in financial instutitons where they follow Clean Code advice and everything are a bunch of functions passing around some big fat objects full of getters and setters, and changing any functionality means adding changes somewhere in the process to check state and alter the state, which leads to loops and if statements everywhere to see which account type it is at each point etc.
But in OOP programming its supposed to be objects sending messages to each other. Every object should know everything about itself, which is what a fat model demands. Any more abstraction than that seems to get in the way.
Popular OOP passes Structs around (they just call them records or POJOs or whatever) through a bunch of "classes" that do x. But you could rewrite the code from Java or C# or C++ into C or Cobol and it basically is the same. Its just imperative code with classes as a nice way of getting rid of globals.
I’ve heard people say it’s a good book if you just ignore all the bad stuff, but how are you supposed to know what the bad stuff is if you’re a beginner? I think it’s time to stop recommending this.
What does "functions should be immutable" have to do with mutating classes?
The best example of a function without side effects would be sin(x). You call the function with an input and it returns a completely new output. The function should be thread safe and easy to isolate because it never touches any outside state.
Someone who does writes quite clean code is in my opinion Tsoding: https://www.youtube.com/@TsodingDaily Strongly recommend his YouTube channel!
It was a good book of its time in that it was influential and encouraged people to think more deeply about how to make code readable, but even when it was published I thought it had some terrible advice.
It's probably still just about worth reading, as long as you ignore all the code examples and appreciate that some of the thinking is out of date, and a lot of the rest is controversial at best.
I also find the style grates, as "Uncle Bob" is far too full of himself and e.g. "rips apart" someone's code to produce a worse refactor.
If you're looking for programming inspiration I love these videos
Jon Bentley - Three Beautiful Quicksorts https://www.youtube.com/watch?v=aMnn0Jq0J-E
Bret Victor - Inventing on Principle https://www.youtube.com/watch?v=PUv66718DII
https://www.amazon.com/Programming-Pearls-2nd-Jon-Bentley/dp...
https://github.com/Svxy/The-Simpsons-Hit-and-Run/tree/eb4b34...
god help the kids reading the trash that makes up 99% of the code out there today.
- Ratio of code:documentation INSIDE the source code
- Directory structure depth is “just right”; not too deep nor too shallow
- Number of dependencies is “just right”; don’t build things yourself, but also don’t import the whole world
- TTLD (Time To Local Dev); how simple is the getting started guide in terms of copy-pasta commands + automation + the right amount of context + easy-to-use tooling
- Code culture; follow industry best practices and make it clear where & why you deviate
My personal favorite one: `make todo_list`
We use keywords (TODO, OPTIMIZE, HACK, etc…) through the codebase and make them easily searchable with make helpers.
Ref: https://github.com/pokt-network/pocket/blob/main/Makefile#L5...
Note that "WEB" here is not "web" as in website.
Knuth's programming style is highly idiosyncratic, and there are many points with which I'd disagree. Furthermore, the choice of Pascal required all sorts of strange compromises. That said, I highly commend reading the source of TeX, to see how a brilliant computer scientist attacked a problem that is not nicely structured, using very restricted programming tools. You will not get much that you can copy into your programs, but it is an excellent, well-documented attack on an interesting and complex problem. (By the way, Knuth's TeX was written long before modern typesetting things like Unicode, PDF, and OTF fonts; you don't want to use Knuth's TeX for modern work; I use LuaTeX. But Knuth's version remains as a useful subject of study.)
Another good thing to study is Lions's commentary on Unix V6. This is interesting because it shows many of Unix's key abstractions implemented in a few thousand lines of code.
Just like if I were to write a book, it wouldn't be in the style of Plutarch's Lives, I'm still glad I read them.
And then read any code through that lens. Then read some different code and contrast it. What did you like more or less? What worked and what didn't? Where did it work and where didn't it?
Remember that code that is fantastic is some ways is often horrible in others (e.g. the legendary fast inverse square root).
Approaching it this way helps one consider the reasoning behind what makes certain code good, and forces one to examine the context of the code, which is also critical. And being opinionated helps you remember to apply those rules in the future.
The fact is someone programming languages - either by design or culture - just encourage unbelievable crap. Obviously, any code can be written well, but if one has to look for a needle in a haystack it's not worth it.
Gitlab, Bitbucket.
I don't really know what to expect. You want to see good production code, and the FOSS community is way way WAY better at this than some of the bubble gum you'd see in a professional setting.
Your question is too general, so I can't exactly give you a specific repository. I could direct you to BGFX[1] for a decent architecture of a cross platform renderer, but if you're not a graphics programmer, that may be a bad exercise, as you'd spend more time learning jargon than studying clean code. Or it uses patterns (or lack of, given graphics programming) that don't apply to your domain.
Good code is different depending on your sense of aesthetics, but also it's purpose.
Good code for an enterprise-life-blood type of system is terrible code for a "let's check if this is an idea" type of prototype (and vice versa).
For a relatively large and mature project, someone might suggest Linux, but it's a bit hardcore in my opinion. FFmpeg is actually really nice, and you can wrap your head around the general idea of how that system works, to the point where you can comfortably add new options and even introduce new codecs and containers in a few days.
I would posit that discarding github out of hand is not a great start. Pick one of the big projects on there and follow what is pushed in. You will notice how others interact with the code. You will start to notice who checks in good stuff. Follow them. Also pick your language. As a good style there could be a bad style in another for example C++ vs Python. Python styles in C++ would drive the C++ guys nuts and the other way around.
I’m sure I’ve seen some HN posts about this article, but I can’t remember the content of such.
In general, have a look at the Standard Template Library or Boost examples directory. Then some unit tests for simple GNU programs similar to what you are building, CLI command source like "ps" for OS interactions, and finally an OS kernel like Linux or *BSD. There are also several online classes offered by linuxfoundation.org etc.
Start with a small SBC like a pi4/BeagleBoard, and learn how to snapshot disk images (you will severely damage things while learning). There are also several open syntax formatting standards published by projects (and companies like Google), that will guide you on the local ecosystem.
Expect a Hazing in some places, as some folks tend to forget they were students once too.
It would also be wise to spend a few days studying security-auditing-tools, as one may learn to mitigate common ways people will try to break stuff. Detection and incident-handling is arguably more important than outright prevention.
Happy coding, =)
* Protocol implementations(e.g. TCP, HTTP, MIDI)
* Smaller compilers(PUC-Rio Lua, Forth-80)
* Commercial video game sources [0]
When studying protocols, you can compare apples-to-apples because the protocol has to work the same way by design, but the implementation can vary. With compilers, you're getting a look into programming in its maximally symbolic form - and every strategy a compiler uses is one you can directly apply to abstract your own code. And commercial video games have another mode of apples-to-apples in that the original release - the dirty, meets-deadline stuff - often can be compared with fan remakes and patches, which have the luxury of an exact specification and no deadlines. To actually ship in industry, you have to accept and know good dirty code hacks, but it's worth comparing them to their counterparts.
[0] https://en.m.wikipedia.org/wiki/List_of_commercial_video_gam...
https://github.com/grantjenks/python-sortedcontainers/blob/m...
Lots of source these days has auto-documentation comments. Good IDEs present that documentation, which helps guess what might be worth diving into the Step Into rabbithole.
Often, from a high-quality framework / library I learn a bunch about handling weird edge cases and about writing code for long-term maintainability. And, I often learn some useful constructs and techniques. (And, it's possible to learn useful things from not-so-high quality code too.)
s/read/train my AI on/
https://github.com/bigskysoftware/htmx/blob/master/src/htmx....
(Tooting my own horn) A Fitbit watchface that I wrote a few years back: https://github.com/GWBasic/Binaryish-Clock
An event/threading library for C#. I keep a fork in my Github because the original source was archived: https://github.com/GWBasic/retlang
Note that both examples are "functionally obsolete." The Fitbit studio environment is deprecated in favor of Android Watch; and if you're using C#, you can should be using Tasks to get similar functionality to Retlang.
I think reading code itself is only valuable when you need to explore a specific domain. Trying to extract coding patterns from an unfamiliar domain very difficult; it comes with unseeable assumptions.
Any code that is used and minimally hated by a large number of people is good code in my opinion. It's valued by its users, and ultimately that is what matters.
It may not be perfectly DRY, use popular abstractions, or whatever people today think of as ideal code, but successful software projects solve real problems every day. The authors have likely done a good job of balancing usability with code quality. And I think that's the best we should hope for.
Even if there're some caveats/hacks in these codebase, it might have a valid cause and you can learn something from how it evolved or how it (dirtily or gracefully) solves the issue.
Specifically something polished that also has complex UI elements like block indicators or embedded modals.
Look where good tech documentation lives.
Flutter is one good example of great docs. Ruby on rails is an anti example (even though the framework itself is great for its goal).
Beast mode. This is a great way to understand more about how the high level code we write actually performs. I learned the basics of compilers long ago but never though to apply it in this way, where you can have a reference implementation against which to test an assembly rewrite. Very cool!
I've basically never been blown away by someone's function definitions or whatever, but I regularly run into well designed APIs.
Maybe it's because the API is the thing that has the ergonomics, but the code is just getting things to work. Not sure.
You would hopefully find that the most compelling aspect of well-designed software systems is the data. In "data-driven" applications, 100% of the application state and configuration can be made to live in a database somewhere. In these scenarios, seeking code examples is not going to tell you much of anything.
My advice is to look at a bunch of SQL schemas (ideally, ones you know to be under successful products) and compare them to the problem area they support. Think about how you would answer questions a reasonable person might ask of that business by way of a query. Then, consider how much code you just now did not write to answer a realistic business problem.
Relational modeling can eliminate entire repositories worth of bullshit code that should have never existed in the first place. Do you want to train yourself to rely upon something that a true wizard can walk in and disaparate at the snap of his fingers?
The problem with with "good code" is that it tends to be code that -- more often than not -- is written by people I wouldn't trust to work on real-world large-scale software development.
On the other hand, any sufficiently large project that was successfully delivered necessarily contains almost exclusively "good code", because if it didn't, it would have collapsed under the overbearing weight of "bad code".
What makes the Android source code interesting reading:
- It's written by some of the best software engineers in the business, by any imaginable standard.
- It is an unimaginably successful project.
- It is codes that deals with the gritty enduring reality of programming in the large that cannot reasonably be addressed by toy samples of "good code". That's where real "good code" lives.
- As a programmer, that's the kind of scale I want to work on. Preferably as one of the principle engineers working on Android 1.0. But Android 14 wouldn't be awful.
Overwhelmingly, it is exceptionally good code. Occasionally it is less than happy code. But the places where it are less than happy are almost more interesting than the places where it's good.
The code is 14 years old. It's been through 14 major releases (34 minor releases). It started on phones with 320x200 displays, with megabytes of memory, and processors than could barely run a toaster. And now it runs on phones with 4k displays with 8GB of memory, on processors that are about 2.4 kilo-Crays.
If you're a junior programmer, every single line is better than what you're capable of writing. If you're a senior or intermediate programmer, there is serious food for contemplation. The question that should be asked at every turn: if I was on the Android 1.0 development team, what could I have done to make this happier code?
And I'd really like to see what the "Less code is better code" guru (a recent HN posting) could do with androidx/fragment/app/Fragment.java and friends. A perfect example of a "Good Code Guru" that I would not trust to work on any of my projects.
---
There is no such thing as bad code; but some code is happier than other code.
-- Herbie Hancock.
Or what Herbie Hancock would have said if he were a programmer instead of a jazz musician.
The runtime/std lib of the PL of your choice. Your text editor or a plugin that you use etc.
1. Busybox. It's basically a collection of common Unix utilities, from common commands like ls or cat, to system daemons like crond and init. It's a good way to learn more about how Linux and other Unix-based systems work under the hood. Busybox applets are pretty independent from each other, so you can just take one and focus on it specifically, digging into the common library code if you need to. The utilities aren't as fully-featured as their GNU coreutils counterparts, which makes it easier to understand what they actually do. Busybox is not a toy project by any means, though, it's used in many embedded devices and leaner Linux distributions, Alpine being the prime example. However, it's written in some pretty dense C, with a fair bit of pointer magic involved, so if you don't understand things like the equivalence between a pointer and the beginning of an array, some things might not make sense.
2. The Go standard library. Unlike many programming languages, Go does not rely on much external code. Whereas Python delegates zip handling to zlib, handling of TLS connections to Openssl and so on, Go just includes all of this in the standard library, and it's all pretty readable Go code. If you want to understand many common algorithms or file formats, everything from sorting arrays to parsing JSON to sending and receiving HTTP requests or common cryptographic operations, all written in a readable style, in a language much higher level than C, just look at the Go stdlib. Go even fully implements everything needed for its own compilation, including linkers and assemblers. I haven't read these parts much, and a lot of that code is transpiled from C, so I can't say how good the code quality is.
3. Serenity OS. It's a hobby Posix-based operating system written in C++, with no external dependencies, not even libc or libstdc++. They have their own homegrown implementations of every part of an operating system, from a monolithic kernel, to common Unix utilities, archive handling, audio and video codecs, common data structures, like vectors (growable arrays), hash maps, locks, mutexes and other concurrency primitives, a custom string implementation, a window server and a GUI library, including a fully-featured event loop system, their own window manager and many common GUI widgets, to actual applications and games. They even have a custom web browser with a custom web engine and JS interpreter. As a rule of thumb, if something it's either in Busybox or in the Go standard library, there's a good chance it will also be in Serenity. Again, their utilities do much less than their non-serenity counterparts and are far less optimized, but that also means there's a lot fewer layers of abstraction to deal with and that the general principles underlying their implementation are actually easier to understand. The fact that everything is in a single repo, neatly organized, written in one language with common conventions, just makes it really pleasant to read. Their code quality isn't always the best, but the fact that it's C++ and not C does make things easier. Even though I haven't actually used the OS (because of accessibility concerns), it's one of these repos that I always have cloned on my computer, and it's the first place I look if I'm curious how a particular feature or app can be implemented.
4. If you're in any way interested in AI, everything written by Andrej Karpathy, notably Micrograd and Nano GPT. There's also Tinygrad, a bigger but still understandable take on Micrograd. Unless you're an expert, you need to watch Andrej's Youtube videos to actually understand the code, but the feeling I got when I actually understood the principle behind Micrograd is one I will never forget. I consider it to be the most beautiful piece of code I've ever seen, it basically embodies the whole principle of what a neural net is in 200 lines of code. Everything else that the big libraries do is basically just implementations of actual models, optimization and glue code, such as for loading data and such. It's often crucial optimization, optimization without which modern neural networks wouldn't be possible at all, but just optimization nonetheless.
5. Everything concerning Elixir, both the standard library, other libraries written in it, as well as open-source Phoenix web apps. Deep down, it's basically a Lisp without the off-putting parentheses. There are a lot of lessons to be learnt there, from the power of macros and the fact that things like "if" can be written in the language itself instead of being a special construct, to the power of pattern matching and the pipeline operator, to the advantages of its concurrency model and functional programming in general.
To generalize beyond these specific examples, if you want to understand something, find a smaller version of it, and try understanding that. The smaller version might just be a git commit from a good few years ago, with much fewer features, but it's better if it is a different, more basic (but preferably not toy) implementation of the same app, feature or algorithm. Don't read V8, PyTorch or Postgres, read Lua, Tinygrad or Sqlite instead.
Instead: write code for whatever fits your fancy, and request experts to give you feedback.