Interesting Codebases
medium.com
medium.com
if i'm not familiar at all with the language, or more specifically how the language architectures the program, i'm just going to be spending a lot of time looking at stuff that probably isn't the meat and bones of the library/app.
take for example, his first suggest codebase, seastar. i haven't done c++ in years (school, using turbo borland) so where do i look? "apps" maybe? nope just seems to be a folder of libraries. ah, probably the core folder. whoops there are 20+ files/headers. should i dig into this assuming that's where most of the code for the app/library is?
i suppose it would probably be nice if there's a site that explains how most programming languages layout their code. e.g) javascript generally is laid out similarly now, as is ruby/rails. so a site to explain the general layout structure would be kinda cool. it's kinda late at night so maybe i'm overthinking this.
I either have something specific in mind that I want to understand, e.g in the case of Seastar, I wanted to understand their reactor design and implementation, so I locate the respective file(s) that implement them and I start from there, and then I just branch out to other files -- I usually keep multiple vim windows open, and I take notes.
When I am not looking for a specific answer, I choose a directory, and then I sort its files by size (e.g ls -lShr *.{h,cpp,hh,cc,java} ). I usually sort by file size in descending order(smallest first), but some times it makes more sense to sort by largest file first, and I start from there. I still map my way around, and if something stands out, I open the respective file in another vim window/tab, and look it up, and then continue with the previous file.
I also have a script that checks what are the most crypto-y file in a directory.
Could you share it? :)
He walks you through his whole code review process starting from cloning/building the project:
grep -r -- 'main[ ]*(' .Once you get it building and running, you can now start making simple changes in order to find your way around.
[0] Developers should pride themselves on how (relatively)simple this process is for their project(s).
Documentation should not be a matter of "every function has a docstring", a description of the architecture is always useful.
Short, readable functions and good comments make it quite easy to follow.
Seriously, he expects people to believe he's evaluated the code of the Linux kernel, the Chrome browser, Postgres, LLVM, Tensorflow (just to name a few, less than half of his list), deeply enough to be able to make statements like "finest codebase in [x, y category] that I've seen", while also being the CTO of a company?
Same is true for Chrome. I was interested in the code for the various UI components, but I looked into other bits here and there.
I 've studied most of Postgres, LLVM and Tensorflow codebases - as in, I went through pretty much most if not all files looking for interesting bits (and finding plenty). I guess there's no way to "prove" anything to anyone, but I don't really care or want to do that either -- it's not about bragging rights, if that's what you implied; I just thought I 'd share a list of codebases I came across that I thought were interesting and worth of other people times.
As for my job, and my work, you may want to check out my Github profile (https://github.com/markpapadakis) -- though there's only public stuff there. I guess I like to spend my free time learning instead of say, watching tv or waste time elsewhere:)
It's hard to even make an observation like "names are meaningful" without a lot of context.
I am interested in exploring codebases now that I have sufficient experience and feel confident (and not initimidated). This will be useful because then I can start from the smaller (simpler) ones and move on to the complex ones later.
Any recommended path, please do suggest.
Some other codebases are so vast it takes a lot, lot longer to understand them enough to feel 'comfortable' navigating them (e.g the Unreal Engine codebase).
Most codebases however are quite small, and within 1 hour, or a few, you 'll be able to understand where to go to find what you need, which are the primary data structures and functions, etc.
I tend to run into new large codebases pretty frequently just from having to work with them. I'd say that happens maybe once every month or two for me, and I'm not even prioritizing this. I literally didn't know Ruby two weeks ago and now I've read through RSpec because I needed to submit a patch to some RSpec tests that were doing something unusual. For a CTO who's listing all the things he's seen his entire career, is two dozen codebases really that surprising?
Classic expert systems, but IMHO not outdated. I think they will make a comeback soon once we understand how to integrate probabilistic reasoning, logic and connectionist approaches.
[0]: https://news.ycombinator.com/item?id=13854431
[1]: https://github.com/google/leveldb
[2]: http://research.google.com/archive/bigtable.html, section 5.3
As an aside, how do you read codebases? Do you read every single line to understand what's going on? Or you get general idea about design/architecture? What are the proven strategies to read code bases?
It is is used internally for sparse matrix representations in Python, R, and Matlab. The entire library fits into 2100 lines of concise yet well documented C code. It is now mostly installed bundled with SuiteSparse, but the link above has the 2006 codebase from the original stand alone library.
These days, it's hard to find time between work and an absolute deluge of interesting stuff to study and play with.
There are a few parts of the codebase that were automatically transpiled from C, but the rest is usually very readable.
https://en.wikipedia.org/wiki/May_you_live_in_interesting_ti...