Tips for Reading Code (2014)
wiki.c2.com
wiki.c2.com
Terms present in the UI (or API docs) of the segment you wish to understand are a good starting point.
Also, how better is it than the code map on Sublime Text or syntax colouring we have anyway?
This did let me work on the code when away from my workstation, but I didn't otherwise find it very helpful. If you have lots of code, some form of code navigation is very helpful.
No syntax highlighting, but it was quite useful. Off the top of my head, the two main benefits are: 1) You free up your computer for the review 2) You get a different form factor to interact with.
Regarding the first benefit, having a printout means you can use your computer to run the application, look up docs, and even (ironically) navigate the code in your IDE.
As for the second, having a printout means you can spread out the pages to compare different bits of code, draw diagrams, connect arrows, circle bits of code for notes, etc.
It was pretty useful.
I can doodle and write remarks on top of the code, but I think much more important aspect of written code is it gives different friction profile than browsing the same code with editor (or on a screen in general).
This "friction" thing is hard for me to put into words, as it's quite subtle feeling. I can't freely search the code for identifiers nor run it, so I'm sort of forced to make notes and remember things, but I can mark parts of the code in different ways, and I can write an alternative pseudocode alongside the original, and I can add a TODO or REMOVE or UNNECESSARY marks.
I do a similar thing for code I need to write, except, obviously, I have nothing to print yet. I write the parts of the code on paper, (what allows me to omit all the uninteresting trivia). Then I can proceed with what I would do with printed code.
However, I'd say it's only worth it if you need to do "deep reading" of convoluted code, or if you're in very unfamiliar territory. Also, it doesn't work when the "printout" becomes too large—a full printout of a 30K LOC project is about 600+ pages. No way you're going to read all that!
* I haven't worked with the codebase in the last 3 months
* I wasn't the author
* I'm heading into a major refactor
Paper is malleable. I can cut it out, slide things around on a desk, to show better relationships.
I can write all over the code without forcing any format or alignments.
It helps me understand the control flow, and the exposed API, faster.
1. actually read the bloody code until you understand well what it does. Don't stop before that, or it won't count.
2. read as many different projects as you can -- different languages, technologies, paradigms, coding styles.
3. Don't spend time on code-comprehension tools, it won't work (speaking from personal experience. The latest was with "Understand C++" from SciTools, a great tool BTW). Master grep.
4. Drawing flowcharts, running under debugger, etc -- all that would be great but nobody allots time for that anymore (neither yourself nor your managers).
5. Don't hope for usable doc for external libraries. Exceptions exist, but it won't be a rule.
I've been chided for not doing this in the past (I just read the code usually) and would actually agree with that person in retrospect.
More generally, studying the code under debugger means you need particular input data -- the one suitable for the part of the code you study. That can take time to get.
Seems weird to say a team was always hung up on ensuring the correctness of their code. I guess it depends on what you work on to an extent.
In fact, isn't this usually the case with any application not yet in production?
As an example, I had a teammate who would get blocked on needing JSON that matched a production set of data. Yet, she knew what the various fields of the data would most likely contain. She ended up writing the code anyway without the "production" data she asked for and it worked great minus a few small changes. It was ultimately just psychological block, IMO.
Since most of the code I read (C# code) can be ran on my platform, I'll sometimes run into that rare piece of code that I can't believe it does what it says it does, in those cases I'll download the repo and debug it and many times be surprised to find there was a gap in my understanding. Reading code helps me find knowledge gaps I didn't even know were there.
My brain will tune out whatever piece of code I've already read once. So if I initially skim some piece of code, figuring out what it -really- does is HARD.
Every subsequent look at that piece of code will be colored by my initial (mis)understanding of how it works
Then after some time when I look at it again, I notice that actually I missed a step in that method which actually did contain the problem, which I did not see by skimming over the piece of code earlier.
This trick is also useful proofreading documents you've read many times already.
4.) I only had to resort to flow charts a couple times so far, mostly with Java libraries to externalize the data and control flow, because it grew just waaayy to large to keep in my head. A particular vivid memory I have here is netty (which also has a tendency for a particularly convoluted and complex control flow realized through modifying pipelines conditionally in many places).
Especially because ag tends to be smart about what it will ignore (e.g., the contents of node_modules, etc.)
I don't know how they turned the most basic HTML page into something that takes 30 seconds of spinner to load the same thing.
I counted 20 seconds to load http://wiki.c2.com/?SoftwareMasterpiece
Edit: opening the same link again is instant. Weird.
I'm on a 250 megabit connection. I haven't seen a page take this long to load in months.
Once a particular page is visited, it does load quickly. But for new ones it's always the same experience. Unless they're buckling under load of some kind and this provides a fancy cache, it's pretty ridiculous.
---
edit: decided to watch the network on a new page, http://wiki.c2.com/?LiterateProgramming for instance. Most everything downloaded in 10s of milliseconds, but the LiterateProgramming XHR request sat waiting for 6.5 seconds. Maybe they are struggling for some reason?
In fact, if I remember correctly, the slow load times are because it loads the entire wiki tree in the background to make linking easier for the engine... But a worse experience for the user.
I'll be hosting a mirror (I think the license allows for that but I'll have to check) and then integrating it with the one hosted on the original domain.
I also kind of want yo set up some kind of edge caching static rendering with loose constraints on how up-to-date pages need to be (because the point is providing a read-only Lynx/CURL experience), but I think this might end up being too much work technically or politically.
I am really frustrated with the new c2.com pages :P
I don't feel this is what Wiki was meant to be.
If you're hot under the collar, please cool down before commenting here.
In lieu of the article's content, I'll spare readers the trouble and provide something better (and faster to load!):
-Read lots of code
There you go, you're well on your way to becoming a code-reading master!
I have no idea what happened, but I presume the new version just doesn't handle the HN effect well enough.
Not at all.
c2 now uses Federated Wiki, which for engineering reasons (I think), loads the entire wiki tree as a key part of the json that serves as the content for every page.
Thing is, these changes come with loads of tradeoffs the newcomers usually are completely oblivious to and the new product is of lesser quality than the original, but they feel good about themselves now :)
http://wiki.c2.com/?OverEngineering
I didn't wait for it to finish loading and actually read it though.
The general content seems on track, if obvious ('find the high level logic', 'draw flowcharts if it's complicated', 'grep for whatever you're looking for') but some comments are out of date ("print the code because there's more code space on your desk" no longer applies, I literally have more square centimeters of monitor than of desk space and that monitor area can scroll!) and some are flat-out wrong ("software doesn't always have an entry point"? Really?)
I didn't comprehensively debug the whole page but in a cursory 2 minute trace in Chrome Dev Tools, I found javascript that makes an XHR request for a 265k text file[1] with 36000 lines. The javascript code then parses the names (which takes ~300ms) and then seemingly does nothing with the information. It doesn't show up anywhere on any subsequent DOM rewrites. The actual visible content that people read is only 10k.
For mobile users, you punish them twice by (1) subtracting 265k from their data plan and then (2) wasting battery power on cpu to parse it for no reason. That one example of useless bloat is probably multiplied dozens of times elsewhere in the lifecycle of the webpage.
[1] don't click on this if you're on mobile phone: http://wiki.c2.com/names.txt
https://web.archive.org/web/20160709111543/http://c2.com/cgi...
The Wiki was remodeled recently: