You Don't Read Code, You Explore It
prog21.dadgum.com
prog21.dadgum.com
Here is the thing, everything you read in a story is supposed to convey imagery of things like what you may have already experienced, they are already internalized, you "see" them when you read as if you were there.
Allow me to use another popular space as an example, music. When you are first reading sheet music, you see notes on a stave, key signatures, different shapes representing different durations. At first you mechanically take that understanding and laboriously turn it into actions on your instrument. But after a while, if you do it enough, the shapes become recognizable as rhythms, the tones in the staves become tones not symbols, and then you stop "reading" music, you look at it and you can hear what it will sound like. And by that time you can make your instrument do what ever you hear.
Coding is not entirely different, at some point you don't see syntax, you see algorithm, you see inter-relationships of data structures, you see flow. After a number of years of coding I got to the point where I could see what code was doing pretty easily (except for obfuscated code which is always jarring on first look). I stop seeing code loops and start seeing iterative processing, if statements are branches on a path.
Anything in words or symbols, is code for something else. Whether its a murder mystery, a symphony, or a sorting algorithm the words and symbols are there to express the idea inside your head you can understand it, I think it is all reading though :-)
Starter guide here for any interested: https://gist.github.com/bitemyapp/8739525
Do not pass go, do not collect $200, go straight to Haskell.
It's also possible somebody has told you Haskell is a declarative language a la Prolog. It is not.
Try this: https://gist.github.com/bitemyapp/8739525
Specifically: http://www.seas.upenn.edu/~cis194/lectures.html
I would still say that programs are quite different from any natural language text in that they require a sustained effort to learn "what really happens here" as opposed to "what general kinds of things are being done in the different parts of the program".
A murder mystery is a genre where "one little detail" often is where the whole direction of the story is said to go. But that's usual a single discreet gotcha, added for theater. A program often consists entirely of such things, added not for theater but because computation can't help but work that way.
I think that's what the GP is getting at. Even with domain knowledge, really understanding a code almost always requires diverging into other parts of the 'text'. You can defer it if your thought process is structured enough, but you still have to do it eventually.
Much the same you can't read code (code review) to look for errors, you'll miss most. You need to run and test the code to really find the errors.
I think that is the take away from the OP.
I've also found it very useful to slightly modify the code and compare my expectation with the actual result of a test run.
Not being able to see what types code deals in, jump to definition, and find usages makes me feel crippled when exploring a new codebase.
One of my wished for programmers everywhere is that tools like Github and BitBucket start analyzing projects and letting you navigate better. I think that could save thousands of engineer hours.
This topic alone is a huge reason I started using Dart and eventually joined the team. Trying to figure out how a very large JavaScript codebase works is so incredibly painful, doing so in Dart is incredibly easy. This is also where very reflective libraries like Guice go wrong, and why, in my opinion, meta-programming should be used very carefully and sparingly.
Although I think Light Table is trying to be like what you're saying, incredibly liberating in it's ability for code navigation. But it's still young.
As for type annotation and static analysis ... you don't need to have static analysis for what you're talking about. Yes, static analysis works best for static languages, but you can have code auditing in dynamic languages. An auditor can dig through and profile your stuff while tests are running and learn about interactions between various entities in your code. And if you're doing anything in a dynamic language properly you're doing a lot of tests.
Take my wish for Github to be more navigable. In a statically analyzable language they can run the analyzer and update their index on every commit. If they have to run tests then they have to basically add an entire continuous integration service which is a _lot_ more work, resources and security risk. Why make it harder than necessary?
Maybe I should try out static languages "done right" (type inference, etc.) before I give up on them.
On another note, being able to read other peoples' code is one of the strengths of Haskell. The Functor/Monad/Monoid/etc stuff becomes a way to know, based on a common vernacular, exactly what kind of interface is being exposed and what sort of data structures you're working with.
In my experience, the majority of existing codebases I've worked with tends to be this way, although there are exceptions where everything is so simple and straightforwardly written that reading them is almost an enlightening experience.
If anyone is interested, I suggest the publications of Anneliese von Mayrhauser and A. Marie Vans from the mid-90s as a possible starting point. They did a lot of work to reconcile earlier theories and paint a more unified big picture. Václav Rajlich is another name to search for, with several interesting publications in the early 2000s.
As soon as you branch off from basic type systems and add in, for instance, subtyping or type classes or existential quantification, this ability is getting weakened. Now, in order to understand what a piece of code does, you need to understand some context: 'which implementation is it?', 'what do the possible implementations have in common?'. The effects of this small complexities add up until there's too much information to be kept in biological memory at the same time, the oldest thunk of information gets purged.
I do think that typed and pure languages have big advantages here: the information that I need is available immediately from looking at the type of the reference to it. If the types aren't funky, I can assume that it will terminate, throw no exception, not suffer from data races -- I can exclusively think about what it does, not how (btw, I'm not having one specific reference language in mind right now, I'm just thinking about what would be possible).
i just tried it and that is a pretty big claim.
e.g. https://sourcegraph.com/code.google.com/p/go/symbols/go/code...
Am i wrong or isn't this pretty much what Intellij calls "Usage"? At least I don't see the difference, could you please explain it?
general feedback: the website is too slow. it takes me a couple of seconds to load each page, rendering the page pretty much useless.
[1] http://en.wikipedia.org/wiki/Data,_context_and_interaction
Nowadays I take it as a given that I'm probably wrong, but I start rewriting anyway. Worst case (and most common case) I have to toss the code. But I learn. Plus there's a different place your brain goes when you feel like you control the code vs. looking at it behind glass.
Code is nearly impossible to read if you don't understand the domain.
If you do it any other way, it won't necessarily make sense. This is really the only way to do it. (Though I'd be interested in hearing other perspectives.)
This was a hard-won lesson for me because we programmers tend to make the control flow of our programs start at the bottom of source files.
I know something like that happens with me now when I look at art after going through art school - Micheal Parson's "Talk about a painting: A cognitive developmental analysis" is an excellent paper on the topic. AFAIK it's only available behind a paywall though:
http://www.jstor.org/discover/10.2307/3332812?uid=3738736&ui...
The example he gives in his talk is the Axiom algebra system[2], which was revised to use literate programming style - the source code, with usage examples, is contained entirely within the books.
[1]:https://www.youtube.com/watch?v=Av0PQDVTP4A, [sldes]: http://daly.axiom-developer.org/TimothyDaly_files/publicatio...
[2]:https://en.wikipedia.org/wiki/Axiom_%28computer_algebra_syst...
http://sherlockcode.com/demos/jquery/
It's recently seen a spike of interest and I've started working on it again. The beta sign up link is still active if you are interested in getting updates.
I just started a new job at a really interesting agency. I got put on to a 12 month old project, a huge web application, that started life overseas, moved back here to Australia, and according to git-blame has then moved through the hands of nearly 15 developers, a solid 70% don't work here anymore (most were contractors).
So, the codebase is a mess. But, with Xdebug and a neat client for it that gives an interactive console when you hit a breakpoint, two weeks later I'm already understanding the twists and turns far better than I ever hoped for!
I'm currently envisioning (and trying to build) something where the types/func definitions are hyperlinks, and they jump to definition in an overlaying window similar to when you navigate in Spotify (the web based player). So you can quickly explore something without losing context.
- To grep it
- To debug it
- To read over it
- To rewrite it