Learn to read the source, Luke
codinghorror.com
codinghorror.com
Jeff can talk about the importance of having source available, but his actions speak louder than his words. He's built a very successful startup on top of a closed-source stack. Having the source isn't as important as it seems, then.
Yes. Recently they even went a step beyond the traditional "shared source" thing by releasing it under the Apache 2.0 license.
http://aspnetwebstack.codeplex.com/
Less so with MSSQL. But it's less of an issue there, because MSSQL provides a very good view of what's going on under the hood to begin with, and Microsoft has done an extremely good job of documenting the whole thing.
Is the Sybase code the biggest problem? Hell no. How about the fact that there is no "standard library" and there are no less than four different hash table implementations written by different people at different times -- only one of which you should probably use, although you wouldn't know it from the "documentation"? That's a pretty big one and almost definitely what I'd characterize as "historical baggage."
That said, one of my the nicer things about developing in .NET is that the library source is so accessible. I can easily step into library code from the debugger if I need to. If I haven't already downloaded the source code for that component then the IDE will automatically grab it for me before stepping in. So for my purposes (just wanting to figure out WTF is happening under the hood), the source code generally feels much more accessible on Microsoft's platform than it does on more orthodox open source ones.
1: http://weblogs.asp.net/scottgu/archive/2007/10/03/releasing-...
. . .a person who has lawfully obtained the right to use a copy of a computer program may circumvent a technological measure that effectively controls access to a particular portion of that program for the sole purpose of identifying and analyzing those elements of the program that are necessary to achieve interoperability of an independently created computer program with other programs. . .
https://www.facultyresourcecenter.com/curriculum/pfv.aspx?ID...
I might not agree with Jeff's choice to demand source code only down as far as his database API, but I've got to admit I've read no more of the source code to MySQL (upon which a _lot_ of my work relies) than I've read of the Oracle's source code. (And I'm pretty sure I've not looked at the Apache httpd source more recently than 1.3 or so)
Many commercial products also have source code available as part of certain deals.
As much as I like open source, there are also other business models where reading the source is possible.
Not so absurd, really. Reading well-written source code is a great way to learn the finer points of the art of programming; it's not just for fixing bugs. In fact, entire books have been published that consist of annotated source code, the most famous probably being Lions' Commentary on Unix and Knuth's "TeX: The Program".
I "grew up" in PHP, and while the majority of the open source code I've read has been difficult to parse, at best, I've read some amazing code. Most from intelligent peers for whom I still hold a deep respect.
It's what has kept me inspired. It's what brought me to eventually appreciate Javascript (server and browser), Actionscript (3), Java, Python, C, Ruby, a few others, most recently including Clojure. And when it comes to amazing code (regardless of language), I will gladly sit and read with more intrigue than the best fiction has to offer.
It's the massive amount of mediocre or less that defines the divide making source code painful. I skim a lot. I look for whitespace, the occasional well-written comment block, proper variable / method naming, things to hint that this was written to be read by another human who likes to read, and from there I dig in deeper with the hopes that I have something worth reading.
That is to say, I LOVE to read code, but it can be difficult to find the code worthy of the smoking jacket and voting-age scotch.
For sure, there aren't a large number of books with this structure. So what?
Nobody reads other people's code for fun.
Not true for me. I /love/ reading code from great engineers. I've learned a lot doing so.Absurd? It's basically what I did recently with ClojureScript One, replacing brandy with beer and deep leather chair with kitchen table and chair. I found it very enjoyable and enlightening. And I'm not trying to brag, I really don't think this is a mark of anything special.
IIRC publishing and highlighting code that was interesting to read was one of the goals that Peter Seibel wanted to tackle with Code Quarterly (which didn't pan out, but still, he didn't think it was an absurd suggestion). I also seem to recall reading code being something a lot of the people interviewed in Coders At Work described as being valuable. And one of the stock questions Seibel asked everyone in that book was if they had tried literate programming ("a la Knuth"), which is really just a way to make large pieces of code easier to be read by someone else. All's to say, absurd it clearly is not.
Writing good source code despite not liking to read source is about as likely as writing a great novel despite not liking books.
And like reading books, it only became enjoyable when you reach a level of fluency. Sadly lots of developers never get there, and so they go through convolutions to avoid reading code.
If "the source code is the ultimate truth", then your source code is indelible---you can never change your implementation because you've given users freedom to depend upon the behaviour produced by any line of it. If you don't want people depending on implementation details, then you need documentation to hide those away.
The problem is that the source can only tell you what a program does, not what it is supposed to do. If you don't know what it is supposed to do, it can be difficult for consumers of the code to know whether some behavior is is intended or a side-effect of the current implementation. Likewise, code maintainers can be prevented from changing the implementation when they don't know if consumers are relying on undocumented behavior. It's more difficult to file bugs against undocumented code; how do you know it's a bug if you don't know what the code is supposed to do?
In brief, good documentation and good code are a virtuous cycle. Reading the source is often necessary but it should be viewed as a failure of documentation.
Of course the source code is the ultimate arbiter of truth. But having a few roadmaps to that source code is _incredibly_ valuable. And as long as the underlying code does "what it says on the box", there's no reason to read the code.
Reading your stack's source should be a last recourse, not the default mode of operation. (Yes, I do read source code of my stack. Plenty of it. Which is why I appreciate any occasion where I don't have to.)
And when I see his "brilliant HN post" mention that suggests that "sometimes, you recompile your compiler", I'd like to smack some sense into people. You really don't. I've been working on low-level software for a loooong time, and I find about one compiler bug a year. I even do have a bit of a background in compiler writing. And yet, the sane choice is to write a small repro case, file it with the maintainers, and write your code to work around that bug, at least in most cases.
I definitely see a hallmark of experienced/skilled coders as not being afraid to follow the trail of code farther than I sometimes have patience for.
For me the biggest one was an FTP library, all it did was figure out when the server stopped sending data for a particular command and then run a Regex over it, populate an array of objects and return them.
Unread source is like a David Copperfield trick, it's magic, once you read the source and know how it's done the magic is lost and you understand what is really going on behind the hand waving.
if understanding the algorithm involves boning up on two semesters of type theory or graduate-level courses in algorithms, number theory and abstract algebra, as debugging problems in modern databases, compilers and high-performance integer libraries would, then having the source code is probably not going to help you as much as you think it would...
I mean, sure: for everyone there are some problems that are so obscure as to be near-impossible. But if you go through life always deferring those solutions (by calling tech support, or giving up, or playing voodoo games until the problem goes away), that set of problems will never shrink. You'll end your career, broadly, just as incompetently as you started it.
If, on the other hand, you make a practice of always digging for bugs, even across library boundaries into "other people's" code, you'll find over time that things like compiler bugs stop looking so scary.
Not having good documentation demonstrates a lack of respect for the user's time. To be a successful project, people of varying skill levels should be able to use it.
In order to compete successfully with closed-source Unix variants, GNU had to have as good documentation as its competitors and the result was excellent documentation (even if Info files were a bit baroque). The result was comprehensive and useful manuals for GNU projects such as GCC, Bash, Emacs and so forth. It's a real shame developers today haven't followed in their footsteps.
Also, how did you get so far in your career as an MS developer with such limited access to source code?
I'm not Jeff - but that situation has occurred multiple times in my career. Along with the more problematic one of there being documentation, and there being serious discrepancies between the docs and the code.
I like to have both by preference, but if I had to pick one I'd pick the source. I can figure out what it does from the code. I can't figure out the bugs from the docs.
Both of these situations outnumber the times I've had large code bases with good accurate documentation.
Also, how did you get so far in your career as an MS developer with such limited access to source code?
I'm not an MS developer, but from those I know there seems to have been pretty wide access to lots of source for some years now - you just can't fix and re-distribute it :-)
From reading source code both for fun and for purpose e.g. like the Android, Minix, QNX, Linux and NetBSD, network protocol stacks, filesystems, web frameworks, CouchDB etc. I got a lot of insights into interesting software technologies and architecture patterns. Good software engineers and architects should be good and fast at reading code.
We moved to J2EE using EJB3, Hibernate, Eclipse RCP. Our application was meant to be 3-tier, but actually it's more like 10-tier. We have hibernate mappings, java model, xml files specyfing possible queries and reports, java EJB3 beans wrapping these xml files, java classes for DTO, xml files specyfing possible views and editors in RCP, and xml files specyfing how to map from query to view or editor. And java classes for custom code in views/editors.
When I want to see what database column is shown in view, I need to start with view class, and descend all those layers down to hibernate mapping.
In our previous qt framework we had one xml file per client view, specyfing columns/sorts/filters/etc just for this view. Our consultants understood these files and changed them when they needed to. Now they would need to understand all those layers.
Now I think more than 2 layers in application is antipattern.
Shameless plug: My startup, http://www.thinkfuse.com is hiring developers who already know how to read the source! Email me at brandon@thinkfuse.com if you're in Seattle and looking to join a bunch of great developers who know how to build cool stuff and have a fun time.
That rings so true here. When I started Python web development, I needed to understand some concept related to middleware and handlers (somewhat foreign coming from PHP). My first thought was to look for blog posts explaining how it works in Django, but that wasn't satisfactory. I took a chance and dove into the Django source code— going against the voice in my head telling me, "You'll never understand it!"— and found myself learning so much. It was great!
In software development, we're taught to abstract everything and only think of the smallest problem, but this sometimes forces us to think of libraries as magic. This was a problem for me as a beginner, but it's been getting better as I've progressed.
I'm specifically thinking of strace, which I've used to diagnose problems that I was having with apache and chromium, among others. I don't think I would have got anywhere from reading the source. lsof is another, though I can't offhand think what I've used it for.
If you do find yourself in the source code, being willing to play around with it is invaluable for working out what's going on. If nothing else, you can insert a printf to confirm that you're looking in the right place.
I don't think it sucks because it's harder to write. I think it sucks because it's not strictly necessary to document your code in order to compile/ship it, and it's easy to justify putting it off. Poorly documented code is a form of technical debt, a compromise between getting it done right and getting it done right now.
(Note that the above refers to browsing the source of Emacs itself, not using Emacs to browse any arbitrary source -- unfortunately the tools for that are still very primitive).
[1] https://github.com/clojure/clojure/blob/master/src/clj/cloju...
https://github.com/clojure/clojure/tree/master/src/jvm/cloju...
I think translating the rest is a slow ongoing project (as performance, etc. gets up to par.)
http://www.zazzle.com/read_the_source_tshirt-235957482677132...
http://www.zazzle.com/read_the_source_tshirt-235248361224605...
This one should do the trick.
Unless it's minified, in which case you might as well be running cat on a binary.
Seriously though, the article does seem to belittle documentation. Documentation gives an API context, it is often invaluable. Even if you have the source with some comments, this often does not give a clear picture right away.
Sure beats trying to post an unformatable block of text into an MSDN forum and then well, nothing happens after that.
I also agree with the other comment here's implication, reading source code from others is a great way to get insight on how to use APIs, how to write idiomatic code and ways to avoid pitfalls.