Joe Armstrong on Programmer Productivity
groups.google.com
groups.google.com
Most time isn't spent programming anyway - programmer time is spent:
a) fixing broken stuff that should not be broken
b) trying to figure out what problem the customer actually wants solving
c) writing experimental code to test some idea
d) googling for some obscure fact that is needed to solve a) or b)
e) writing and testing production code
e) is actually pretty easy once a) - d) are fixed. But most measurements of productivity only measure
lines of code in e) and man hours.For me, b) is a bottleneck. c) is far and away my favorite part. Nothing like green-fielding . . .
c') tooling: learning/debugging the latest test tool, test runner, deploy tool, source management system, CI system, etc. etc.
You have to either:
a. stick to old and/or boring and/or obviously inferior tools and programming languages, or
b. expect to pay the price of the time lost by programmers joining your team to learn the tools that you have chose or "worse", to actually survey all the options and pick the best tools for the project
...I know, the evolution of technology needs diversity in everything, including tools, so that some form of "selection" can actually push forward the best alternatives, but just as in nature and biology, "evolution" has huge prices to pay (in biology the most tragic prices are mortality and cancer, in software engineering I don't know what their analogues are, but I dread to think that we're already close to facing these "prices"...).
I can think of at least two further difficulties that come with the explosion of available tools:
1. It’s hard to know whether a tool is actually worth using at all. Many tools sound appealing because you’re familiar with the problem they aim to solve but not yet familiar with all the quirks and edge cases and hidden costs of using the tool instead.
2. Many tools look promising in their early days, when they are prime examples of enthusiasm driven development. It’s another thing to ask whether the tool is still going to be in active development five years down the line, when there are newer and shinier things to work on. Even if there are, your 100,000 line code base might still need that security fix or one of the 10% of missing features that were on the roadmap when you first started integration.
I’m all for using the right tool for each job, and I’m certainly in favour of using good tools rather than doing everything the hard way, but figuring out which tools those are gets harder as tools proliferate.
"Experiments that show that Erlang is N times better than "something else" won't be believed if N is too high."
That is exactly it. I've had those arguments where I was forced to tell a manager about another project I did for another company, a similar project where things went quickly and smoothly, using technology the manager was not familiar with. And every word I spoke was met with disbelief.
Then, when your manager inevitably does not believe you, ask him if he thinks you missed something (no), then ask him if he trusts you (yes), then contemplate the face of cognitive dissonance.
Learning Cocoa or Cocoa Touch, otoh...
I'll have to check erlang out. I only hear good things.
A) Writing a "Hello, World"?
B) being able to fix a simple bug in a program?
C) Writing a small program that solves some domain problem
D) Writing a program that solves a "real world" problem, making use of the programming languages' strengths and conforming to standards (something that other people would let you commit into a repository unchallenged :) )
I'm certain A) takes only a few hours, and B) can be done in a day.
As Norvig says, you can do C) if you write in the new language like you would in a language you already know (not much of a problem if we're talking C# vs Java).
But D) IMO takes weeks/months.
Ignoring OTP, my ability to use the language proper was pretty much complete within a week, owing to its simplicity and my prior background in functional languages.
This is in contrast to C, which I have been using for around 15 years now, yet am still learning nuances of the language. (Things like: which integer operations are undefined on negative numbers; how integer promotion works with shift operators; how "restrict" interacts with scope.)
C is almost a fractal of nuance that does take years of experience to comprehend; Erlang has no nuance. (The closest thing to nuance I can think of is the relationship between integers and floats; and even then the takeaway is "it just works; don't worry about it". The only language in which I've seen the number hierarchy handled more cleanly is Racket.)
Norvig it covers this well in his Teach Yourself Programming in Ten Years (http://norvig.com/21-days.html) essay.
One problem I have, which Norvig doesn't talk about, is that I'm currently moving from one ecosystem (Microsoft) to another (the Java world), and it's not just about the language, it's everything, all the toolchain is different and unfamiliar.
The way of doing things is quite different too.
Coming from a strong background of similar languages? Sorry, you're wrong there. Perhaps you aren't familiar with Erlang; it is the "everyman" of functional languages. The things I think you'd call "idioms" (not really sure what you mean as that's a very loose term), let's say recursive programming, pattern matching, carry over as-is from any number of other functional languages. Pattern matching? It's a mix of OCaml and Prolog. Terms? Imagine Scheme had tuples in addition to lists. Etc. The only "new" concept is the message system, which if you know anything about networking, is pretty straightforward.
In other words, if you know pretty much any other functional language, it's almost trivial to translate basic programs to & from into Erlang knowing little more than its syntax, which as you concede, is simple.
OTOH, if you're coming from, say, PHP, sure, learning Erlang will be difficult as first you have to understand functional programming. But you're asserting impossibility, so counter-examples are moot.
Ten Years… I've been coding over twenty years in more languages than I can count… I have a good idea how long it takes to learn a language well. Maybe I'm biased because it's the latest language I've studied, but Erlang was the first non-toy language where I skimmed the reference manual, said "huh, that was all unsurprising", and started coding effectively.
-- Neal Stephenson, Zodiac (1988) http://en.wikiquote.org/wiki/Neal_Stephenson#Zodiac_.281988....
(I'd include the previous paragraphs in the quote, to show the conversation the protagonist actually has about this, if my copy wasn't on loan to a friend right now, dammit! If anyone else can post that, that's be cool.)
"We integrated with Facebook a few months ago, it should be easy to do it again."
Famous last words. I've lost count of how many times I've become an "expert" at some aspect of Facebook integration to have it be completely different just a few months later. Google is also really bad about this, at least in the past year or so. It almost makes me want to quit webdev and be a dba or something.
DBA it is then.
I would like to explore to find a sane and reasonable approach and not be driven by artificial deadlines guessed at three months ago but by todays business need.
i would like to actually be measured on business value generated, not lines of code written.
I would like to improve and simplify, deliver real value and savour the joy of actually creating.
this may of course explain why I feel totally unproductive at times
Slow food isn't about slowness for the sake of slowness. It's just about being realistic and interested in the long-term. It's about saving energy and mental stress.
One beautiful expression of this ideal is in "Domain-Driven Design."
had a rant recently on
able to move quickly and easily; able to think and understand quickly
Agile should have been called Adaptable or some other name that has less implication of speed and more implication of accomodation of unknown or changing requirements.
No putting that toothpaste back in the tube at this point though.
Seems to be paying off quite well. What other database can ship reliably every year with major new features and all kinds of improvements on many axes? And people still say the code is readable and the reliability is solid.
People should try and find their "natural speed" when coding, and figure out other personal special abilities they have instead of "speed", stead of trying to keep up with the fastest guy in the room and being afraid to give estimates that are x times longer than his, and managers should accept that people work and think at different speeds and "slow" != "stupid", as most Americans tend to use "slow" in everyday lingo.
I have seen so many furious sprints towards the 'finished' system with so little knowledge of what actually needs to be built. The only part that has changed recently is that there is now typically a unit test suite along side that correctly tests the wrong thing.
http://www.metabrew.com/article/rewriting-playdar-c-to-erlan...
I expect the 'N' value varies wildly depending on what you are building, and whichever language you are comparing against.
Changing the language changes your thinking about the problem. If you translate a program from erlang to C++, the result might be shorter and more reliable than if the first version is in C++ and you rewrite it in C++. It also might not be, of course, but the fact that it could be muddles the comparison.
I worked with someone who produced very elaborate designs for the most trivial tasks. Unfortunately, his code was more clever than he was, so it often didn't actually work. Whenever I re-wrote something he'd written in the same language, I managed to do it in 1/10th to 1/15th the number of lines, while adding the "actually works" feature.
1. memory management is hard - this is mostly a solved problem because the difference between C and $GARBAGECOLLECTEDLANG was great enough that it outweighed most differences in programmer ability. The mass move to web servers put paid to the need to develop for Windows APIs and Linux Syscalls and the a erage software project actually got better. (well ... worked)
1.a. business rules work better in languages that treat functions as first class objects and even better in languages that are logic based and really well in DSLs. most programmers don't reach the first so ... well we await the next mass change if languages.
2. Choice of algorithm. For loops are nice. just not everywhere. O(n^2) hurts. You can see the same in Unindexed table scanning queries.
3. Metrics. measure your own programs. write dynamic do s that update profile information in them. if you are not measuring then ... how do you know you improved? passing another hundred tests does not tell you if those tests measure what is valuable. tests are regression. metrics are progression.
4. I'm ranting too much today, blood pressure is rising :-)
I didn't do a full C++ rewrite of course. I think it's fair to say that whenever you are doing lots of concurrency/multithreading/evented stuff in a language like C++, the savings will be enormous when you switch to a language that has a nice concurrency model (ie, not just threads and shared mem and mutexes).
I expect if I rewrote either codebase again, it might shrink even more.
So this is similar to benchmarks: if you don't write a long paragraph about methodology, SLOCs and benchmarks are misleading
But what if we actually built things that we already agree in advance to throw away? Maybe even start from two or three different plausible designs, and push on each one for a while until one seems to be winning? Then rewrite it, learning from the other contenders and from the prototype itself, all before shipping.
The obvious answer is that it would take too long. But I'm not so sure. They say designing software takes too long, also, but when I spend a few weeks designing, the implementation ends up going smoothly and hitting the target. Usually the features that are designed thoroughly at the start of the release hit the target, and other "quick" features that are added in later, sans design, take longer than the designed features and end up pushing the release out. Any overruns or missteps in the implementation of the well-designed features seem insignificant in comparison to things that skimped on design work.
Of course, what's done in practice tends to differ, unfortunately.
Exactly... I've never really seen this happen beyond what I would consider "experimental hacking". Maybe it happens somewhere.
The key is to not be ashamed of it.
http://www.cs.umd.edu/class/spring2003/cmsc838p/Process/wate...
In the late 80's and early 90's, rapid prototyping was very fashionable in the 'anti-waterfall' software engineering schools of thought. The idea was to use dynamic languages with good development tools (various Lisps, Tcl/Tk, Smalltalk) to rough out the idea, which would then be rewritten in 'real' languages.
Of course this tactic failed successfully! We shipped the prototype. Nowadays, it would seem absurd to develop a web application (for example) written in a 'real' language like C++ or PL/1.
The trouble at the moment is convincing the management that a) this is productive work and b) it really is time to (incrementally) rewrite this code - it's beyond saving. From their point of view it looks like we want to scrap something that sort-of works most of time and do all the work from scratch. It's difficult to convey just how much psychic damage the current code is causing to someone who isn't buried in it day-to-day.
That makes me think of a more modest way to do these experiments that might return more reliable results: use the same team twice, but have them solve the problem in their favorite language first. That is, if A is the pet language and you want to compare A to B, write the program first in A and then in B. This biases the test in B's favour, because A will get penalized for all the time it took to learn about the problem while B will get all that benefit for free. Since there's already a major bias in favor of A, this levels the playing field some.
Here's why I think this might be more reliable. If you run the experiment this way and A comes out much better, you now have an answer to the charge that the second time was easier: all that benefit went to B and B still lost. Conversely, if A doesn't come out much better, you now have evidence that the language effect isn't so great once you account for the learning effect.
Also, whichever of A or B goes third would still have an advantage over the one that went second.
(I say "optimal" rather than "shortest" so we're not tempted to sacrifice clarity for concision.)
The only downside, and I think adding the language C tries to defuse it, is that whatever the team writes first will always stick. For example: if I wrote a version with dynamic typing (say in Ruby), then redid it with static typing (Haskell), of course I'm going to try to reuse types. The extreme example (Greenspun's rule) is if my pet language A is Lisp: regardless of what B is, the team could try to write a half-baked lisp runtime on top of B. The style carries over, and sometimes it doesn't translate exactly. I don't know how to solve this.
We're standing on the shoulders of giants. Underlying most of our environments is Unix (or Linux, same diff), which is basically the same as it was 20, 30 years ago. We're also running on http, an astoundingly good design. There are lots of other basics, none of which are all that complex, and all of which are finely tuned.
More to the point, a) represents a practical limit. If our environments get too flaky due to poorly understood configuration or bugs, we ultimately can't get programming done. But not pushing to where it hurts some means not using the latest, greatest tools that can amplify our power as programmers - the same tools that let us have thousands of times more software than we used to, without losing any more time to configuration/bugs than we did decades ago.
And beyond that, a lot of the business opportunity in the industry lies with running along just behind the bleeding edge, being firstest-with-mostest to the new and powerful technologies. So unless you're in a safe business relatively independent of new tech, you're going to bleed a bit.
He said thousands of times more lines in each piece of software, not thousands of times more software.
>That's because today's software is significantly less broken than the software of yore.
As a per-line measurement, yes. As a per-program measurement, not even close. Nor as a per-feature measurement, because those 1000x locs are largely going into abstraction layers at the bottom of the stack and chrome at the top.
Pardon the intermission, but this is one of the areas where node.js shines IMO. I can read the complete source code for a very complex application, or at least know that each module has a reasonably-sized source and is readable if the need arises. Very small modules and using composition is encouraged, not restricted to a fringe community, and you also get to share tooling and libraries with the browser. All that on top of a friendly, functional language. It really feels like a step forward.
Now if only npm could shadow repositories...
require('my-dependency')
returns a value, rather than aliasing identifiers into the current scope. So you end up with code like: var myDep = require('my-dependency')
exports.doSomething = function (x) {
myDep.doSomethingToAnX(x);
// more code here
}
The Python import (and Go and probably others) works in a similar way, except that in node, even the module names (the argument to require) are local to each module. That is: `require('my-dependency')` can return different values depending on the location of the file that called it. This means your project can depend on two different libraries that both depend on conflicting versions of "my-dependency". The mechanism by which this is accomplished is simple and straightforward. Another nice side-effect (as compared to Ruby) is that you can easily isolate a copy of any given dependency to monkey-patch or otherwise modify it, without affecting anybody else. (Not true of built-in globals such as String or Function, but that's a shortcoming of JavaScript, not Node)Of course, the event+callback architecture and preference for writing everything in a manual continuation-passing-style makes working in Node suck for a host of other reasons, but the module system is really quite nice.
If someone wants not yet released features, they can compile to JS/node with source maps (i.e. Traceur, Coffeescript). This also solves the cps/callback issue (i.e. Iced Coffeescript).
Npm is also incredibly easy to publish to (one line in the CLI). That leads to a ton of packages released that would otherwise hit friction in release process in other package managers.
The problem is we don't do similar things over and over again. Each new unsolved
problem is precisely that, a new unsolved problem.
And if tasks are similar, and they can be made mechanical (or at least, commonalities can be factored out), it should be turned over to the machine itself.Have the same feeling.
If you made me write with Eclipse, I'd probably be horribly inefficient. It's not that Eclipse is a bad tool, it's just that it lends itself to languages that I'm not super fond of writing in (ie. Java, ActionScript, etc.) and it's much different than the way I'm used to working (with vi, git and a command line). If you transplant a C# programmer into the vi and Python world, they'll usually run away screaming. Neither is better, they're just different.