What We Do and Don't Know about Software Development Effort Estimation
infoq.com
infoq.com
An explanation I like, from michaelochurch:
Let's say that you have 20 tasks. Each involves rolling a 10-sided die.
If it's a 1 through 8, wait that number of minutes. If it's a 9, wait 15
minutes. If it's a 10, wait an hour.
How long is this string of tasks going to take? Summing the median time
expectancy, we get a sum 110 minutes, because the median time for a task is
5.5 minutes. The actual expected time to completion is 222 minutes, with 5+
hours not being unreasonable if one rolls a lot of 9's and 10's.
This is an obvious example where summing the median expected time for the
tasks is ridiculous, but it's exactly what people do when they compute time
estimates, even though the reality on the field is that the time-cost
distribution has a lot more weight on the right. (That is, it's more common
for a "6-month" project to take 8 months than 4. In statistics-wonk terms,
the distribution is "log-normal".)No need to talk about distributions, gaussian or otherwise.
When I put together plans and estimates, I always take a lot of care to separate out those things which are linear and those things which exponentially impact other things within the schedule, along with the sort of inflection points. I may not know where I'm going to roll a 9 or 10, as they may crop up anywhere, but there are certainly areas where they are more possible and less possible.
In a sane world, at least. Can't do much with a black swan.
I think that's a better strategy than making wild guesses and ultimately falling behind schedule, but at the same time maintaining cadence, which buys you power to sometimes say "hard things are hard, I don't know when it'll be ready to ship".
Perhaps in a twisted future where we estimate project cost before deciding which projects to take on, we might discover our estimates are much better.
A related pathology is trading technical debt for speed, every time, on every project. The debt will be paid.
People have a recurring delusion that the world is shaped by their wants.
Part of the reason estimates are inaccurate is because there's that business disincentive to be accurate.
Hofstadter's Law: It always takes longer than you expect,
even when you take into account Hofstadter's Law.
The best advice I ever got on project-time estimation (from a biology postdoc) was: make your best, most honest best effort, and then double it.When I make projections with a spreadsheet, I have a cell that copies my grand total of all costs and call that copy "unforeseen costs". I always hate bidding that high at the start, but the estimate ends up being close to right surprisingly often.
This article says 30% overruns are common, which is within my former boss' +100% bounds.
The other nice thing about doubling your cost estimate is it prevents you from catching the winner's curse and landing an overly-stingy client. Plus if you really can keep costs within your spec for the project, then you win extra profits. You'll never win that "game" if you don't leave room for error.
http://alistair.cockburn.us/The+magic+of+pi+for+project+mana...
I think people should start using confidence intervals. Then the upper bounds become more realistic. If you need to estimate roughly with 90% condifence intervals, then a developer can communicate the uncertainity: I think task A will take about a week. At least 2 days. And no more than 3 months.
You can immediately see that it's probably best to either a) work with this tasks a few days and make a new estimate based on the acquired knowledge b) or if that's not possible, try to split the task to smaller subtasks to identify which parts are the most uncertain.
As others have pointed out in those large projects it's better not to make up-front estimates and just build as much value as possible for a fixed cost, using agile principles. However, that's typically not how large software projects are sold (or bought). Fixed price almost always means fixed scope. I'd like to know of any large software project sold to a customer in truly agile fashion (no fixed scope determined in advance). To me it sounds like a software development unicorn: you hear about it, but you're never the one building it.
In general when under-estimating the project you can make it:
1) on time and within planned resources,
2) with all the planned functionalities and
3) without sacrificing quality.
Pick any two.
Evidence suggests otherwise. Sure you can estimate +/- 100-200% early on but that isn't what anyone is aiming for in a software project. Even detailed plans of repeatable (non-trivial) software projects do not result error bars that anyone really desires.
The problem is that most companies don't record this data. Start today!
My point was that it is the planning step that is extremely difficult, not the estimating one. With most real-world projects the project plan must follow changing requirements (based on external input or on things you have learned during development). It is extremely unlikely that the original plan will (or should) be followed to the end.
"A tendency toward underestimation of effort is particularly present in price-competitive situations, such as bidding rounds. In less price-competitive contexts, such as inhouse software development, there are no such tendencies - in fact, you might even see the opposite. This suggests that a main reason for effort overruns is that clients tend to focus on low price when selecting software providers - that is, the project proposals that underestimate effort are more likely to be started. "
Is that really correct? Are there studies that shows that inhouse projects (or not fixed-price projects) do not underestimate systematically as opposed to fixed-price client projects?
http://meta.stackexchange.com/questions/19478/the-many-memes...
Edit:
One of the authors of the references has several articles available here:
https://www.simula.no/people/magnej/bibliography
but not the one referenced, although there are several newer articles.
4. T. Menzies and M. Shepperd, “Special Issue on Repeatable Results in Software Engineering Prediction,” Empirical Software Eng., vol. 17, no. 1, 2012, pp. 1–17
http://menzies.us/pdf/12stability.pdf
EDIT:
Several of the author's articles are available here (click the PDF links):
https://www.simula.no/people/magnej/bibliography?b_size:int=...
From looking at Google Scholar it looks like there are many newer articles on Software Estimation that the OP does not reference so may not have read.
Edit 2:
This paper talks about Monte Carlo Simulation:
https://www.simula.no/research/se/publications/Jorgensen.200...
Estimation cost (of doing the estimation, not consequences of estimation) not mentioned. Is estimation itself significantly costly relative to subject of estimation?
To the extent open source works relatively well as a development practice, how much of a role does suppression of estimation play (assuming there is suppression; harder to even pretend to hold anyone to an estimate without a contract, so why bother)?
It is feasible for the team to claim that it met the estimate, and it is feasible to have all indicators green on the day the deadline is met. Simply do less design, less refactoring, less thinking, less tests, less collarborative work, less engineering...
There isn't much mention of the estimates you can go for:
1) Accurate but not reliable
2) Reliable but not accurate
I absolutely loved this line.
When a contractor gives you an estimate of how long and how much it's going to take to do a remodel, he's invariably on time and on schedule, right?
And when Boeing spends billions of dollars on a new plane, they have on ready and on budget, right?
So why can't software people do the same?
Oh wait, complex, badly defined projects tend to run late and over budget. It's not that complicated. Spend 3-6 months defining all the details of your new web application, promise not to change anything on the fly, don't ask us to make it work on IE 7, and by the time we do 3-4 of these, we'll be able to give you a good estimate.
Well, no, they don't.