Another view of this would be: progress in this industry is made by those who succeed while sticking to their guns and not optimizing their infrastructure for the available worker pool.
Another view of this would be: progress in this industry is made by those who succeed while sticking to their guns and not optimizing their infrastructure for the available worker pool.
I have a strong suspicion that people smart enough to do advanced meta programming on their own are not smart enough to collaborate with equally skilled coworkers on code that uses all those abstractions to the max. Even collaborating with your own former self is more difficult that writing new code.
So I think _some_ dumbing down is inevitable. The question is how that can be enforced. Choosing a dumber language to bludgeon everyone into submission is the nuclear option so to speak.
I often think of this Dijkstra quote about abstraction: "The purpose of abstraction is not to be vague, but to create a new semantic level in which one can be absolutely precise".
But I think there's something missing in this thought. Abstraction, especially when its precise, is also a form of compression. To understand what a particular abstract expression means in a concrete case requires decompression, which may require a lot of mental effort.
These are admittedly extremely half-baked thoughts...
[1]: Once a certain point has been reached; theory does predict diminishing returns.
Do you mind sharing a source for this? It seems trivial to generate a programming language (Brainf*ck) that performs very poorly.
A more likely statement: all commonly used languages are more or less on par.
I have read quite a few studies, and I always come away thinking that the methods used are inadequate to support any conclusions at all.
But does that mean there is no causal effect of programming language choice? I don't know and I'm not making any claims either way.
I was talking about abstraction in one person's mind vs abstraction shared between a group a people. And what I said could actually be the reason why the choice of programming language has little effect (if that is the case)
I'm also not making definitive claims, but we can't ignore observations, either. While small effects can appear or disappear depending on methodology, big effects are easy to find and hard to hide, especially in an environment with strong selection pressures. Therefore, the most likely explanation to why no big effect has been detected -- either in studies or in industry -- is that there isn't one, and that should at least be enough to make it the working hypothesis. So before explaining effects that have yet to be detected we should establish that they exist at all, something we have so far failed to do.
There may well be a small effect, and there could still be a large effect, although the chance of that is small unless we have a really good explanation to why we haven't found it.
1) It is inconsistent with the personal experience of most developers. Few would claim that C is just as productive and safe as some garbage collected language (provided that both are suitable for the task at hand), or that a SQL group by statement is not more productive to write than the equivalent procedural code.
2) The methods used to study the subject are completely unfit for purpose.
3) It is inconsistent with the few results that seem at least somewhat well supported because they are simple enough to measure, such as the constant bug count per line of code.
I disagree. I've been programming for about 30 years now, and most reports I've received -- both from programmers and managers -- are that language doesn't matter; certainly not much. Carefully selected personal "feeling" report are worthless (especially in biased forums), as that would produce even more evidence that homeopathy is effective than that programming languages make a difference.
> Few would claim that C is just as productive and safe as some garbage collected language
I agree, but that's where the many caveats come in. First, my argument is more one of diminishing returns (following Brooks's theory). I.e., languages make an increasingly small difference. C may have been a big improvement over assemby, and Java may have been a smaller but still quite big improvement over C, but now differences are quite small.
Indeed, the effect of GC and memory safety vs C was one that was almost immediately detected by industry (once mature, performant, etc.) and triggered a huge shift. We do not see such shifts now. Companies move among mainstream languages, among non-mainstream languages, and between the two classes in a way that does not suggest a big effect at all.
> or that a SQL group by statement is not more productive to write than the equivalent procedural code.
That's a completely different matter. I don't think anyone believes that writing a 1MLOC program in SQL is easier than writing a similar one in Java or Python, even if the Python/Java one is 1MLOC, and the SQL one is 200KLOC. When the program is small, there can be large differences, but that's a completely different problem (and also explained by theory).
> The methods used to study the subject are completely unfit for purpose.
I disagree, but it doesn't matter. Big effects are easy to find and hard to hide. And in an environment with strong selective pressures, they're found even when no studies at all are conducted.
> It is inconsistent with the few results that seem at least somewhat well supported because they are simple enough to measure, such as the constant bug count per line of code.
It is not. First, the difference in LOC is not as big as PL fans claim (it can be big for small programs, not large ones). Second, you may want to review those results.
I don't necessarily disagree, but I'm a bit skeptical about the way in which you use the concept of diminishing returns. If I'm not mistaken then Brooks used diminishing returns in relation to headcount. And this is in fact how economists use the term as well - adding quantitatively more of one production factor.
But that doesn't apply to qualitative changes like introducing new production methods and tools, which is what programming languages are. New programming languages are not simply "more programming language".
If we agree that there have been observable effects of new programming languages in the past, then we cannot exclude the possibility of the same happening in the future.
>First, the difference in LOC is not as big as PL fans claim
What PL fans claim is a bit of a fluffy benchmark, but there are undoubtedly significant differences in LOC, certainly significant enough to make an economic difference.
>I've been programming for about 30 years now...
Shockingly, me too :)
Oh, I'm referring to "No Silver Bullet" and diminishing returns in the sense that languages at best help with "accidental complexity", so the less of it there is, the less languages can help.
> New programming languages are not simply "more programming language".
But there are inflexible theoretical limitations on their utility (even without Brooks). We know the effect of different programming constructs on expressive compression and reasoning costs. The closer we get to the limit, the less benefit we can have.
> then we cannot exclude the possibility of the same happening in the future.
I'm mostly pointing out that it's not happening at present. I'm less sanguine about future advances because of various limitations (and, BTW, Brooks's predictions were called overly pessimistic by PL fans at the time and it turned out they were too optimistic), but I won't rule out another major breakthrough, or possibly two.
> but there are undoubtedly significant differences in LOC, certainly significant enough to make an economic difference.
I don't think I agree with your first assertion (although that depends on what we mean by "significant" here), and I certainly disagree with your second. We simply have not been able to detect or induce such an effect. Here, too, there are theoretical results showing that increased expressiveness cannot lead to cheaper reasoning, which is not an obvious result even in the worst case (because there are far fewer "compressible" programs than non compressible ones).
>Here, too, there are theoretical results showing that increased expressiveness cannot lead to cheaper reasoning, which is not an obvious result even in the worst case (because there are far fewer "compressible" programs than non compressible ones).
I find that very interesting. Do you have a source for it?
* http://www.lsv.fr/Publis/PAPERS/PDF/Sch-aiml02.pdf
* http://www.lsv.fr/Publis/PAPERS/PDF/DLS-jcss-param.pdf
* https://www.cis.upenn.edu/~alur/Zohar03.pdf
What's important to put those results in context for those who are not familiar with the subject is that, in the context of complexity theory, the "model checking problem" does not refer to the complexity of a particular model checker algorithm, but to the inherent complexity of the problem of deciding whether a program M satisfies some property đťś‘ (i.e. whether M is a model of đťś‘ in the formal logic sense). The "model checking problem" is the closest to a mathematical description of what we usually mean when we talk of "reasoning about a program."
I've summarized these results and more here: https://pron.github.io/posts/correctness-and-complexity
Consider JVM and Java: that you can write Java seemingly without any knowledge about JVM is an illusion - you already learned a lot about how things like JVM work under the hood when first learning programming[0]. Learning more about JVM itself lets you write more efficient code and debug better.
Consider functions: it's near-impossible to contain all the information necessary to use a non-trivial function in its header[1]. That's why good code contains comments and other forms of documentation describing the abstraction in more details. Even then, it's not always enough - sometimes it's really much easier to understand what you need by reading the source[2].
Consider any appliance - be it a car, or a radio, or a dishwasher. You can use it to it interface, to some extent at least. But knowing what's going on under the literal hood really does help with use, and especially helps when something goes wrong. If you don't know what's hidden under the abstraction layer, any failure will likely be incomprehensible to you and leave you helpless.
--
[0] - In theory, it might be possible to learn some programming without learning anything about hardware or hardware-emulating abstractions. In reality, I've never seen it or heard of it, and I suspect that the simplest mental model of code execution is isomorphic to somewhat simplified computer or virtual machine.
[1] - Function name + name of its arguments + types of its arguments and return value, if available.
[2] - Then again, implementation code may not capture the entire abstraction either. That's why good code often features comments inside the implementation, explaining the rationale behind some of the less obvious code parts.
I think there's a misunderstanding. What I'm calling decompression has nothing to do with digging into the implementation of a particular abstraction. What I mean is merely tracking down all the indirections in the public interface and breaking them down to their concrete meaning in a specific case.
Some languages allow for a lot more moving parts than others. For instance, in Java obj.otherName means that otherName is an instance variable (or theoretically a class variable) and access time is guaranteed to be constant (tbd: caveats). It's not redefinable as it is in C#, Python or Swift. So there's one less thing to track down but also less flexibility.
Or take an extreme example like Go's range loop. It's defined only for a handful of builtin data structures. When you see a range loop, you know what it does without following any further indirections. It can be extremely annoying and create a lot of friction when you have to work with custom data structures. But it can make other people's code easier to read than in other languages.
Or take something like this:
a == b
This simple expression has a far greater number of possible meanings in a language that supports operator overloading and generics than in a language that doesn't.So transforming an expression's possible meanings into its concrete meaning is what I call decompression. There may not be a need for looking anything other than public interfaces.
Also, you're talking about abstractions purely from the perspective of users. But we are often creators and maintainers of abstractions as well when we model a particular problem.
You might be interested in this book, Patterns of Software, which brings up the issue that inheritance in OOP isn't about reuse but about compression and how the implications of that are why it's not had that much success in following through on the "reuse" promise. http://dreamsongs.com/Files/PatternsOfSoftware.pdf
I'm an unabashed Clojure fanboy and will talk the language up to anyone who will listen, but one thing that I've been thinking about lately is that good judgment on abstraction use is necessary for using such powerful tools. I don't think the place to tackle it is by being "smart enough to understand," but rather by being "disciplined enough to not abstract sometimes." The exciting thing is that discipline doesn't require you to be super smart, just to think it's worth caring about.
Also half-baked, but I think it's something.
It's not that much harder to enforce than it is to enforce conventions in less powerful languages, and the convenience of working in a well-designed language is so nice.
There's no reality here, just myth. That switching to Java would drive quality up makes at least as much sense as it driving it down.
If you're referring to the particular languages, then the reality is that no big effect has been found for the choice of programming language on software quality. As lack of evidence in favor of a big effect is evidence of its lack, the most probable explanation is that no such effect exists (it's possible a small effect exists, but we don't know in which language's favor). And, indeed, while there is a theory explaining (and even predicting ahead of time) why there is no big effect to the choice of the programming language, not only is there no empirical evidence suggesting there is such an effect, there is no theory to predict it, either. That language has a big effect on quality is, at this point, no more than wishful thinking among a minority of developers who are big programming language fans. It's a bedtime story. (I am not saying that language didn't ever have a big effect or that it never will, just that both theory and observation show that there isn't one currently among "reasonable" languages in common use; also, I have used and I like both Java and Clojure)
And if you're only referring to the priority of making the project more accessible, then I don't understand your comment at all. There are many, many more experienced Java developers than experienced Clojure developers, and experienced developers are at least as likely not to bother learning a new language in order to contribute to a project (although perhaps for different reasons). Picking a popular language makes the project more accessible to experienced developers.
What has been shown to drive quality up or down is process. If you have a good process, then making the project more accessible can help, and if you don't then you're screwed anyway.
Do you mean "defect rate" or something else?
Substantiate that. I say it's garbage.
<https://www.cse.iitk.ac.in/users/karkare/courses/2010/cs653/..., and this BTW is from 1994 so this evidence has been around for ages
Language LOC Documentation Lines Dev time (hours)
(1)Haskell 85 465 10
(2)Ada 767 714 23
(3)Ada9X 800 200 28
(4)C++ 1105 130 –
(5)Awk/Nawk 250 150 –
(6)Rapide 157 0 54
(7)Griffin 251 0 34
(8)Proteus 293 79 26
(9)RelationalLisp 274 12 3
(10)Haskell 156 112 8
(Edit: all aligned with tabs but looks like they've been stripped. See paper instead)comments of note:
(2) The Ada solution was written by a lead programmer at NSWC; in this sense it represents the “control” group. The developer initially reported a line count of only 249; this was the numberof lines of imperativestatements as reported by theSun Adacompiler, anddid notinclude declarations, all of which were essential for proper execution. The line count of 767 is based on the actual code in and does not include lines with only termination characters.
(4) TheC++solutionwaswrittenbyanONRprogrammanagerafterhavingfirstwrittentheAwk solution described in the next paragraph. In addition to these 1105 lines of code, the developer also wrote a 595-line “test harness.” No development times were reported [Note: C++ has evolved a lot since then, but perhaps so has haskell]
(10) Intermetrics, independently and without the knowledge of NSWC or Yale University, con-ductedan experimentof its own: theHaskell Report was given to anewly hiredcollege graduate, who was then given 8 days to learn Haskell. This new hire received no formal training, but was allowed to ask an experienced Haskell programmer questions as issues came up during the self study. After this training period, the new hire was handed the geo-server specification and asked to write aprototype in Haskell. Theresulting metrics shown in row 10 ofthe tableare perhaps the most stunning of the lot, suggesting the ease with which Haskell may be learned and effectively utilized.
Also mentioned was "The use of higher-order functions [in haskell] is noteworthy" so language features did help. Actual evidence.
> As lack of evidence in favor of a big effect is evidence of its lack
yeah right. Let me rephrase that for you - Lack of knowledge by an HN poster is not evidence of lack of evidence.
> What has been shown to drive quality up or down is process
Evidence please? My experience is that process can be rigorous but useless crap. My old boss said "process is not sufficient to produce quality, but it is a necessary"
There's plenty more in that PDF, please read it.
Evidence of what? This is not a study of software development at all, but of prototyping using a program with no more than a few hundred lines. I'm talking about effects on software development. The issue of prototyping/tiny programs is a completely separate one.
> Substantiate that
I don't know how I can substantiate the lack of evidence, but here's a recent failed attempt to find evidence: https://arxiv.org/pdf/1901.10220.pdf
> Lack of knowledge by an HN poster is not evidence of lack of evidence.
True. While I have been following the subject for at least the past 20 years (and I have read your PDF several times already, including shortly after it was published), it is certainly possible I have missed something. If there's any evidence of a large effect you believe you've found, let me know.
> Evidence please?
Sure. Here's code review, for example (I don't have time to discuss each paper -- as they're quite different -- but they all paint a similar picture). BTW, while programming language studies are debating effects in the 0-15% range, these report effects in the 30-80%:
* What We Have Learned About Fighting Defects, 2002 -- https://www.cs.umd.edu/~mvz/pub/eworkshop02.pdf
* The Impact of Design and Code Reviews on Software Quality: An Empirical Study Based on PSP Data, 2009 -- https://www.pitt.edu/~ckemerer/PSP_Data.pdf
* Large-Scale Analysis of Modern Code Review Practices and Software Security in Open Source Software, 2017 -- https://www2.eecs.berkeley.edu/Pubs/TechRpts/2017/EECS-2017-...
* What Types of Defects Are Really Discovered in Code Reviews?, 2003 -- https://ieeexplore.ieee.org/document/4604671
* Modern Code Reviews in Open-Source Projects: Which Problems Do They Fix?, 2014 -- https://www.testroots.org/assets/papers/2014_beller_bacchell...
* The impact of code review coverage and code review participation on software quality: a case study of the qt, VTK, and ITK projects, 2014 -- https://dl.acm.org/citation.cfm?id=2597076
* Best Kept Secrets of Peer Code Review, 2013 -- https://smartbear.com/SmartBear/media/pdfs/Best-Kept-Secrets...
* Modern Code Review: A Case Study at Google, 2018 -- https://storage.googleapis.com/pub-tools-public-publication-...
If we summarize our current state of knowledge (not myth) it is this: process matters a lot; language matters little, if at all (with all the caveats I mentioned in other comments, such as diminishing returns etc.).
> then the reality is that no big effect has been found for the choice of programming language on software quality
Then I show some evidence giving small programs rapidly developed, which you then say doesn't count because it's 'tiny' and 'prototyping'. You also don't define quality which gives you lots of wriggle room.
Nope, 1000 line programs aren't tiny. They aren't industry monsters but you can't dismiss them because it contradicts you. It's not proof, but it is strong evidence.
...arxiv paper... Interesting, thanks. It would have been helpful to have posted that in your original post.
However from <https://soarsmu.github.io/papers/A_Large_Scale_Study_of_Mult..., and note this paper is not mentioned, as in not debunked, in the above arxiv paper.
"Bhattacharya et al. study four open-source projects which use C and C++, i.e., Firefox, Blender, VLC Media Player and MySQL to understand the impact of languages on software quality [20]. They compute several statistical measures while controlling for factors, such as developer competence and software process. They find that applications previously written in C are migrating to C++ and C++ code is often of higher quality, less prone to bugs, and easier to maintain than C code."
and from that same paper:
"As can be seen in Tables 8, 9, and 10, the mean de- fect density values for the C sets can be up to an order of magnitude higher than the mean values for the C++ sets."
> Sure. Here's code review...
I was talking about your claim about languages not mattering, I did not mention code review. All these papers are about code review. This is relevant to my code review comment (edit: I meant process comment), they don't say anything about language vs bugs (unless you wish to point out a paper that does).
> process matters a lot
And I agreed with you, I just said it didn't deliver quality, it just prepared the ground for it. With enough process you can close any hole in a language, but it gets exponentially expensive.
Only if you insist on uncharitable reading. I'm talking about software development; if I mistakenly assumed that could be left implicit, I'm sorry. And ~500 line programs are positively minuscule. JQuery is ~50KLOC, and an average business system is ~5MLOC. We're talking four orders of magnitude between those prototypes and a rather average industry system size (systems commonly run to tens and even more than 100 MLOC).
> It's not proof, but it is strong evidence.
I don't think it matters if it's strong or weak evidence, as it's not even about software development.
> the C sets can be up to an order of magnitude higher than the mean values for the C++ sets.
As I mentioned in another comment, C is not a good example. It's a ~50-year-old language, and the theory that correctly predicted that languages won't make a big difference was based on diminishing returns. I.e. not that no two languages have ever been or could ever have a big difference, but that over time the ability to affect quality with language would severely diminish. The original prediction of that theory was that no 10x improvement would be made by a single language improvement over a decade, and was called overly pessimistic by PL fans. It's not been over thirty years, and we have doubtfully made a 3x boost with all language features combined.
If you want a more precise statement: no theory or empirical evidence supports the claim that a reasonable choice among production languages developed in the past three decades or so has a big impact on any measurable bottom-line metric.
> I just said it didn't deliver quality
But those papers show that, unlike language, process does have a big impact on quality.
> With enough process you can close any hole in a language, but it gets exponentially expensive.
You are now making an unsubstantiated claim that contradicts a substantiated claim. Programming languages (with the caveats above) have not been found to have a big effect (nor is there a theory that suggests they do), while process has.
And the value different languages may add is entirely about software development.
> I don't think it matters if it's strong or weak evidence, as it's not even about software development.
Programming languages are not about software development? Oh do go on.
> As I mentioned in another comment, C is not a good example.
A study of C vs C++ strongly indicates your claim about languages is false and suddenly C is not a good example? I'm trying to believe I'm just misunderstanding you but it's getting more difficult.
> The original prediction of that theory was that no 10x improvement would be made by a single language improvement
This isn't what you said. Let me remind you "If you're referring to the particular languages, then the reality is that no big effect has been found for the choice of programming language on software quality"
The goalposts aren't being moved, you've just bought them plane tickets to barbados.
> If you want a more precise statement: no theory or empirical evidence supports the claim that a reasonable choice among production languages developed in the past three decades or so has a big impact on any measurable bottom-line metric
Another claim. Show me the study that says that.
> But those papers show that, unlike language, process does have a big impact on quality.
It can if done properly; it does not do so automatically. You can have code reviews that are of little use because you have no spec to review the code against. I have been in that very position. I don't dispute the value of process but it's a necessary but not sufficient condition to bring about quality. I also agree code reviews are good, if done properly.
ME >> With enough process you can close any hole in a language, but it gets exponentially expensive.
YOU > You are now making an unsubstantiated claim that contradicts a substantiated claim.
If I use C I have to worry about garbage leaks, double-freeing pointers, out-of-bounds accesses etc that can be detected with code reviews. If I use python, I never have any of these problems so a code review need not check for these. That's a lot cheaper.
(Edit: weird stuff happening to my post, may appear twice)
I can't repeat all the caveats every time I say something. I assume readers believe my comments are at least reasonable, even if they disagree with them. For convenience, I restated the current state of knowledge more precisely in my previous comment. If you're interested in what I have to say, assume I'm not an idiot, and perhaps we could have an interesting discussion, and if you think I'm an idiot, then there's no point in arguing at all.
> Show me the study that says that.
Show me the study that says otherwise. I don't need to provide evidence for the lack of an effect (although I did). The lack of an effect is, at the very least, the default hypothesis in this case (as in many others) as there is no theory suggesting we should believe otherwise (while a theory that explains why languages increasingly have smaller effects has made correct predictions).
> It can if done properly
That's your hypothesis. We've been trying to find evidence of that, or others like it, for a long time -- both in academia and in industry -- without success. We simply do not observe that the choice of language today (among reasonable ones etc.) has a big effect, either in studies or in industry practice. But if you've found a language that can drastically reduce software development costs and/or increase quality at scale, that discovery can be easily translated to billions of dollars. Go ahead and make them.
"In terms of programming-in-the-large, at Google and elsewhere, I think that language choice is not as important as all the other choices: if you have the right overall architecture, the right team of programmers, the right development process that allows for rapid development with continuous improvement, then many languages will work for you; if you don't have those things you're in trouble regardless of your language choice."
Could there be other reasons? Maybe they rushed?
This is in the context of software development in general, not this specific project per se.