In practice I find that static typing as done by, say, Java and C#, excels at uncovering very shallow bugs. Its benefit is the most obvious when you do very little automated testing. The tech industry's dirty little secret is that automated testing is done infrequently and badly, so it looks pretty effective.
I find that if you try to compare the cost of building sophisticated tests in a high productivity language (like python) with the cost of writing code in a low productivity/high type safety language like haskell, there's a pretty blurry cost/benefit trade off. If you take out productivity as a factor, haskell looks much better.
Nonetheless, the idea that static types (and, to a lesser extent, formal proofs) are some sort of silver bullet persists. There is no silver bullet.
A good counterpoint is Go. Despite being far more familiar with Python (it’s my professional/work language), it’s static types actually help me write correct code more quickly than I can with Python. Your mileage may vary, but I think anyone who has given Go an honest shake will find that it’s at least in the same productivity ballpark as Python. Note that this isn’t considering tooling or deployment, or performance where Go specifically excels over Python (and also often due to static types).
Of course, with sufficient investment in testing, you could get the same confidence with Python that you get in static languages, but writing tests is a productivity cost as much as boilerplate or a pedantic compiler. And with static languages, I often find that I can write fewer tests for the same confidence (in fact, I often prototype in Go and backport to Python).
On that note, why would you consider Haskell a low productivity language? People who are well versed in it seem to find it an exceptionally productive language to work with.
Types also help document code. Just like tests.
The point is that both are investments and both have different payoff matrices. Sophisticated are often better at preventing obscure logical bugs. They're also good at uncovering obscure not-bugs and preventing code from getting out until the compiler is satisfied. Bad unit tests also do this.
Haskell simply takes longer to write than other languages. Given two developers of equal skill and experience, anyway. I partly attribute the relative paucity of haskell software out there to this.
Further, while I agree with your “different matrixes of payout”, I think static types are a very low investment and they have a respectable payout in terms of preventing bugs but also in terms of documenting code and facilitating tooling (such as autocomplete or documentation generation) and they also permit easier code changes than comprehensive unit testing (even “good” unit testing). Of course, I’m not advocating that static types completely obviate tests; rather that they obviate some of the tests; however, there is no clear answer as to how many or which tests are obviated; it’s very circumstantial.
> Skepticism is not pessimism, however. Although we see no startling breakthroughs, and indeed, believe such to be inconsistent with the nature of software, many encouraging innovations are under way. A disciplined, consistent effort to develop, propagate, and exploit them should indeed yield an order-of-magnitude improvement. There is no royal road, but there is a road.
Alternatively, usage of Rust may not be correlated with fewer bugs overall but I'd expect there to be fewer of the use-after-free variety.
I actually heard Aribus software is written in C++. If true it should scare the shit out of anyone.
If you don't believe the results of the study, please provide the counter evidence.
Besides, all I wanted to do was to point out that you shouldn’t derive any conclusions from that study. I don’t need any counter evidence for that (apart from the evidence that the study is flawed of course).
Instead, common sense is actually the best substitute we have.
Point is, code culture matters much more than language.
The advantage of this is that if you end up not wasting too much time "building the wrong thing". Let's say that you took one form of I/O and built massively strict validation in and then realized later that you should have taken an entirely different form of I/O for your subsystem. All that time building in validation on that useless part of code was a pointless waste.
I don't have any stats, but my gut feel is that on average 40% of code can end up being tossed in this way (in some projects it's 100% =).
Prototyping speed is, additionally, not just useful in reducing the cost of building the right kind of code, it's useful in reducing the cost of building the right kind of test (a really underappreciated facet of building mission critical systems).
In my younger years I used to believe that for mission critical systems "building the wrong thing" was somehow less of a problem in code because you could fix requirements and do architecture upfront with some sort of genius architect. Turns out this was wrong.
I'm curious as to where you have observed this, because my experience has been exactly the opposite: even in circumstances where "dynamic" might lead to more readable code, C# developers are loathe to use it, to the point where it's very hard to find it in idiomatic C# code.
I'm parsing your comment to mean, "it's easier to write correct code in compiled languages", but this is not obvious to me, or anyone who's, for instance, written any C at all.
* Ahead of time compiled languages reduce flexibility but allow the compiler to do more reasoning about the system.
* Statically typed languages allow the compiler to reason about the assignments you make and methods you call. If the language is also AOT compiled, you avoid crashing at runtime.
* Dynamically typed languages significantly reduce boilerplate and avoid the mental effort of expressing an idea in a rigorously typed manner. It's the "hold my beer" approach.
There's a usual confusion of banking in this case, but if we're talking about popular dynamic/memory-managed languages, you can put an equivalent of:
try:
actual code
except:
log.exception
At a high level and be fairly sure things will be ok in a long run even with failures along the way.On the other hand, most static-typed systems handle failures explicitly.
I think C is a bad example here. We've come a long way since C. We know how to do better.
Unless you want to include the ecosystem as well. C + valgrind + PVS + clang-analyser do make it easier to create correct code.
(1) Due to ecosystem richness, the need to collaborate cross-functionally, and other reasons.
EDIT: grammar gremlins.
Before test suites, I would have agreed with you.
The biggest value to me is that you can easily do huge refactorings with a lot of confidence. That in turn means you can keep redesigning your code and frameworks long after they would ossify into legacy code no one dares change in a normal project.
That's not to say I disagree that much with your point.
While static typing does catch code errors, in my experience they are mostly errors that would be caught somewhere else in the testing process. Memory safety issues are rarer, but also have a penchant for showing up unexpectedly and for crashing the entire system in unrecoverable ways.
Obviously, the best would be to have both kinds of safety, but people program safety critical systems in C++ all the time, and deal with the consequences. At least you can catch a TypeError in Python and handle it.
Catching TypeError in python and handling it is not a substitute for static typing.
https://stackoverflow.com/questions/3270680/how-does-python-...
"CPython implementation detail: Objects of different types except numbers are ordered by their type names; objects of the same types that don’t support proper comparison are ordered by their address."
Seems like a completely reasonable thing to do?
????????
None (as far as I remember) of the common "wat" examples about JS are interpreter implementation details, but specified behavior.
I'm mildly curious why CPython chooses to implement it this way, but if I had to guess: I'm assuming it is to provide a stable order between objects of different classes with the same hash() value. Hash-value being what I suspect is what the quoted answer misrepresents as "their address" (the address being in CPython the default fallback for objects that do not implement a hash() function), and being a good candidate to establish an arbitrary order.
I wouldn't say "100" < "2" really counts, since, alphabetically it makes sense (an equivalent would be "baa" < "c").
You realize that many mission-critical workloads run on Python with great success, right?
Or rather, how would you define mission critical?
If you write str1 == str2, it compiles but probably does not do what you want (object identity vs string comparison).
(edit: to clarify, this does not mean Java is not a good language overall, every language has its own set of shortcomings and as someone points out in the replies, this behavior is consistent with everything else in the language)
Frankly, if this is the most meaningful complaint you can come up with, it might be indicative that you haven’t learned enough about it to have an informed opinion.
As a non-Java developer I have to ask why you think this is a reasonable design choice. It's a massive footgun for the a casual polyglot Java developer. I can imagine just forgetting I can't do == if I've just switched from being immersed in another language.
Which is surely the important bit. I can just use "==" in C#.
In fact, from a certain perspective the two flavors get treated identically. Comparing two integers with ‘==‘ will be true if the contents of the variables are the same, and comparing two references will be true if the pointers they contain are the same value. It just so happens that for primitives identity and logical equality are the same thing.
Regardless, none of that changes the fact that it's a bad design decision, because people regularly get it wrong.
The problem is that in Java, there are no user-defined value types, only primitives; and primitives can't have methods. So if it were a primitive, you'd have to write "String.length(s)" etc. Also, all other Java primitives are basically bit sequences that are interpreted in one way or another, but that wouldn't be the case for strings.
.NET/C#, though, doesn't get that excuse. It totally could have defined String as a struct with an internal char[] field, and then there would be no question of value/reference equality for its ==. But that would also mean that you couldn't use null for strings - and they didn't have nullable value types back then. I suspect that, plus the overall mindset carried over from Java, is what won the day.
Side note: there's no hard and explicit distinction between reference and primitive types in C itself, nor in C++. If Python is "everything is an object reference", and Java is "everything is an object reference except for primitives that are values", then C++ is "everything is a value, including object references". Thus, there's no ambiguity with == in C++ - it always compares values, it's just that sometimes those values are pointers.
I think sometimes that perhaps the implicitness of object references that is so common to OO languages, that I think was introduced by Smalltalk, is a mistake. It's interesting that Simula-67 didn't have it - although it was very Java-like otherwise, having only primitive values and references to objects (i.e. no objects as values, like C++), it distinguished the two consistently by using distinct equality operators (= vs ==), and even distinct assignments (:= vs :-).
Or, alternatively, treat everything as an object reference, but make == do implicit dereferencing as well, as Python and VB do; and provide a completely distinct operator for reference equality, such as "is". Python has a problem in that regard in that it has a default implementation of == for all classes that compares references, and so it ends up used as reference equality in practice sometimes. It would be better to have no default for == at all, just as there's no default for other comparison operators; this is what VB does.
It might also be best to stop talking about value and reference types altogether - what's actually important is the presence or absence of identity. Then everything can be an object reference, but not all object references can be compared for equality (and in practice, under the hood, the implementation can just skip the references and copy the data itself - or not, depending on mutability and size).