Why I don’t like Dynamic Typing
win-vector.com
win-vector.com
But, in practice, some dynamic languages just don't seem to suffer from hidden type bugs. One of those is Clojure. Instead of packing everything into complicated types, you are mostly working with maps and sequences of maps--just raw data--with higher order functions. There isn't a lot of room for type bugs. Additionally, without any input from the developer, Clojure does some type-inference all on its own. You can enable warn-on-reflection such that the Clojure compiler will tell you anywhere in the code where it's going to have to use reflection to figure out how to invoke a Java method or property on an object. You can fix them with type hints such that you get performance on par with statically typed Java.
Clojure protects you from name typos you might see in other dynamic languages and accidentally clobbering existing variables through the fact that most variables are really not variable at all. You bind a value once in a lexical scope and it can't be mutated. That's more of a benefit of functional languages and lexically scoped languages, than statically typed ones.
It may boil down to a matter of taste, but I don't see the need for fancy refactoring tools like "method extraction" in a language like Clojure. The methods already don't live in classes, so they hardly need to be extracted from them :). If you rename a method in Clojure, the compiler _will_ complain in all the places where you are trying to call a method that no longer exists.
EDIT: actually, the one thing I miss most from statically typed languages is the immediate error highliting in the IDE. The REPL alleviates the need for that, but it was still cool immediately seeing files and lines highlighted in red as soon as I type in some error. I wish Clojure could do that.
Maybe, but they would be imprecise and they would require developer supervision. When the tool doesn't have any types, it can't reason about the code accurately.
Here is an article that explains why in more depth:
http://beust.com/weblog/2006/10/01/dynamic-language-refactor...
I'm also skeptical of your claim about Clojure symbols: can you not select a field using a runtime-computed values, thereby making perfect static refactoring intractable?
You could select fields using computed values in any language using reflection, but no, 99 times out of a hundred you use explicit symbols ("keywords") in Clojure.
At the same time, I think that programing with complex generic structures is a functional anti-pattern anyway: if your type is "Map (Foo, Bar) (Map Int (Set String))", and it's not somehow encapsulated (as in a class or ADT), good luck. In Python you would hide this in a class, but what is the Clojure solution (genuine question)?
This really doesn't appear to be as much of a problem as it was for me in Python and Haskell. Perhaps it's because the accessor syntax "feels right" so that manipulating even complex structures of maps and vectors seems natural, and the code won't look too convoluted.
(In Clojure, maps are functions of their keys, vectors are functions of their indexes. So:
(def a-map {:a 1 :b 2})
(a-map :b)
=> 2also, when using keywords as keys, the reverse also works:
(:b a-map)
=> 2This basically makes normal operations on maps look much like using accessors on Haskell ADTs.)
The reason? When the action was sent over the network, the boolean was converted to a Boolean.
On the same project, I would also frequently suffer from name-typoes -- in structure keys (deftype and defrecord didn't exist yet).
A dynamically-typed language is still dynamically typed. A lot of basic analyses are still undecidable. Lexical scoping doesn't fix that.
That sounds more like an impedance mismatch between Clojure and its underlying platform.
For example:
Confusing ints with booleans.
Confusing ints with pointers (no longer done, thank goodness).
allowing assignments to have a non-void type ("x=1.0" should be of type void, not type double).
Poor choice of assignment and test equals (would be nice not to use "=" but to force ":=" and "==").
Pverly aggressive type coercion and autoboxing.
Integer division using the same symbol as floating point (many end-users want 5/2 to be 2.5 as the answer, mostly you want it to be 3 if you are doing loop-foo).
Making "else if" and if-braces mere convention.
Allowing null way too many places.
Automatic variable decleration.Having programmed in both kinds of languages, I have to admit I love not thinking about the type. I don't buy the claim that static typing saves much in the long run (in terms of debugging or writing better code). In 15ish years of using both static and dynamic typed languages, I think I've been burned two or three times at most by not having static types—it's just not a big problem.
Using static types creates a self-fulfilling environment in which getting the right type is crucial (and thank goodness we have a modern IDE to help with that).
I've done a good bit of programming in Java and that was painful, with relatively little benefit. In fact, for a pretty long time, Java made my an adherent of dynamic typing.
Then I tried Haskell and realized that a good static typing system is very useful. It not only catches mistakes--I've worked on very language interpreters in both Python and Haskell, and I constantly made preventable mistakes in Python--but actually makes your code more expressive. Take the read function, for example. (It is basically the opposite of toString.) I can't think of any way of writing it in a dynamically typed language because it has to know what type it's supposed to be to work--it's polymorphic on its return type. And this is very useful: if you have a string, you can just read it in and it will be of whatever type you need (as long as you've defined how its read).
Haskell has some other advantages over Java (e.g. much more lightweight types and no messing about with classes and subclasses), but I think the important point is that static typing done right is not only safer but also sometimes more expressive than dynamic typing.
- Statically checked and removed at compile-time if possible
- Checked at runtime otherwise
- Not checked at all if you compile with high speed and low safety, yielding close-to-C level of performance (complete with memory faults and weird bugs when things go wrong)
Furthermore, I can turn on this high speed/low safety mode in specific portions of my code (inner loops), while compiling the rest with high safety.
SBCL + SLIME + Quicklisp is a seriously awesome combination.
Edit: Software such as pylint are nice too if you're going to go dynamic only. That helps catch lots of bugs that you'd not otherwise notice.
Then I learned Haskell, which I was more open to because I had heard I didn't need to write down any type information. With Haskell I learned that static typing was great because it created a documentation for your code that was self enforcing, and as Haskellers love saying, if the code typechecked, it probably works as expected.
Now I prefer working in a statically typed environment. But for none of the reasons presented in this article, which I think is just presenting information that those used to dynamic typing will turn their head at.
I will say, however, that weak typing can be used to overcome the first complaint. For example, in Python, if you were to do apply a function that resulted in integer overflow, you would get an automatic conversion to long integers. (The actual example doesn't really apply in this case, since to make that function work at all in Python, you'd have to to make `n` a float, which would fix the problem anyway).
Edit: And yes, I know that Python is a strongly typed language, but implicit conversion to long integers is still an example of weak typing.
No, it's an example of a properly defined integer type. Letting "integers" silently overflow is weak typing. If you explicitly want wraparound arithmetic (by proving bounds for performance or wanting implicit modulo), then explicitly specify int16/int32/int64.
"Weak typing" is a broad concept, but that's not one of the ideas that it covers. In any case, Python's conversion of integers to long integers is not transparent. It is an actual change to a different type, which is an example of weak typing.
Edit: mindslight's comment adds another possible interpretation of "weak typing" (integers which overflow), further illustrating my point.
No keeping APIs in my head or losing flow to a google search, hit dot, see what I can do.
Everything beyond this is gravy.
var x = "Hi!";
x.
I'm not sure how Komodo does with that one without any JSDoc notation. I'm know that there are editors that do a slightly better job (for instance, Visual Studio probably has the best Javascript autocomplete I've ever seen) but still, it really pales in comparison to how well it works with statically typed languages.
But it doesn't go much further than that. For example parameters inside functions don't have intellisense and return values from functions rarely do, but sometimes.
For parameters, you have to help visual studio out by documenting what the parameter does using XML code comments (e.g. /// <param name="arg1" type="String">Arg1 description</param>).
A best case scenario for dynamic typing is either global inference analysis, which is slow and painful, or relying on test/runtime information, which is unstable and incomplete.
That's just a fact, man.
http://rope.sourceforge.net/ is pretty painless. Doesn't seem like you've actually used it.
Cute library though.
Please don't comment on what Rope can or cannot do unless you actually understand how it, or Python for that matter, works.
Complaints like "refactoring" or compiler error checking however are the oldest FUD in the book.
> Until real software engineering is developed, the next best practice is to develop with a dynamic system that has extreme late binding in all aspects.
> --Alan Kay
At this point it doesn't make much sense to say that C# is strictly anything. It's perhaps the most obsessively multi-paradigm language out there.
http://wekeroad.com/2010/08/09/csharps-new-clothes/
In case you plan on skimming that and missing the details the tl;dr is that Rob, pushes `dynamic` as far as he can and it still comes up far short of the extensibility and expressiveness of Ruby constructs.
However, I don't find that a problem. I tend to prefer the correctness of statically typed languages these days.
What makes you say that?
First, I happen to have worked with this fellow before [Hi, John! Long time no see. Fun to come across your post here.] and I'm sure he's quite sincere.
Second, having spent a bunch of time lately in dynamic languages (Ruby and JS mostly), I do indeed miss the automated refactoring tools that existed for Java. It's really nice to be able to rename a method everywhere across the code base without worrying that something else got mangled. I also miss the magic documentation that a type-aware IDE can provide; I spend a lot more time digging through layers of code to figure out what a particular parameter is expected to be.
Not that I want to switch back; I agree that a language with strong type inference is the way to go. But man, I miss the tools.
There are generally very few situations where you have a single method from a single class sprinkled across an entire codebase for a good reason. That always indicates high coupling and design problems. There are even fewer reasons to rename a well designed method. Constant renaming and high-coupling are bad habits that seem to be commonplace when working in strictly static languages and yet they are known to be bad habits even there. At least dynamic languages discourage these anti-patterns. Not to belittle your friend, I found myself falling into these traps constantly when working in C#.
Honestly, what I'm saying is that these problems just are not nearly as significant in say, Ruby, as they are often made out to be by static typing proponents. It's a bit disingenuous to point to these trivialities as reasons why static > dynamic. If it were true then there would be little benefit in dynamic languages. Clearly there is.
More generally, the ability of modern IDEs to manipulate code structurally, and not just texturally, has been eye opening for me.
The theory that you can get good design up front depends on having both a stable problem and a stable solution. Most places don't have stable solutions, because creation of software is usually an exploration of the solution space. (If you don't need to explore the solution space, that often means you should just buy something off the shelf.) Quite a number don't have stable problems, either. Some people have innovative competitors; others are doing startup-ish things.
As Keynes said, "When the facts change, I change my mind. What do you do, sir?" I think that applies to design as well, method names included.
I am reserving most of my replies for the comment section of my own blog, but I couldn't resist say hi to you. I am in fact sick of typing in type labels (so strong type inference is great, it is just finally becoming available to us masses).
Also I would like try defend refactoring. Those that don't think it is important have not seen it used well (there are a lot of trivial uses of refactoring). I remember pair programming with Brian Slesinsky (when the two of you were coaching). We wanted to change some deep functionality of the project. Brian slowed the moment down for a bit and then said something like: "we call this refactor method, then introduce this error here, fix it here and then call this last refactor to clean it up." He was definitely not only doing what the IDE supplied but thinking in a very deep way how to trick the IDE into correctly performing a big change it was not designed for.
I don't think you should post it here if you aren't prepared to discuss it here. That seems disrespectful to the community.
For example, I'd like to hear your response to this objection: http://news.ycombinator.com/item?id=3633596.
Yet it seems to be what the evidence is telling us.
http://squab.no-ip.com/collab/uploads/61/IsSoftwareEngineeri...
But as an argument for one paradigm vs. another it doesn't really stand up, because the essay takes its own conclusion (that large systems would be easier to maintain if they were more like Squeak) as a major premise. Him being one of the inventors of the platform, I don't think we can just take his word for it. At a minimum, what we'd really want to see is a large successful enterprise system built on Smalltalk to serve as an instructive example. To my knowledge no such system exists, so we can't really take the paper, insightful as it is, as much more than hopeful musing.
That said, the "bind really late" approach has seen a lot of success. Just not quite so pervasively as the "in all aspects" that Kay advocates in the paper. Nowadays, the standard way to build large systems is to build completely independent modules and couple them on fluid interfaces. Text streams in Unix, REST APIs, and even SOAP are clear examples. What we don't see, though, is a whole lot of reason to think that the languages and run-times on which these modules run must also support late binding and hot-swapping of code at the micro scale, or that we're really suffering for lack of it.
To be honest it doesn't seem that you are arguing against this concept that strongly, you're just not overly sold on it. I agree it's good to be sceptical, just as I'm sceptical there is all that much value in static typing in most cases. Having used a statically typed language for years, I've personally found the value to be little and the cognitive friction to be high.
I wasn't using the word 'advertise' to suggest he was getting paid to push it. I was using it in very much the same sense in which you use the word 'sell' here:
> To be honest it doesn't seem that you are arguing against this concept that strongly, you're just not overly sold on it.
To which I respond: Correct. I'd even go so far as to say I'm not arguing against the concept at all.
Considering that all sorts of approaches to software development continue to be extremely popular, including among very smart people, it just seems crazy to me that people get so acrimonious about such issues. Programming is a very wide and varied field. If two developers try something and come away with differing opinions of it, isn't it just possible that those different opinions are both informed, but informed by different experience resulting from working in different problem domains? I'd submit that developers don't give each other nearly enough credit when they offhandedly dismiss each other's sharing of their own practical experience as "FUD".
To bring it back to Kay's paper: He's got some very interesting ideas, but the specific examples he talks about are problems that just don't cause me any stress in my day job. Yet he implies that people in my problem domain are suffering for not using his preferred programming paradigm. Now I don't want to accuse him of attacking me, and I certainly don't want to attack him because I believe that paper represents learned speculation and not a stake in the ground. But it remains true that he's not necessarily coming from the same place that everyone else is, and his experience of what works well does not necessarily translate into a universal best practice.
Does such a system exist to anybody's knowledge? I've heard that there's a good bit of Smalltalk on Wall Street, but companies don't like to advertise it because it's a competitive advantage.
"a fully-integrated, model-driven, automated silicon wafer fabrication facility (fab) ... ControlWorks managed fab saved TI the equivalent of 1.2 fab lines last year. What does a fab cost? roughly $1B!"
http://www.google.com/search?q=ControlWorks+smalltalk
Kapital, 70 developers in 4 locations in 2005
"in a statically typed language the language would force.. coercion"
But R just did an implicit coercion, right there. You gave it ints, it gave you back 1e12. I don't know R well enough to understand what the int() call does, but there's no reason why a dynamic language can't perform implicit coercion. Python does:
>>> type(1000000)
<type 'int'>
[define sumXX]
>>> sumXX([1000000,2000000,3000000,4000000,5000000])
55000000000000L # <--- type long
I could wrap the elements of the list in int() calls but that would change nothing.The whole point of dynamic languages is that you don't have to think about the types of your numbers. They're just numbers. If you have a use case that requires "4-byte integers, goddammit!" (and there are many valid ones), why yes, you should use a static language and think about overflow. If you don't want to think about overflow, the languages that go to the greatest lengths to shield you from it (lisps) are all dynamically typed.
* The omission of variable declarations has nothing to do with dynamic typing. Some dynamically-typed languages certainly do this (Ruby and Python). Some dynamically-typed languages make variable declarations optional or optionally-required (JavaScript and Perl). Some dynamically-typed languages require variable declarations (the Lisp family), and so do not suffer from the identifier misspelling problem.
* In my experience, Eclipse and other IDEs are not completely reliable when it comes to refactoring and variable renaming. Yes, they are good at it, but at the same time, they all offer preview modes and encourage the user to double-check the IDE's changes. I have seen Eclipse fail in renaming identifiers in Java code. For Lisps, I have found http://brian.mastenbrook.net/display/26 to be just about as effective as Eclispe's Java renaming.
As for the first point, the misuse of the word "macro" makes the entire argument rather difficult to address. It seems to boil down to taste.
True, but an IDE will refactor accurately much, much more often on a statically type language than a dynamically one. Refactoring dynamically typed languages is pretty much impossible to do automatically when you don't have type information:
http://beust.com/weblog/2006/10/01/dynamic-language-refactor...
IntelliJ's ability to refactor Ruby and JavaScript is a joke and really just amounts to global searches and replaces, and wildly inaccurate results. I typically get out a plain text editor and ditch IntelliJ altogether.
Like I said in my comment above, this has nothing to do with IDEA: dynamically typed languages are just technically impossible to refactor automatically without the developer's supervision.
For sake of argument, let's say that's true and then ask - How much does that actually matter in practice? We have unit tests don't we?
Here's an example -
A very large Smalltalk application was developed at Cargill to support the operation of grain elevators and the associated commodity trading activities. The Smalltalk client application has 385 windows and over 5,000 classes. About 2,000 classes in this application interacted with an early (circa 1993) data access framework. The framework dynamically performed a mapping of object attributes to data table columns.
Analysis showed that although dynamic look up consumed 40% of the client execution time, it was unnecessary.
A new data layer interface was developed that required the business class to provide the object attribute to column mapping in an explicitly coded method. Testing showed that this interface was orders of magnitude faster. The issue was how to change the 2,100 business class users of the data layer.
A large application under development cannot freeze code while a transformation of an interface is constructed and tested. We had to construct and test the transformations in a parallel branch of the code repository from the main development stream. When the transformation was fully tested, then it was applied to the main code stream in a single operation.
Less than 35 bugs were found in the 17,100 changes. All of the bugs were quickly resolved in a three-week period.
If the changes were done manually we estimate that it would have taken 8,500 hours, compared with 235 hours to develop the transformation rules.
The task was completed in 3% of the expected time by using Rewrite Rules. This is an improvement by a factor of 36.
from “Transformation of an application data layer” Will Loew-Blosser OOPSLA 2002
However, I do feel though that IDEA made some poor choices and tried to get refactoring and autocomplete capabilities in JavaScript when they probably should have backed off, and IntelliJ's performance suffers (sometimes greatly) because of it. As a simple example: if you use ExtJS and have it loaded in your IntelliJ solution, try to rename a local variable named ownerct. IntelliJ will completely lock up for about 15 minutes. Why? Because ExtJs uses the variable "ownerct" throughout, and IntelliJ is mindlessly sucking in all those references in. Of course you can set up your IntelliJ to avoid this situation, it's just an example.
I was inclined to say this as well when I was critiquing the first point (http://news.ycombinator.com/item?id=3633351). But on further thought, just about the only dynamic language that flags assignment to a typo'd variable name is scheme[1]. So it seems like a significant correlation.
[1] Not common lisp. Does clojure?
use 5.010; # or 5.012 or 5.014 and soon 5.016With regard to destructive assignment, most Common Lisp compilers warn on setf calls when executed against unbound symbols. Clojure doesn't have even have assignment in the sense discussed here (refs are bound before use, and the compiler enforces this).
However, after I learned Java and C++. I found that I really preferred static typing, and the error checking you get from the compiler.
I've been learning Clojure, but I'm thinking of switching over to Haskell for my next project mainly because it has static typing (also because I'm intrigued by QuickCheck).
(I recommend learning both languages tho)
"Okay, this function will take two arguments, one will be an integer, and one will be a string, and it will return a list of strings."
It helps me reason about my code since I'm never going to want to, say, add an integer to a string, without some intentional type conversion in there.
I'm diving into haskell at the moment and really liking it. I always notated to myself what types a function took and returned, even if it was just in a comment, but with haskell there's a notation for it, and the compiler is pretty sophisticated about making sure that type-wise my function is doing what I'm expecting it to be doing.
+1 to the refactoring example.
programming requires discipline. using the discipline required to do X as reason to not do X isn't going to get you very far.
It is not necessary to add tests specifically for type errors. Unit tests for basic functionality will usually catch them.
The unit tests that were written with the original procedure are likely to only cover usage scenarios that are expected by the person who just wrote the procedure. Someone who's not in that headspace can easily be a lot more "creative" about coming up with surprising ways to try and use the code. So easily that it might even happen by accident. Like, say, as a result of a simple refactor. So in a reasonably-sized project, unit tests for basic functionality end up being a rather short Maginot Line for type errors.
And static typing would not necessarily catch this bug anyway. A C++ template or a polymorphic Haskell function would fail in exactly the same way. What is called for here is argument conversion (A C function or java method declared with double arguments would convert its arguments automatically, but no reason you cannot convert explicitly in a dynamic language.)
Chapter 27 of Code Complete (2nd ed.) discusses this. So does http://news.ycombinator.com/item?id=3037293 among other HN threads. Also, if anyone has access to Capers Jones' "The Impact of Program Size" from Programming Productivity (1986) and is feeling generous, please cite the relevant findings. I can't find it online.
It is not just about completing a word. The completion can be a pattern of several lines of code with placeholders where the IDE jumps and waits for you to enter one or two characters that it completes again.
The resulting experience is that things just flow. The IDE frees your mind from details like name spelling, api method list, exact language syntax and usual idioms.
And it is probably just the beginning. The completion is still pretty basic when you think about it. At some point maybe, IDE will switch to a rule engine to manage thousands of completion rules.
Although not quite the same thing I've been moving towards more rigorous typing through Moose and wrapper classes in my Perl development work/
I believe I would prefer compile time type checking, but when working with a lot of dynamic data structures (say, JSON or XML) there is an ease that dynamic language toolkits can bring that a lot of the time evade more formal languages.
In general, I just disagree. Strictly typed just isn't superior to dynamic typing at all. It's a different approach with its own pitfalls, and more than enough of them.
Dynamic languages DO have refactoring tools. It's just that not many people use them.
That typo error WOULD have been caught with a static analyzer or with unit tests.
Yes, if you supply a function with a type you haven't tested for it might not work. Solution: test for that type, and validate your inputs.
The fourth argument that debugging is more expensive is not substantiated. I find dynamically typed systems easier to debug, especially when they don't have four hour compile times like some statically typed code bases do. Being able to more quickly change, and rerun code in dynamically typed systems gives them a big debugging advantage. Many dynamically typed systems even let you easily change code at run time (which I know is possible with statically typed systems too).
I think more major reasons would be dependent on the types of problems people are solving.
Also... vi forever!