Why I Prefer Dynamic Typing Over Static Typing (2017)
smashcompany.com
smashcompany.com
He seems to consider static typing equivalent to "locking down" an architecture, that when you use static typing you're locked in to a certain way of doing it, and you're stuck with it. In my experience of working with large codebases with both dynamically and statically typed languages, the opposite is true.
Static typing frees you up to make changes to interfaces in a way that dynamic typing does not. For instance, if i want to add or remove a parameter to a function in a statically typed language, it's trivial. I just do it, and the compiler goes "ok, this function is called from these 4 places, so just fix those" and you now are assured that everything will work.
That doesn't work in dynamically typed languages. There's no reliable way to change an interface and be confident that you haven't broken anything (+/- extremely good linters, but those aren't perfect in the way a compiler is). This leads to a "never change this function, you can break all sorts of things" attitude that simply doesn't exist with static typing, where the compiler can just tell you what you need to fix.
There are perfectly valid arguments in favor of dynamic typing (it's arguably easier, often faster to work with etc.) but this particular one is extremely silly. The exact opposite is true.
You are talking about the ability to refactor. In my experience with major refactors, you need a test suite, not the type checker. Why so specific? Very simple, when I did this in the past, the compiler tended to be happy way, waaaaayyyy before the test suite. Had I relied on just the compiler, I would have been in a very bad place.
> "locking down" an architecture
Simple refactors are not "changing the architecture". Static type systems tend to be fairly specific about the kinds of architectures that are expressible, and they tend to be "you can have any architecture you like, as long is it's black, er, call-and-return". Now we tend not to notice this, because we are so used to everything being call-and-return that we don't even notice it.
"We Don’t Know Who Discovered Water, But We Know It Wasn’t a Fish" -- Marshall McLuhan
Once you do notice, and you do want to build alternative architectures, the straight-jacked becomes very constricting.
One benefit of static typing is that it's like a spell checker. It can't catch all errors but there's a whole class of errors it prevents you from making.
I always wanted something that's in-between both: an "analyzer" that warns you of oddities, but doesn't prevent running the program. You could put indicators to stop notices about things you know will be otherwise flagged.
Indeed. Like fopas.
Did you mean "faux pas", or it is something I didn't know?
I will not participate in Hacker News Static vs Dynamic typing debates.
I will not participate in Hacker News Static vs Dynamic typing debates.
I will not participate in Hacker News Static vs Dynamic typing debates.
I will not participate in Hacker News Static vs Dynamic typing debates.
I will not participate in Hacker News Static vs Dynamic typing debates.
Stay strong, Me
Besides, this thread still has only 65 comments and it will likely have thousands before the day is out:
That being said, every now and then I find them a good refresher in software politics, with a side of insight.
The debate fundamentally comes down to 'under what circumstances is the extra cost of working with a static type system worth it for the benefits'.
For a 15 line helper script of the sort I've just written, Python with its dynamic typing is (to me - a proponent of static typing) the fastest thing to work with due to the low overhead. To me there is a threshold of project complexity where the cost-benefit ratio flips. Is that actually true? Does that vary for all people?
Here you go: "A large-scale study of programming languages and code quality in GitHub"[0]. I've only read the abstract (are we OK on sharing sci-hub links? [1]), spoiler alert, according to this research some languages are more defect prone than others but the effect is small.
[0] https://cacm.acm.org/magazines/2017/10/221326-a-large-scale-...
[1] sci-hub.se/10.1145/3126905
It's a total cost analysis that's interesting - and there's a lot of costs. Cost of development, how that varies depending on required time to market, cost of bugs (fixing them, and loss of sales due to customer reliability concerns), cost of complexity as software has to be maintained over time.
Yeah, good point. I read an old research where they tried to test the "productivity" offered by different languages, but then it became difficult to asses the experience of the different developers...
I suspect the same about procedural vs. functional programming.
And I suspect that's why discussions about them wind up the way they do. It's just obvious which one is better - to you. It's so obvious that it doesn't even seem to need proof. And therefore it seems that everyone who sees it differently must be ignorant, an idiot, or a troll.
Or so I suspect.
Nothing described is impossible, or even hard, for a statically typed language to do.
Perhaps I misunderstand them, but author seems to be under the impression that static typing doesn't allow for type conversions.
> In static-type languages such as Java, I’m forced to go with either #1 or #2, and they are both bad options.
That's simply not true.
But I don't think this particular issue has much, if anything, to do with dynamic vs. static typing. It's mostly about how clean and elegant Python's way of handling JSON is compared to the black hole of viral enterprise-y overdesign that is Jackson, Java's de facto standard library for dealing with JSON.
With Python, JSON is deserialized to, in essence, a dictionary that maps keys to values, where the values are either strings, arrays of values, or dictionaries.
With Jackson, it's exactly as awful as TFA describes.
But I can easily imagine a Java library where JSON deserializes to some "JSONEntity" type that can act as either a string, a list of entities, or a string->entity map. And then I could interact with that in the same loose, don't-force-me-to-worry-about-details-I-don't-currently-care-about way that I do in Python. Perhaps one even already exists. It doesn't seem a far-fetched dream - I've used similar libraries for other static languages.
TBH, I haven't looked, because, even if I found one, I'd still be compelled to use Jackson, anyway. Because Jackson has been around forever, and Jackson is now so deeply entrenched that nobody working in the Javaverse can even fathom life without it. Suggesting such a thing would be heretical.
https://fasterxml.github.io/jackson-databind/javadoc/2.2.0/c... and related classes/functions.
The linked article, however, is a whole other level of incorrect thinking and drivel.
This is rather a broad generalization that doesn't seem accurate to me. It's at least misleadingly phrased if it's not intended to imply that this is the main reason or primary domain where static typing is used. Performance is a major reason static typing is preferred in many domains like games development for example where legal regulations are largely irrelevant.
What I see between each approach is tradeoffs. Which one has the net benefit may ultimately depend on personality, domain, and shop practices. It's one of those architectural holy wars that may never end.
I think personality is often the biggest issue. Some people like dynamic, some like static. You can make some rational arguments for your preference but in the end we know that successful projects have been written with either paradigm.
They're both manageable, really.
My problem is nearly every type system breaks down when you try communicate with something outside of the closed system (e.g. API calls, 3rd party integration, etc).
An API returns an unexpected data type, Missing a field that's supposed to be required, Starts returning a number as a string because they changed their ID system, etc.
It always seems to end up in types where every field is optional - defeating the purpose of having types in the first place.
However, the problems you describe also happen in dynamic languages and they have just as much potential to cause problems. The biggest issue with dynamic languages is that you don't get immediate errors when something changes. The system will just keep doing its thing, until it doesn't and you discover that something has changed and has already been propagated to other parts of the system.
I don't think dynamic languages are always better in these cases. They basically let you defer the error/change handling, but with the potential of much bigger data integrity/quality problems when things eventualy do go wrong because of changes in an integration.
> It always seems to end up in types where every field is optional
“Always”? Not by a long shot. These situations where this is the case can be isolated and dealt with explicitly. You may object to the explicitness but — again, for reasons of robustness — proponents of static typing prefer this.
I didn't work much with static typed languages lately, so I'm really puzzled by your comment. If I expect an int in the age field of a JSON response and the server suddenly sends me a string, maybe with value "NaN", how do you check for that at compile time? I expect the program to either crash or manage the error at run time (try/catch, die-and-restart a-la-Erlang, etc)
Maybe JSON doesn't count as "statically typed API", but how many of them do common Internet services expose to the public? And how to we make sure that the server or any in between proxy doesn't misbehave and violates the contract?
Does that make sense? I can show you some code if you’d like.
Either way, you need to handle such failures, both in statically and in dynamically type-checked languages. The only difference is whether the handling happens explicitly throughout your code or automagically by an API that gives you back a well-behaved object. And both of these ways are possible in both static and dynamic languages (but static languages tend to go for the latter, whereas dynamic languages tend to make you do the former).
In statically typed language, at least when troubleshooting it's an easy guess that the error is highly likely to be at the external boundary of your application code since the rest of it type checked.
Let me turn this around: if you're saying possibly-missing fields inevitably propagate throughout the types in your codebase, it sounds like you're saying malformed input propagates arbitrarily deeply into your core logic. Wouldn't you rather catch that at the boundary?
Possibly. If it's a key part of the application, like Credit Card processing or auth - absolutely. If it's simply auxiliary or supplemental information - probably not. The thing for me is the type system shouldn't decide whether or not a certain piece of information is critical to the system. That typically falls to deeper logic in the system. If the type system isn't deciding, then it's pretty much meaningless (as everything would be optional).
For example, I'm using a Stripe JS integration to allow customers to manage credit cards in app. Stripe messes up or changes something without me paying attention. Instead of returning a complete object, Stripe starts returning only the credit card ID.
This isn't a big deal if we simply need to know that a record exists. We can still tell the user has a card and we have an ID to attempt a transaction against server side. We might not be able to display details about the card, but users could still submit a transaction (likely against their last used card).
However, it is a big deal in a credit card management interface. Without supplementary data, all of the cards would look exactly the same and the UI might break if expected fields aren't present.
----
I don't want my type system deciding what the use case is or how important one case is compared to the other. The actual implementation needs to decide if it has the information it needs to do its job.
No, that is in fact the stated purpose of the type system. If you don’t want to make such a decision statically, you mark a field appropriately (e.g. via an Option type).
If everything is optional, then you might as well not use a type system in the first place.
data Stripe1_0 = {
field1: type1
field2: type2
}
data Stripe2_0 = {
cc_id: str
}
data Stripe = Stripe1_0 | Stripe2_0Most of the time optional fields:
* Encode that your module doesn't care about such-and-such fields, so the parser / type checker should just ignore them.
* Are used to encode variants of the same data that slightly changed in another module [id changed from number to string], so you should use an union type.
In practice, designing a statically typed ecosystem for wire data is hard [0]. Grpc has basically gave up on validating field presence altogether [1]. These failures are directly caused by the immaturity of the ecosystem [deep decode batches of unrelated messages in middleware, WTF?]. It would be a mistake to throw the baby with the bathwater and blame the mathematically sound idea that your module can only process data of a certain shape, encoded by a suitable type signature.
[0] https://capnproto.org/faq.html#how-do-i-make-a-field-require...
[1] https://github.com/protocolbuffers/protobuf/issues/2497#issu...
It's more of an issue that you have unreliable 3rd party vendors that you buy from/use.
It happens very, very rarely - but the problem is when it happens it can happen suddenly and be very, very hard to address properly.
I'm interested in what an implementation of option 3 would look like. Is it just some regex to find the relevant characters in the string?
It seems a bit flame-y to see this as "dynamic typing is better than static typing", though. -- It's more of a suggestion that a technique like "cast to a type" where you don't need all the info in the type risks breaking when it doesn't need to. It's better to make fewer assumptions when relying on things you don't control.
If the external API changes, code breaks whether you were "dynamic" or "static". "Just change the code" is as applicable to code in either case.
That said, it's kinda weird that an API can be both unreliable enough as to not conform to a schema, but still reliable enough to trust the data you want to get out of it.
Wait what? Why would you ever work directly with the data that you grab in some foreign API? You map the data that you actually need into some datastructure of your own making, that you have control over, and that hardly ever changes - unless you want it to. So when the API changes (or even worse, you have to exchange it for whatever reason - it's foreign after all and not under your control), all you have to do is change the mapping process, and voilà, you're back in your own safely typed world.
Things like null-ability, base-type, and max length should be supplied by the database and therefore not echoed in source code per the Don't-Repeat-Yourself principle (DRY), unless a custom deviation is necessary.
If the screen validation needs to know such info, it can get a majority of it from the database dynamically. But, our stacks don't typically do this for unknown reasons, and thus we reinvent the field attribute wheel in app code.
I posted a question similar to this on HN recently: https://news.ycombinator.com/item?id=19146439
Automation is good, right?
With dynamic typing, you need to manually write more unit tests to get the same level of confidence. That's just not worth it for anything important.
And in languages with type inference, you largely don't even have to do the annotating.
Programs have (static) types, data does not. Static typing removes correct programs from the language. Try this in some ML-like language:
let selfapply f = (f f) in
let identity x = x in
selfapply identity 42
If you transcribe this to (e.g.) Lisp, it will return the number 42. The program is short, simple, and safe, but well-regarded static type systems can't cope.From his Java's example, I can't see how he can fix with dynamic language. He still needs to validate the API response and throw exception even he uses dynamic language. Unless he uses something like functional "Maybe", so he can safely process the response.
It makes it very difficult to comply with Postel's Law.
Not really, since “handling” doesn’t mean “throw errors”. But regardless, Postel’s Law is often seen as a failure in hindsight. See the “Criticism” section on its Wikipedia page. In general you don’t want to comply with Postel’s Law.
But the whole process of creating Beans or immutable POJOs and then decorating them with annotations seems to invariably lead you down a path where, if you're dealing with a particularly hoary API, you've either got to create a ridiculous number of single-use DTOs to handle all the edge cases, or define a manageable number of DTOs that unfortunately also place more stringent requirements on the input data than are strictly necessary.
This isn't always a bad thing, because sometimes "any implementation of a product" is vastly superior to that product not existing, and works well enough, often enough, that any edge cases can be ignored safely.
Things like null-ability, base-type, and max length should be supplied by the database and therefore not echoed in source code per the Don't-Repeat-Yourself principle (DIY). If the screen validation needs to know such info, it can get it from the database dynamically. But, our stacks don't typically do this for unknown reasons, and thus we reinvent the field attribute wheel in app code.
(I posted a question similar to this on HN recently: https://news.ycombinator.com/item?id=19146439 )
I'd like to see a data-dictionary-driven approach where most of field info comes from a data dictionary, and only deviations need to be defined locally. (Or perhaps create and mark dummy columns in the data dictionary that are slated to be used for local customization.)
I suppose one could push a button and have static classes or definitions generated for the source code from the data-dictionary, and thus still have a statically typed system. It's still duplication of field attribute info, but at least it's systematic duplication.
I used to use systems/tools that integrated the database, business logic, and UI; and they reduced the amount of typing and code rework by a large amount. The multi-layered approach may give us cross-vendor flexibility, but at the cost of more code diddling to wire up the interfaces between all the layers. Layer independence seems to increase duplication of field-related info. Back then I could spend more time on requirements and analysis and less on code wiring and re-wiring. Now we waste too much time dealing with low-level repetitious grunt work.
[1] https://courses.cs.washington.edu/courses/cse341/18wi/videos...
[2] https://courses.cs.washington.edu/courses/cse341/18wi/videos...
So I guess both sides are right but it’s good to think about what suits you best and then embrace that.
Funny, that's the same reason I prefer static typing.
>In How ignorant am I, and how do I formally specify that in my code? I said that I liked to add run-time contracts to a function as I better understand it. When I first write a function, I may not know for sure how I will use it.
Which is irrelevant. You still know that the "customer_name" argument to it will be a string and the "age" will be an integer, no?
>For the programming that I do, I am often creating new architectures. Therefore I need dynamic typing.
Non seguitur.
>Consider dealing with JSON in Java. Every element, however deeply nested, needs to be cast, and miscasting leads to errors.
Another bad argument. First, if you don't cast, and expect an integer like 55 but get a string like John, your program has an error. The difference is that in a dynamic language your program will continue to run doing stupid things instead of crashing in the miscast.
>Given JSON whose structure changes (because you draw from an API which leaves out fields if they don’t have data for that field) your only option is to cast to Object, and then you have to guess your way forward, figuring out what the Object might be.
You can trivially have a JSON parser that returns a list of map of the 6-7 valid JSON types (maps ("objects"), floats/"integers", strings, booleans, null, arrays).
You don't need to deserialize to specific object structure/schema as is the default mode in some parsers, if that's what you want to avoid.
If you dislike the verbosity of getting nested elements in Java, that's a Java thing, not a static typing thing. Heck, it's very easy in Haskell (what with lenses and stuff).
With a language that supports Optionals and try or has a map index operator and list/map literals, it's more like JS or Python than Java.
If you just want to extract a deeply nested element, and not care what it is, but still use it, I'm not sure how that works.
Do you expect the object at e.g. resp["x"]["y"]["z"] to change types randomnly?
>Static typing does not live up to its promises. If it did live up to its promises, we would all use it, and the fact that we don’t all use it suggests how often it fails.
"Proper weight does not live up to its promises. If it did live up to its promises, we would all be on it it, and the fact that many are obese suggests how often it fails"
How about some uses of dynamic typing are for specific, constrained, use cases (e.g. glueing, quick scripts) and it increasingly breaks down with larger codebases and multi-developer projects?
Larger codebases and multi-developer projects either don't use dynamic typing (most huge projects the world relies on, from compilers to browsers, and from OSes, to all kinds of embedded hardware, games, desktop apps, etc, are done in one of C/C++/Java), and others (e.g. large websites) are moving towards some kind of types (e.g. Typescript and Flow).
I've seen integers sent as strings and others as integers in the same JSON. Datetimes stored as strings in three different formats by the same application in the same column of the same table, etc. The real world is really messy. Writing a new program is definitely quicker with a language like Ruby or Python for me (especially Ruby.) Working with Java is like walking with a cast on a leg.
But you know, every project and every developer is different, that's why we are using so many different tools. The OP accounts for that in the "There are many arguments in favor of static-typing" paragraph.
You still need to know what they are in each case, regardless of dynamic vs static typing (it's a matter of whether the language has weak vs strong typing or overly accommodating coercion).
If you get an integer value as string for example, Python won't let you add it to another integer you got as integer. So you'll need to know what they are to treat them accordingly in your processing even in a dynamic language.
In such case, you can also read them as a variant type on a static language, that's eg. one of the 7 allowed types in JSON (you don't know whether a JSON primitive might be int or string, but you know it wont be suddenly a Date type or a Foobar instance).
So what? Just use a veryFlexibleUint64: https://github.com/sapcc/limes/blob/d51ecdbb763318da146b9f32... ;)
Yes, that is why Java and C++ are such tiny unknown little used languages
Seems fair enough?
1. Refactoring. In, say, c++, you change a signature, or a member name, and your IDE provides you with a neat, comprehensive list of errors to fix, if you're even performing the change manually. Contrast that with python, for instance, and I have to rely on what appear to be heuristics if I use an IDE to propagate a change, or otherwise repeatedly hit run to change everything that I've broken, and pray I didn't miss any code paths that'll bite me later. Maybe I'm just not writing pythonic code, but I find it rather difficult to maintain large codebases in python for this reason.
2. Focus. I realized that having a constant stream from a trusted linter makes it really easy to stay on task. I can throw out a bunch of scaffolding and then fill in the blanks based on the constant stream of errors, pre compilation. Sure, in pycharm I can run code analysis to get a list of probable errors, but given the nature of dynamic languages, there'll be plenty that's missed and plenty that isn't necessarily right, and not having a project wide linter in real time isn't the same.
3. Understanding APIs and library internals. There is a TON of use information that comes for free with explicit static typing, especially for scientific/numerical code. In python I frequently find myself having to look up simple function use because arguments, names, and returns can be ambiguous. In a typed language, I know exactly what goes in, what goes out, and how to structure my data appropriately. The difference in coding speed is huge, and I don't have to run the code nearly as often to make sure I'm doing things right.
Why anyone would choose python for large projects eludes me!
"Mypy is an experimental optional static type checker for Python that aims to combine the benefits of dynamic (or "duck") typing and static typing." [0]