Again, that's not what research shows. Type checking bugs are most definitely not the most common ones and typing doesn't really help with runtime bugs unless it's a strongly typed language.
I'm all for types, but after using TypeScript and seeing everyone (including the article) touting it as the solution, it simply isn't all that better than vanilla JS.
I did spend countless hours having to wrangle typescript when integrating with third party libraries, searching for types that were not always there, manually extending/creating library types and so on.
So no, it's not as clear cut as some say it is.
If we want to argue for types, then we should use strongly typed languages that prevent these bugs (if you can manage to create the right types, of course). Examples in TypeScript are simply not going to cut it because that's a horrible example of a typed language. I mean, it's not even a language.
Edit: above all, hard research. If there's research for this, why isn't that in the article?
So Typescript then? The most common recurring bugs on my team is front-end devs not knowing the shape or property names of types written by other devs (backend, other font-end, themselves a month ago). If we moved to Typescript they would get those little spellcheck-like underlines pointing out these errors before they attempt to commit, rather than customer's pointing it out.
Sounds like an argument against untrusted input, I just want Typescript so at least my own team becomes trusted input.
This is not a facetious question, so bear with me: what is a "type checking bug"?
Do you mean "a bug in language L that would have been caught by the typechecker T, if only L had it"?
That's the only meaning I can think of that makes sense [1] (ok, thought of another one [2]). But what does it mean that "most bugs" aren't of that kind? Is it that "most bugs are of a kind that cannot be caught by any conceivable typechecker"?
--
[1] We can safely disregard the other possible interpretation of "a type checking bug is a bug in the typechecker itself" ;)
[2] Maybe "this expected a list but I passed it an integer instead"? I do occasionally trip over those with Python, but I think most typecheckers are more advanced than that, and can check for more interesting classes of bugs.
That's the most basic bug that typed languages catch. But for me, those are extremely simplistic. I don't need a type checker to know that. If you follow the code path you already know the types you're dealing with.
What strongly typed languages can provide are types with constraints. E.g. an age is not only a number, but a valid number greater than zero (maybe with an upper bound?) that is attached to a person. That kind of strong guarantee will definitely catch runtime bugs if you pass in some other numeric value that is not a representation of an age.
Languages like TypeScript can't do that. Take the example below. It compiles as if there's nothing wrong:
type Age = number
type Person = {
age: Age
}
const person = { age: 10 } as Person
const someRandomNumber = 10
function printAge(age: Age) {
return console.log(age)
}
printAge(person.age)
printAge(someRandomNumber)
This is the kind of bug that a good type system would be extremely helpful with. In my entire career it's definitely a lot more common than passing a string by mistake. This can be extended for having a type that differentiates an `Email` from a `ValidEmail`, an `ActiveUser` from an `User` and so on. Those are, in my view, where a type system can truly help us catch bugs. But even if you have a type system that supports that, nothing stops someone from simply using `number` instead of `Age`.In any case, that's why I don't think it's as clear cut as people put it. Yes, in very rare occasions TS will tell me I accidentally allowed something that can be undefined pass (excluding the many times it gets it wrong). However that doesn't come for free and that is what makes absolutist claims for either side unhelpful.
Structural typing is great when programming functionally (i.e. with immutability) when the most important thing is the shape of inputs and outputs of functions, instead of named objects (like Person) and properties. In a functional program, for example, "printAge" would likely be called something like "print" or "printNumber" since that is what it is doing to its input _value_.
I think a lot of the misunderstanding I've seen recently around TypeScript (like from the Rails creator) comes from the misuse of TypeScript - if you use TypeScript in an object-oriented way, its going to be significantly less helpful.
I don't follow this one. I've never seen anyone use TypeScript with an OO approach aside from, ironically, the .NET folks.
The code I wrote has nothing OO with it and we can already see the issues. The majority of TS I've ever worked with was written for React and it still would benefit greatly from nominal types as you call them (thanks I didn't know that terminology).
I don't see everyone misusing TS. For me, it's simply a very limited language as far as typed languages go. As a result, it's a shame that it's what is being touted as a good example of why you should use typed languages.
Then, if you have converted lots of Javascript into Typescript, you have probably found several type-related bugs.
On one hand, this makes you (me) trust Typescript more, and feel safer knowing that your types are correct.
Later, you discover that a few of the declared types are wrong! You also still get "this" wrong in callbacks.
You could say that Typescript lets you try to get your types right, but you have to do the heavy lifting yourself.
In your example it would be:
type Age = Opaque<number, "Age">
And then you wouldn't be able to pass a random number to printAgeCan you point to that research?
I’m skeptical of this research. Typescript prevents runtime bugs for me day in day out. It also allows me to produce much more elegant and flexible solutions which would be infeasible in plain JavaScript. When people say there is no significant benefit compared to JavaScript it almost feels like I must be on a different planet or something. I would love to see what kind of code these people are working with, or how they’re attempting to leverage the type system.
What does this have to do with static types? High-level languages, typed or untyped, are generally memory-safe. They'd need a very good excuse not to be. If Haskell isn't a "real" statically typed language, nothing is.
TypeScript is not this; it's a type checking layer over a dynamic language.
If there's more, I hope people reply here with it!
The evidence behind strong claims about static vs. dynamic languages - https://news.ycombinator.com/item?id=16287083 - Feb 2018 (1 comment)
The empirical evidence that types affect productivity and correctness - https://news.ycombinator.com/item?id=8594769 - Nov 2014 (25 comments)
Which is what I would expect: a proper research on the topic would be terribly expensive, requiring multiple teams reimplementing complex projects in different languages and separate distinct expert panels evaluating development and results.
Regardless if it's no meaningful difference or no good research, the point remains: we simply don't know, so these discussions are about opinions for either side, not facts.
But the basic argument is that a programming approach — having a compiler and a type system is not a “style” btw — that employs development time tools to reduce the burden of runtime operational tools, afford greater application of a wider set of optimization techniques at runtime, and also add to the information bandwidth of source code via type annotations is reasonably expected to be more rigorous than the other approach.
As for the move between runtime to development time tools, it still misses the cost of it.
We could move all our code to Haskell and have absolute guarantees for a lot of the common found bugs, but we don't because it's costly. And I don't mean rewrite cost, I mean the cost of its own complexities.
Nobody argues that typed languages aren't more rigorous, but that's not the only variable we care about.
So where does that leave us? Opinions are one option - comparative views to other ‘industrial’ age type of activities may be informative.
I propose to you that “we moderns” live in a typed world. It is not strongly typed but it is typed. One could argue that that is a side-effect of physical world artifacts produced at scale. I would be interested in hearing the argument as to why that near universal phenomena* does not apply to software, in your opinion.
(* Industrial production at scale and emergence of standards)
My issue is a practical one. Using limited typed languages like TS has several drawbacks for little benefit. Using a strongly typed language like Haskell would add a ton of greatly needed rigor and correctness, but it's also not without huge drawbacks. Same goes for dynamic languages.
It's not a question of whether or not we should model our software based on our worldly types, it's about how strict we should be about it and the benefits and drawbacks that come within this spectrum. For that reason I argue there's no single general answer to this and claiming there is one is nonsense.
It's likely not the most common bug family to reach production or to be published on github so it will be visible to researchers doing their studies (survivorship bias…), but they are definitely the most common bug I'm facing when using an untyped language, and the majority of developers shares this sentiment.
Of course you'd argue that personal feeling isn't science, but studies falling to survivorship bias aren't good science either…
I should have prefaced that with "IMO". That particular sentence comes from my experience/career, so YMMV.
When I'm reading code, I know what types I'm dealing with. Even more so if I'm creating that code. But then again, I worked with dynamic languages for a long time and this could be a skill I picked up because of it.
Even when dealing with third-party code, or joining an existing project?
> Bad programmers worry about the code. Good programmers worry about data structures and their relationships.
Of course I still have to read and learn that. But that's the same with or without a typed language. Whether your data is an unstructured hash or a carefully typed structure, you need to know it.
Way too many people rely on ctrl+space based programming, figuring out everything as they go. One aspect that dynamic typed languages forces into us (at least it does to me), is to learn the data structure and relationships more thoroughly.
Also, if I am using an API another team is providing - I don't want to have to go and trace through their code.
Anecdotal story: We recently went through an effort to introduce strict-type checking for our typescript backend. As part of that effort, we introduced concrete types for every API accessible across a service boundary, which caught 6 mismatches between the documentation and the implementation (e.g. return Promise<bool> vs Promise<Instance>) and several hundred issues with missing null or type checks that would result in an "InternalServerError" instead of a meaningful error message.
We started rolling the strict version out this week, and we can already see a meaningful improvement in the sentry logs.
TypeScript actually performed quite well in the last academic study I read on this debate (like 6 years ago).
The real problem with web APIs is that there is always some lossy conversion between type systems as we cross boundaries. So we can't really make some sort of closed system assumption. A typical web app may interact with dozens of API services, some run by third parties. And maybe you can trust their API docs, but maybe not really, and they're subject to updates anyways, so you always have to be on your toes.
Even internally, you can't really control all the type info from end to end. Even the most monolithic systems will have some sort of abstraction leak when going from JSON -> Object -> Relational storage. Even largely monolithic systems will typically break off some functionality (like email sending, websockets handling, etc) as a separate service. The boundary creates a co-evolving connection between separate services run by separate teams, with separate upgrade cycles, even without full microservices buy in. And that creates potential runtime type errors when mapping between these layers.
Even if you've somehow plugged all those leaky abstractions and your tight type system has handled all the edge cases and is provably correct: the user will teach you otherwise on the UI layer. User input can vary wildly, and everything from device capabilities to personal disabilities to network throttling to authentication to using weird ISO characters to file sizes to strange input devices and legacy systems with their own quirks, will completely throw you off at some point.
So no matter how type safe your language is, you'll always have to deal with the untyped and unpredictable user layer. And the reason why JS has been so successful there is because of how flexible it is. It's not as painful to make quick tweaks with JS as it is with a type system like Rust's, for example.
JS, for all its quirks, made a good amount of trade-offs for its target platform.
Type errors are almost always the easiest types of errors to fix. What really will get you is debugging interdependent systems and services, and that pesky user layer. But people will spend extraordinary amounts of time maintaining complex type systems just so their OCD can be satisfied about believing, that at least for a moment, if their program compiles...that all is right with the world for that brief moment just before you deploy to production.
Oh were that the case. I do think types are essential for mission and life critical systems, but testing is even more essential for those cases, so everything should already be thoroughly covered. For consumer apps, however, is it worth the cost in velocity?
Static typing is where you significantly decrease your development speed and significantly increase your bug count in order to increase the performance of your software.
But even then that's only as bad as not using any types at all, the statement that statically typed languages create more bugs is just plain ridiculous - there's a reason that space-faring agencies and defence contractors exclusively use statically typed languages with very strong rulesets on how code is written.
(30+ years of experience here working on million+ lines systems).
In any large statically typed program, 2 out of the 3 lines exist to make the compiler happy rather than the software developer.
Static typing is a terrible terrible way to develop code once you start measuring it objectively.
I use Scala and the syntax is pretty slim to get very rich typing. Love it compared to non-static typing.
Dynamically typed programs have much simpler structures and are much easier to reason about and test than their static typed counterparts.
The larger the program you write with static typing, the worse it becomes compared to it's dynamically typed counterpart. The internal complexity of statically typed programs grows at a much higher rate. This increases development times and decreases program correctness.
Through I don't think I can really do a fair comparsion with a functional language like Scala. I've only used Scala for 2 weeks on a single project. Functional languages tend to increase code correctness by trading away performance.
Consider a simple program that adds 2 numbers: a & b together. In a dynamically typed language, this would be a + b. In a statically typed language, once you consider overflow and underflow, you will likely have a hundred lines of code.
This creeps up again in JSON parsing where the types really are defined at runtime.
But the most general case is people creating complex and hard to understand types in their code, which they then export to unsuspecting developers to use. e.g. The type of stuff that forced C++ to introduce the auto keyword.
https://stackoverflow.com/questions/52376716/c-overloading-o...
literally a single line of code to overload an operator.
anyway, if you don't consider error states in your dynamic implementation and you do this in static then yes, it could be simpler. But static doesn't force you to consider error conditions - just like this C++ example does not.
and yes, if you fall back on javascript notions of default cross-type arithmetic (1 + "1" = ?) then you will get something, and if you take great care to structure your code you may get something useful as a side-effect, but, this is not meaningfully different from the error-state for most people. Getting [] or "11" isn't the expected outcome unless you've come to expect that quirk of javascript, and is something that a static compiler would rightfully complain about - because a user asking for an undefined operation computed across incompatible types seems like a classic type error. If it's not, then just define the operators that are meaningful to you - and you could define it as an interface if you wanted it to work with a bunch of classes etc.
Again though, the complaint elsewhere that "a lot of this debate just ends up being dynamic programmers who are unaccustomed to using static types at all and think it must be so super burdensome all the time" seems pretty accurate. Defining custom operators is not something that occupies even 1% of my time in any static language.
Another example is A is -9223372036854775807 and B is -10. Again you will get an incorrect result.
Once you finish fixing these bugs it will be at least 100 lines long. The dynamically typed version of the code is just A + B.
"a lot of this debate just ends up being dynamic programmers who are unaccustomed to using static types at all"
Static typing is objectively the inferior approach. It's just that people are too lazy to learn how to code with two different typing systems.
Overall statically typed programmers only code at 1/3 of the speed as their dynamically typed counterparts. Of course, if you have only ever done statically typed programming, you are not going to understand how bad it is in comparsion.
> 9223372036854775807 + 10 :: Integer
9223372036854775817
>>> numpy.int_(9223372036854775807) + 10
<stdin>:1: RuntimeWarning: overflow encountered in long_scalars
-9223372036854775799It's a little bit simpler than anything you will get in an statically typed language. It works for integer and floats of any size. It even works for lists!
You can't do that in a statically typed language without a ton of code.
Oh, look. Something that statically typed languages solves.
In fact JavaScript numbers aren’t magic, they are 64bit doubles, which have their own strange behavior and special cases, many of which could be considered “incorrect”. You’re acting like dynamic typing solves problems that it doesn’t solve.
JavaScript is terrible because it was built over 3 decades by different competing browser vendors.
If we're discussing weakly typed languages then this argument is more valid, but that opens up to a new class of bugs strong dynamically typed languages don't have.
Anyone who says the conclusion is obvious have probably not thought this through.
Dynamically typed languages are more powerful languages that can handle problems like this correctly.
Much too often, any discussion of typing systems often boils down to specific traits around someone's favourite language. These specifics are obviously important enough to warrant a lot of skepticism around too general conclusions from studies.
Bascially we're all recounting anecdotes, so let's be honest about that. I don't doubt yours, I just don't think they necessarily reflect fundamental truths about programming.
You're explicitly comparing apples and oranges here. Are you implying that it's never necessary to check for overflow and underflow in dynamically typed languages, and that it's somehow mandatory to do so in statically typed ones? Or, in case your argument is "numbers in dynamically typed languages do not overflow", are you aware of BigDecimal in Java or sys.maxint in Python 2?
Yes, running arbitrary code at compile-time can produce arbitrarily complicated results if you're not very very careful. This is not a type system issue: lisp macros are just as dangerous.
> abstract base classes
Again, nothing to do with types: your complaint is with Java-style OO. Which is garbage, but for historical Java-specific reasons. Using Java as your prototypical example of a typed language is like using PHP as your prototypical example of a dynamic language. There's no such thing as a feature good enough to rescue a bad language. That doesn't mean no feature of a bad language can ever be good.
> generics
Generics produce strictly simpler code than copy-pasting the same implementation over and over again. That's their entire purpose. A generic function isn't a way to give the same name to different pieces of code (the way that inheritance is, for instance), it's literally one function: `mapIntList` is `mapFloatList` is `mapStringList` is `mapIntListList`, all the way down to the compiler output. Unless you're doing some pretty serious performance optimization, giving them different implementations is a bug. And you wouldn't do so in a dynamic language either!
> non-trivial user defined types
These exist as part of your program structure whether you like them or not, the only question is whether they've been made explicit.
> Consider a simple program that adds 2 numbers: a & b together. In a dynamically typed language, this would be a + b. In a statically typed language, once you consider overflow and underflow, you will likely have a hundred lines of code.
Considering overflow and underflow in the static case but not the dynamic case is stacking the deck. Either you're using arbitrary-precision numbers or you're not.
Here's Haskell code for adding two numbers:
add :: Num x => x -> x -> x
add a b = a + b
If you don't want to worry about overflows, you use `Integer`s, which are arbitrary precision. If you're worried about performance, you use `Int`s, which are machine integers.You can complain that this isn't "really" the addition code, since we're calling out to the `Num` instance, but the same applies to dynamic languages. Python's `a + b` is calling out to `__add__` in exactly the same way.
> This creeps up again in JSON parsing where the types really are defined at runtime.
They are not. The type of arbitrary JSON is (using Haskell again)
data JSON =
| JSONNull
| JSONBool Bool
| JSONNumber Double
| JSONString Text
| JSONArray [JSON]
| JSONObject (HashMap Text JSON)
(You don't have to use a hash map, of course: decoding the raw text stream to JSON is a separate step). You then check that it has the structure you expect by implementing a function `jsonToFoo :: JSON -> Either MyPreferredErrorType Foo`.That just makes no sense. Both programs have the same underlying structure, but it's hidden in dynamic languages. Having no compiler and type system makes reasoning much more difficult, there is no discussion about that.
> The larger the program you write with static typing, the worse it becomes compared to it's dynamically typed counterpart.
Again, that's just not the case: try refactoring Python code without type hints in a large code base. It takes a lot more time and effort and you can never be sure you caught all the cases affected by changes. The compiler tells you about every change needed to be made to arrive at a working state.
>Both programs have the same underlying structure No they don't, this is just a massive gap in your knowledge.
> try refactoring Python code without type hints in a large code base
I've done that on multiple occasions, the test suite is fine, thank you very much.
That's simply not true.
> No they don't, this is just a massive gap in your knowledge.
Based on your responses here I don't think you really know what you are talking about.
> I've done that on multiple occasions, the test suite is fine, thank you very much.
Well, if you say so.