Duck typing is safe (2020)
jerf.org
jerf.org
Funny, not too long ago I encountered a logging library with a writer that expected each write call to pass in a fully-formed log message, which happened to be in a JSON format. That meant that if you passed the writer to anything that expected a properly byte-oriented stream - like, say, a JSON serialization library - the writing would typically be done in chunks, and each chunk would be sent in isolation to a remote server that was expecting fully-formed JSON, and that server would silently drop the malformed data. It was weeks before I figured out what was going wrong.
This would be a great refutation of the article if it had happened in Go, but no, this was Rust, and the library authors had explicitly marked their type as `impl std::io::Write` without understanding why that wasn't appropriate.
I guess the moral is that semi-sensible isn't good enough: the real danger isn't that you end up shooting a physical gun; it's that you shoot your video game gun in a subtly wrong way that takes ages to track down. Failing to compile is loads better.
Silently dropping log messages, without even logging that logging is failing, is an even worse defect since it is also terrible without typing errors, e.g. in the case of clients earnestly trying to pass JSON as text but making some mistake with quotation, escaping, commas etc because it is text.
So the logging library is inadequately designed and not "semi-sensible" at all.
The silently dropping bad input, yeah, that part was just bad developer experience on the vendor's part.
This OurJSONLogRecord still surfaces the undocumented assumption, "Oh, we need the entire JSON document, we didn't realise anybody would want to stream data" whereas the Write trait does not.
How can it be that in so many programming languages, adhering to a contract means nothing but "these methods have similar calling signatures" and "I pinky swear I implement the semantics of std::io::Write " ?
Haven't the Ocaml people entered the next metaphysical level of understanding? When they speak of "types" they don't mean "uint32_t" or a class.
:(
Name the library so people can avoid it? Or did they fix this and drop broken versions?
(edit to add: the feature was in beta at the time and we were early adopters, so Support not knowing how to field issues about it was sorta understandable. Their engineers seemed cool when I had the chance to talk to them)
And to be clear, the system would be broken for the plaintext providers too, just more obviously broken as you'd presumably have a bunch of JSON tokens (or segments of a format string or whatever) show up as individual log lines.
This is a joke I’m not getting, right?
Somebody is forgetting that comparison and arithmetic operators are an interface on a type, and that almost every popular language (even many that are ostensibly “statically typed”) have ubiquitous nightmare scenarios surrounding implicit casts.
And of course, it’s also common where data comes from an i/o source and then is magically assigned a type by an ORM or deserializer or framework helper.
Or when working with large projects with lots of modularity and abstract interfaces, where the names of parameters, methods, and members become increasingly generic.
People work around these issues and learn how to avoid them with good practice, but they happen every day.
I have heard that this is one of the things that Scala 3 fixes, but I haven't done any programming in that yet.
For what it's worth (and in the context of the article), Go doesn't do implicit conversions like this: you get a compiler error if you compare an int32 to a uint64, for example, and you always have to convert one to the other type explicitly. That goes for comparisons and arithmetic (neither of which are overloadable in Go).
Well, C/C++ implicit cast are very, very bad, but this doesn't mean that another language can't create sensible 'implicit cast' rules. No int <-> unsigned implicit cast but allow intY <-- intX implicit cast when Y>=X (same for unsigned).
I haven't seen a single SRE say, "oh, you use X language, with Y type system? Ok cool, you don't need us then"
When I write TypeScript, I have so few runtime errors that I'm often in disbelief. Sometimes I have zero runtime errors on my first test of an application.
Contrast this with my own writing of JavaScript (same coder, same runtime, etc.) and it's very clear. JS needs many, many unit tests that TS just doesn't.
Following your logic, why should we test our code before deployment?
Everyone should push to master and deploy to production. Things will break anyway even if we tested our code.
The goal is to catch as many bugs as possible BEFORE it runs in production. And even if you just catch a single bug, you have already won. In particular if the type system is balanced in such a way so there's not much additional effort involved (e.g. by using type inference).
For example ELM is a programming language that cannot crash due to it's type system. I mean it can crash off of exceeding system resources or through the FFI, but it cannot crash any other way.
The argument the OP is making is that the difference achieved by a type system is negligible, but I would argue that in some cases this is true, but ELM is a case where it's not.
A successful test only proves something for a single test case. In fact it is impossible to do the equivalent of type checking in the runtime code of your program unless the runtime code has the ability to self "reflect" on it's own source code.
There are also complex type systems that can prove correctness and completely eliminate the need for tests all together but this style of programming is really challenging and time consuming.
Primarily it has to do with making refactorings much, much easier.
The "new object" paradigm for instantiation (instead of Class.new, or some other generator), the static keyword, the final keyword, no global scope (eg for singletons), the inability to do non-inheritance mixins without 3rd party libraries, a full pre-runtime to bypass some of these issues in the form of Spring. Many of these work against testability and JUnit just can't keep up. It's horrific and it's not going to get better anytime soon.
Which is in line with the famous “No Silver Bullet” article.
When you have an outrageous amount of effort to (what amounts to manual) testing in closed environments before release, you're fighting the language but you just don't care.
> Empirically I would say that it does very very well on maintainability’s count
FAANG companies don't even release data related to how program defects are detected and addressed (publicly), so the assertion is literally baseless. Given enough time and effort, you can overcome the inability to do simple things. That doesn't make the tooling better, just the perception. This is common, historically, in tech (re Mechanical Turk).
Especially the global scope bit. Even statics fields are too much global scope in my opinion, not too little.
They harm testability for me, because I have to use java to test them. There's a larger discussion about "what is testing" and "how to properly test" which I'm not going to get in to here, but would be happy to talk about in another forum.
JUnit (and all the inferior alternatives) can't properly mock them or verify how/when they are called (despite running in a closed test container, which is how JUnit executes them). If you have ever tried to deal with native methods (or FAANG libs that like to use static final methods), there's a practical barrier to how much state you can avoid in testing. That barrier is java's ecosystem itself (JVM + syntax + tooling + ethos).
A large part of Java's verbosity is optional, but it feels to me that people do it because "That's how you write Java!".
But you don't HAVE to write an interface, abstract class, and factory for every class. You can just...write a class. You don't need to write a class that has nothing but an "invoke" function. You can just...write a function. You don't need 10 layers of abstraction for simple cases.
I've seen people do this (rewrite a Java app in JS) and then go completely off the rails, instead of, you know, writing it just like they would have (read: already did the first time) in Java. Bonus points when it's followed by the conclusion that JS sucks. It hints towards irrationality on their part. It's as if the structure and order imposed by Java is the _only_ thing keeping them on the straight and narrow, because they themselves can't be trusted to act responsibly and simply follow the same principles on their own when no one is around forcing them to.
But the reason I plan on going down this road is that I don’t want to program the complex business logic twice, it is quite important that both backend and frontend reliably calculate the same results. It is much easier to handle it through a compiler that will uphold Java’s semantics, than trying to write the same thing in JS - where you might easily do float arithmetic when you didn’t intend.
I object to this specifically as it relates to JS because the sort of code that mainstream "JS" programmers write nowadays ends up being passed off as idiomatic even though it really shouldn't. If you look at well-written JS in very large JS codebases with high standards (like Firefox, at least as it was circa ten years ago, or Apple's Web Inspector[1]), then you realize idiomatic JS looks a lot more like Java, C#, and Objective-C than what you find getting pushed to GitHub today—and (very soon after) discarded, since the half-life of an NPM-powered codebase is something like 2–3 years.
The problem is that the style of development associated with the NodeJS+NPM crowd has cannibalized the JS ecosystem, so almost anyone gazing in on the situation today is going to have a _very_ skewed view of what "idiomatic" means.
The perverse thing is that you will encounter more of an impedance mismatch with JS-the-language if you actually try to follow the crowd. This is a big driver for toolchain/framework churn—because they're all fighting the language instead of just going with it.
1. https://github.com/apple-oss-distributions/WebInspectorUI/se...
But that said, if starting from scratch, I think a Node/TS backend and a TS frontend is a much, much better choice.
This is such an odd claim. Mixing up a string for a list of strings, which both satisfy the iterable interface, happens all the time in python. Does the author, and his acclaimed fora simply not use python?
But those are actual valid and useful implementations of the interface, nominative typing wouldn't save you.
Not sure what scenario you're imagining here.
I'm including functions like min() there, technically it can take a string argument, but I've never wanted it to.
I cannot today find the bug, but as I recall an early version of Golang's stdlib included some HTTP handler that would "interface upgrade" your Reader to a ReadCloser and close it.
This could be bad if you expected functions that accept Reader to not call close!
We experienced this bug at Twitch and found it quite troublesome.
I think Structural is safer and better overall, and the best exponent is the relational/sql model.
What make it good is that structural types are more about the data itself, and if the code is around that concept, is very productive and safe in practique.
What make Duck typing unsafe is that is easy to "break" the type in runtime (ie: monkey patching, adding or removing things) and this is what I expect structural to retain it safety better...
How about a function like add()? It might mean something like BigInteger.add, or something like List.add, or Set.add, and these could all satisfy the same interface while behaving very differently.
For example, how did you accidentally get an instance of "Set" in the code path doing arithmetic on "BigInteger" instances, and why didn't that "Set" fail earlier during the other operations that are most likely being performed on the BigInteger (parsing it from a string, doing other arithmetic on it, etc)?
Because I got it as a return value from someone else’s function, and then this was the next thing I did with it.
It’s not hard to come up with everyday examples.
This is exactly the point of the article. Extremely extremely unlikely.
But if you have a real world example of this, I'm sure the article's author is interested.
I wasn't claiming it would. I was just saying Shoot() isn't exactly a common function name that would collide with anything to begin with, so it's a bit of a strawman as an example before you even need to list out multiple criteria. I'm pretty sure I've never written a Shoot() function in my life, let alone worried about it colliding with something. Compared with add(), get(), setValue(), read(), etc. which are incredibly common and thus more realistic candidates. That's all.
IME a lot of libraries that do duck-typing (for instance Python) make very little effort to explain what methods I'm supposed to implement or what the contract is supposed to be for those methods.
It also forces users of the API to explicitly state in code that they're intending to implement that interface, as opposed to just implementing a handful of seemingly random methods whose purpose will not be immediately clear to readers.
Also, AFAICT duck typing is incompatible with the approach taken in Rust or Haskell where the definition of a type can be decoupled from its implementation of various interfaces (called traits in Rust or type classes in Haskell), and where methods or attributes in different interfaces implemented for the same type can have the same name and not clash, because the object isn't the namespace for its methods. I strongly prefer this approach.
This has another interesting beneficial side-effect, when embraced: it often leads not just to better API design, but better design in general. Defining an interface necessarily means thinking about it. Which in turn generally leads to thinking about how it’s consumed, at least a fuzzy picture of the code implementing it. All of which can also percolate out into other parts of a system.
In scientific programming I've often wanted that, to push asserting particular constraints on values to the caller and/or the type system (whether that's compile-time or runtime type checks). But those constrained types would always match the interface of their unconstrained forms.
Needless to say, this code would be very verbose and ugly - what you would actually want is dependent typing, which solves this problem at the type system level.
Here is the implementation for Nat in Idris, for example:
data Nat =
Z
S Nat
With dependent typing, you can specify constraints on the values of function parameters, and have the compiler prove that only such values can be passed to those functions. For example, the following function only takes non-zero natural numbers, since the S constructor used in the parameter implies that the element can not be Z: addTwoNonZero : (S n) -> Nat
addTwoNonZero = (+) 2https://en.wikipedia.org/wiki/Design_by_contract https://en.wikipedia.org/wiki/Dependent_type
Contracts are ways to specify invariants for the runtime to verify. For example you could write a function with a contract saying that it takes a string containing all lowercase characters and returns a string containing all upper case characters. As far as the type system is concerned it just takes and returns strings but the runtime won't let a caller give you a value you don't want and won't let you return a value that doesn't match what you promised.
Dependent types essentially let you do the same but at the type level. For example you could have a function whose type says it takes a list of length N and a list of length M and returns a list of length N+M. Or maybe a type that represents a tuple containing a number N and a list of length N. Your code wouldn't type check unless the type system can prove the types. There's a lot of overlap with theorem provers in this space. Agda, Idris, and Coq are some decently popular languages in this category.
I find dependent types to be extremely powerful but sometimes the added overhead of figuring out how to encode an invariant into the type system is overwhelming.
declare class Branded<Meta> {
private meta: Meta;
}
type Brand<T, Meta> = T & Branded<Meta>;
type UInt = Brand<number, 'UInt’>;This is not a tradeoff, where I sacrifice development speed for safety. I gain speed, because types help guide me, and catch errors I would otherwise need tests for.
I've used JavaScript and TypeScript professionally for more than 7 years now and I'm still discovering new ways structural typing can break systems in novel ways. It's not like you don't benefit from the reduced need to type everything everywhere, but there are still holes that can leave subtle bugs in your system.