An Old Article I Wrote: "What To Know Before Debating Type Systems"
cdsmith.wordpress.com
cdsmith.wordpress.com
This is, of course, a scale - for instance, Python, generally viewed as a strongly-typed language, allows you to say something like
foo = 'lolsomestring'
if foo:
[code]
where a string is implicitly converted to a boolean (as opposed to using bool()). However, an exception is thrown when attempting to add 5 and "5", where a language like PHP will happily provide an integer result.In your view, it means this. In someone else's view, it means something else. Unlike "static", "dynamic", "explicit", "implicit", "structural", even "duck", all of which people more-or-less agree about (well, sort of), there's this problem with the definition of the "strong" or "weak" typing.
1) If you use them in conversation without a definition, there's a very high chance that they mean something different to the listener. (Unless of course it's a listener with whom you've already had this conversation.) 2) If you stop to define them, the whole conversation gets derailed to argue about the definitions.
That is why people say these words "have nearly no meaning at all". Does that help your confusion?
So, if you want to communicate about type systems, your communication will contain less misunderstanding if you simply omit the words "strong" and "weak". And for myself (although this may be too snotty a move for you), if someone else uses them, I refuse to hear them, and ask them to use different words and define those words.
My experience, in practice, is that when people say "strong" and "weak", I do indeed find that as cdsmith says, "strong" mostly corresponds to "it makes me comfortable".
But even people who are very comfortable in C, do not say that C is strongly-typed.
In my experience, people agree on which languages are strongly- or weakly-typed, and that can be used to reverse-engineer the definition, based on the properties those languages differ on. Assembly, C, and PHP are all weakly-typed, at the very least, and their proponents agree with this. They don't agree on why they are weakly-typed, simply that they are.
The original definition I ever heard for weak-typing, back when C and Assembler were the only things that had it, was "a weakly-typed language allows you to take the raw, in-memory representation of data of type A, and perform operations on it as if it were data of type B that had an equal raw, in-memory representation." The most famous example of this is Quake's fast invsqrt(), where the bytes making up a float have int operations applied to them, as if they were bytes making up an int.
Of course, PHP doesn't have this capability, but we still call it weakly-typed. So the definition must have moved on from its previous strict form. What does PHP allow you to do, that gets it called weakly-typed? This:
echo "(" + 5 + ")"; # prints "5"!
PHP, here, is taking the Strings "(" and ")" and interpreting them as numbers for the purposes of the addition operator. However, the strings "(" and ")" are not valid numbers—but PHP is fine with this, and uses the default numeric value of 0 to represent this invalidity.The similarity with the original definition, is that the data is transformed from one type to another, not implicitly, and not that the raw representation of the data is used, but rather without a guarantee that the datatypes have the same capacity for informational entropy. In other words, casting "(" to a number silently loses data, just as storing an int in a char variable in C silently loses data. This fact seems to be universal to all languages that get called "weakly-typed", and non-existent in any language that has never had that epithet applied to it.
It has nothing to do, notice, with whether a language will implicitly cast values—as long as all implicit conversions happen upward to types that have "room" for all the information in the original representation (char -> int, int -> String, int -> float) the language is still strongly-typed. And notice that the datatype itself being lost in a typecast (String -> Object) cannot be a reason to call a language weakly-typed because, in an interface-like cast like this, all the data is still retained, and its original form can be reclaimed simply by casting it back (Object -> String).
It's not your fault though, the problem is really just that 'strong' and 'weak' are just fundamentally unsound nomenclature.
As I said, we all know, and agree upon, which languages are "strong" or "weak", whatever those terms mean. Thus, the two words really do partition the world in some way; they do real Bayesian work. We just have to figure out what that work is.
> there's plenty of implicit downcasts that are still 'strong', like boolean tests on data bigger than 1 bit or gt/lt comparisons between different numeric types
Notice, above, that I said that explicitness isn't a part of the definition, as much as people like to think it is. In fact, it could be completely explicit at all times that you're casting (int)s to (char)s—and that would still be weak typing. The real property a type-system has that tells you that it is "weak" is that it allows casts between types that do not have valid surjections. "(" has no number it maps to—and yet the type system pretends it does. IEEE754 Infinity has no integer it maps to—and yet the type system says that's alright. (Note that in this special case, the hardware notices this loss of information, and sets a hardware exception flag.)
Operations retain strong-typing if they have well-defined one-way transformational characteristics, such as when casting (int) to (bool) using C's "!= 0" branching-requirement. "x % 6" compresses the field of integers into a set of six elements—but it does this in a way which maps every x to a valid new value on the "x % 6" ring. However, Infinity does not have a valid (int) representation—any value the computer does choose to represent it with will be "wrong."
To sum up: if, in your language runtime, there exists at least one type-transformation-implementation that goes like this:
try
a_entropy = f( a )
B.construct( a_entropy )
catch( DomainError e ) # A is not a well-defined B
B::SomeDefaultElement # who cares!
end
...then you have weak typing.The type-system property you're dancing around is soundness. A well-defined term that you'd know if you'd read the fucking article, which was written explicitly to teach it to you.
As for your example of what you consider 'weak' type-checking, I can see only two ways to make it 'strong'. The normal thing to do is a dynamic type-check resulting in a runtime exception if there's an invariant — in Python int(float("Inf")) raises OverflowError and int(float("NaN")) raises ValueError — would that satisfy you as 'strong'? The other option would be to use BigFloats and a fancy static type checker to ensure that you can never construct a zero value that could be processed as a float. Hilariously, GHC takes neither of these approaches: http://hackage.haskell.org/trac/ghc/ticket/3070 http://hackage.haskell.org/trac/ghc/wiki/Commentary/CmmExcep...
You missed a (perhaps subtle) point: the thing that has "well-defined one-way transformational characteristics" is an operation (an algorithm), not a type-system. A type-system can be thought of as a set of operations, a digraph specifying all the ways to go from each type to each other type in your system. This set has the property of being strong by default—but becomes weak if any of the operations is non-surjective. This is separate from a sound system, which is one that does not permit non-surjective operations.
For an example of the crucial difference: I can program in "a strong subset of C" (C with all the non-surjective functions/casts removed.) I cannot program in "a sound subset of C." There is no such thing.
You'll find programming language theorists using the phrase "a sound subset" with no qualms.
Point taken.
Python's type system uses interfaces, so strings are effectively also of type sequence, which is also of type boolean. Any object can just define __nonzero__() and be a boolean. That's all the bool() builtin does — it just tries to call __nonzero__() or __len__() — http://docs.python.org/library/stdtypes.html#truth-value-tes...
foo = 'lolsomestring'
if foo:
[code]
> where a string is implicitly converted to a boolean (as opposed to using bool()).Umm, no. You're assuming that the condition of "if" is necessarily a boolean.
It's not clear that it is. It looks more like "is the value not an element of [None, 0, 0.0, False, {objects of 0 length}]" and a couple of others. (There's a strange exception for certain objects.)
The fact that False and True are Python constants of a "boolean type" does not imply that booleans are privileged wrt conditionals.
Indeed. Conditional branching generally needs a boolean value - do we go down this path or not?
It's not clear that it is. It looks more like "is the value not an element of [None, 0, 0.0, False, {objects of 0 length}]" and a couple of others.
Right, and any of those things, when converted to a boolean are... False. The things that don't convert to True.
I don't know how the language itself is constructed, but it would surprise me if the code for bool() and if-statements were logically separated. Can you imagine the havoc that would ensue from, say, bool(None) returning False, but if None: causing the branch to execute?
There's a dozen more definitions, though.
String str = ""+5;
"" + 5
is defined by translation to new StringBuilder("").append(5)
which statically resolves to the method StringBuilder.append(int), whose implementation explicitly converts 5 to a string.Java does offer a counterexample to "strong typing means no implicit conversion", though:
String s = "foo";
Object o = s;
I just implicitly converted a String to an Object, which the GP's definition disallows.Of course you can fix up the definition for Java by adding the proviso "except upcasting a derived type to a superclass of that type", but the fact that you have to add a language-specific proviso indicates why there's no accepted common definition.
But at least with that proviso our definition works for all object-oriented languages, right? Well, actually it still doesn't even work for Java, because since Java 1.5 you can say:
int i = 2;
Integer i2 = i;
So Java needs an exception for upcasting, and for autoboxing, and we probably should worry about exactly what type we think 'null' has...Contrast this to something like Java where all code is littered with shit like foo != null and foo.length != 0.
There is only one thing I find missing — I think it would be prudent to avoid the use of "typed" or "typing" wherever possible. I much prefer to refer to properties being "typechecked" or systems using "typechecking" — it's otherwise quite easy it forget that the types are only manifest when and where they are observed.
My personal preference is to identify the language is question, which is usually the important and relevant point anyhow. Strong vs weak typing? Too vague. Haskell vs Python or Java vs C++, now we've got something we can sink our teeth into and we're not just spinning in space.