Let the Type System do the Work
javadocmd.com
javadocmd.com
I realize that variable naming and type systems are two different things, but it seems many programmers never realized the point of type systems (expressing the type for the programmers sake) because they only ever saw it used to distinguish abstract primitives. For a long time, I had trouble understanding that hungarian notation wasn't a type system, because they seemed to do precisely the same thing - that's how limited my understanding of types was. The kind of style seen in this article was alien to me for a long time, but it was enlightening to realize that the type system is there to help me out, not just make me type a bunch of unnecessary crap.
tl;dr: I really need to learn Haskell.
float x;
and this: kilogram x;
The former, as you say, only tells you about the underlying representation. It says "this is a float, so the computer should store it in such and such a way." That's fine, but it doesn't tell us enough.The latter example is far more useful. In my imaginary language, the declaration implicitly tells the computer to store the value as a float, because the kilogram type has been defined as such elsewhere. But that's not all it does! It tells us and the compiler that this float represents a real-world quantity measured in kilograms. It prevents us from mistakenly passing kilograms where a pounds were expected, or seconds where kilograms were expected.
On the Haskell front, you might be interested in the Dimensional library, which does just that. It also works elegantly with multiplying and dividing units. E.g. if you have a miles value and an hours value, you can divide to get a miles per hour value.
Thanks for the pointer - I'll be sure to check out the library as I dive into Haskell over the next few months.
So that's why so much Java code ends up looking like this:
//class FurryThing implements Thing {}
//class RubberThing implements Thing {}
FuzzyThing furbie = new FuzzyThing();
RubberThing bouncyball = new RubberThing();
DoStuffToThing( Thing(furbie) );
DoStuffToThing( Thing(bouncyball) );
In fact, it's not uncommon to see things like class Furbie extends FuzzyThing, so you end up with very specific types that don't necessary behave differently, but do add some supposed semantic precision to the code. Unfortunately, you don't get much benefit from it in Java, so it just ends up being a hassle, resulting in lots of programmer-forced unsafe casting and just as many runtime errors as you'd have got with simpler types. This is not a fault of the language, just poor implementation which rightly or wrongly appears to be idiomatic in Java.The type system in, for example, Go is much simpler, and allows you to create semantic types that actually are the underlying type instead of creating a class wrapper (called boxing in Java, e.g., boxed int is class Integer, which is an object containing only an int), which means you can define a type like 'kilogram' that's really just a float, but the type system won't silently cast a 'kilogram' value to a float or vice-versa because it knows they are different things.
I haven't got into Haskell in much detail yet but it just basically takes that a few steps further and allows you to define the characteristics of each custom type based on axioms, that describe things like whether it's commutative or associative, the range of allowed values, etc.
This kind of axiomatic semantic typing is clearly a lot more powerful, but as yet it hasn't really found its way into mainstream languages yet. Go takes the view that it's too complicated for what is intended to be a low-level language. Rust looks interesting because it aims to be equally low level, but provide some of the same rich typing you'd find in Haskell. Unfortunately it's far from stable, but it'll be interesting to see where it ends up.
I think that people hate the implementation and syntax associated with semantic type in Java, not the fact that semantic typing is widely used in Java.
I think both much of the recent golden age of dynamic languages and the more recent resurgent of cleaner, less-heavy-syntax statically typed languages has been motivated by the perception that static typing in Java (and similar languages) has too high a cost for the benefits it provides.
Luckily, more and more programming languages are making creating these simple types easier. Haskell:
newtype Kilogram = Integer
Rust: struct Kilogram(int);
In fact, this blog post, while a bit outdated, shows an application of this idea, to solve string encoding issues for HTML templating: http://bluishcoder.co.nz/2013/08/15/phantom_types_in_rust.ht...Which, thankfully, don't add any runtime cost (as should be expected).
As an aside, you don't actually need to typedef them at that point, but it saves you having to write "struct ..." everywhere so it's typically worthwhile.
newtype Kilogram = Kilogram integer
The issue with doing this of course, is that it's mostly useless to computation. You lose the ability to use mathematical operators because you're no longer an instance of Num, and even if you create the Num instance, or use -XGeneralizedNewtypeDeriving, you can't multiply a Kilogram by a Meter/Second for example, since the arguments to (*) must be of the same type. One would need to use a generic "Measure" type instead, where the unit is some metadata attached to it, and the Num instance implements the typechecking on units.As a user of a library/API I need to know if I can rely on it being precise or not.
I'm not certain, but it would appear Fahrenheit support is not built-in. There's a module called NonSI:
https://hackage.haskell.org/package/dimensional-0.13/docs/Nu...
I would have expected Fahrenheit to be in that module if it were supported.
But I presume you could add Fahrenheit yourself, if you were willing to learn a bit about the library's internals.
For gory details: https://en.wikipedia.org/wiki/Affine_transformation
To say it in more detail: To transform between Celsius and Kelvin, you just translate. (That operation is supported in Dimensional.)
To transform between Fahrenheit and Celsius or Fahrenheit and Kelvin, you translate and scale.
We know that Dimensional can scale, because it supports things like miles to kilometers. We know it can translate, because it supports Celsius to Fahrenheit. Is it the combination of scaling and transforming that makes Fahrenheit impossible?
f = m*c + b, which is both a scale and a transform.
f = m*c + b
where m is simply equal to 1. Why wouldn't Fahrenheit conversion be the same thing, just with different constants?
Yeah, you're correct.
--------
Per [1] it seems that since both would be named types you wouldn't be able to assign one to the other without explicit conversions. If my reading is correct, then that addresses my concern. It's things like storing a `float` to a `inch` type that can happen without explicit conversion as long as `inch` has float as its base type.
[1] http://golang.org/ref/spec#Properties_of_types_and_values
type synonym/alias.
newtype Kilogram = Kilogram Double
which has no runtime overhead but does enforce the fact the Kilogram is a distinct type. {-# LANGUAGE GeneralizedNewtypeDeriving #-}
newtype Kilogram = Kilogram Double deriving (Show, Eq, Ord, Num, Real, Fractional, Floating, ...)
And remind ourselves that Haskell is really quite expressive.Declaration:
struct Kilogram {float v;};
Instantiation: Kilogram x = {1};
or Kilogram x{1};https://gist.github.com/bitemyapp/8739525
And as the other commenter mentioned, Dimensional is indeed a cool library :)
When building code for the long term, you really want to focus on techniques that make it impossible to misuse an API, whenever it's possible to do that. This article nicely calls out one such technique.
Generic typing can be used to (try to) force the consumer of an API into correct usage. Judicious use of final and abstract in the land of Java can be used force/guide eventual overriders of an abstract class down the right path -- they'll get compilation errors if they don't at least implement the right methods.
Whenever you find yourself writing an assert method, take a beat and figure out if the type system could have turned that into a compilation error instead.
Compilers know a lot more than just what is represented in "the type system" (as it's typically thought of, anyway). Control flow (basic block analysis, &c) for instance. Which isn't really taking away from your larger point - certainly, things represented in the type system are things the compiler will know about.
"Whenever you find yourself writing an assert method, take a beat and figure out if the type system could have turned that into a compilation error instead."
Possibly with a _Static_assert (or static_assert) as if C11 (/C++11)! Obviously there are substantial limitations still, but it's great to be able to pull more to compile time cleanly.
Unless I have to provide support, add features, scale my application, write bug free code, hire additional developers, or explain my thought process to anyone else.
> No problems here. The first noticible crack in the system comes after a few weeks vacation away from this code. You come back, and you want to change the starting point of the player. Will you remember which Vector2 to change?
player = Player(position=Vector2(100, 100), size=Vector2(50, 50))
The problem is not typing here, it's lack of named arguments.For example when you have one function returning (pos, size) and another function expecting (size, pos).
>>> from collections import namedtuple
>>> Vector2 = namedtuple('Vector2', ('x', 'y'))
>>> PlayerOptions = namedtuple('PlayerOptions', ('position', 'size'))
>>> opt = PlayerOptions(position=Vector2(100, 100), size=Vector2(50, 50))
>>> Player(**opt._asdict())
The order arguments are passed doesn't matter anymore. You can also enforce this calling convention with a constructor signature like this (in 3): >>> class Player(object):
>>> def __init__(self, *, size, position):Both of the following would compile successfully using named arguments but one of them is semantically wrong:
Player(pos = v1, size = v2);
Player(pos = v2, size = v1);
Using a compiler's static type checking enforces correct semantics more than named arguments. Of course, this only works if the programmer creates new differentiated types so that they are no longer identical (which was the crux of the article.)
Dynamically typed languages are expressive, quick and fun.
But strict static typing is like a seat belt, it's annoying, but it just might save your life.
Also ceremony code is there for a reason; that reason is not to annoy you, it is to make sure that those who come after you know what it is that you have done and why.
Also: comments.
Also: documentation.
Also: UML.
The problem with saying "static typing" without further precision is on one axis it ranges from C where the steatbelt is made of paper to ATS or Idris where the "steabelt" is a zero-zero ejection seat, and on an other axis it ranges from Java where you have to braid the steatbelt from raw fibers any time you sit down to MLs or Haskell where the seatbelt magically appears around you.
Well, only if you don't need nominal subtypes.
I will say though that newer models of OOP, focusing on interfaces and composability of behavior, seem to have quite a bit of steam left in them.
[1] https://news.ycombinator.com/item?id=7618933
[2] http://raganwald.com/2014/03/31/class-hierarchies-dont-do-th...
Regarding C, it is weakly typed so it doesn't care whether you feed it apples or habanero chillis, the static type system isn't for safety, it's for improving compilation speed (which was often a big headache in the 1970's).
It sticks with the methodology of C: if you make a mistake you pay for it dearly.
For example, kilogram as wrapper around float?
newtype Foo = Foo Int
map (\(Foo x) -> x) listOfFoos
which should be a no-op at runtime because the representation of Foo x and x are the same, but due to the newtype, the traversal does need to occur. [1] discusses a way to avoid these problems. Basically these newtype should only ever be needed during type checking, and should really be erased after that; once you've type checked everything, then you know your program won't ever treat something which is a non Foo Int as an Int, so you can then get rid of the compile time tag and gain better optimisations.[1] http://research.microsoft.com/en-us/um/people/simonpj/papers...
As an aside, Scala does allow for compile time only small types in the form of AnyVal's. They won't help in the examples in the article, but in the simpler kilogram wrapper around float it will and there will be no performance penalty as it will be elided by the compiler.
You can always find the slow part and change it to use raw floats/ints/etc.
But as people already pointed, in practice compilers are never optimum. When it's important (that is, rarely) it's good to test, or at least read the compiler's manual.
On the other hand, if you give data to a C-programmer, they can locate its position in RAM. Give an assembly-programmer a simple data-type, they can point to a specific register(s) in the processor.
Your codebase will "say" more, and therefore your tools will "know" more, and therefore your tools can do more for you.
Also, although it makes the code much more verbose, named parameters (as used in, for example, Objective-C) fix any eventual bug with refactoring methods' signatures, right?
I think a good rule of thumb is to ask yourself what the "dimension" or "unit" the parameter has. (See also: Dimensional Analysis.) You never want to pass 4.33f radians into a function expecting 4.33f newtons!
A few more examples:
Distance in meters, Distance in feet, Speed in m/s, speed in yards/minute, Acceleration in m/s^2, Acceleration in cm/s^2, Mass in grams, Weight in pounds, Force in newtons, Force in dynes, Temperature in Celsius, Temperature in Fahrenheit, Angles in radians, Angles in degrees....
"But I don't do physics programming!", you say? Well, there's plenty more:
Time in seconds, Time in milliseconds, Distance in pixels, Distance in inches, Numeric ID of a Foo, Numeric ID of a Bar, Bytes in UTF8, Bytes in ASCII, US currency in dollars, US currency in cents, Euros, interest rate (yearly), interest rate (monthly)...
And that's not even getting into industry-specific dimensions that might exist.
(Again, asked another way, is it ever a problem to create a "small type" if it follows that rule? For some codebases it might mean having a "small type" per unique parameter on signatures.)
IMO the biggest problem isn't having too many distinct types, because they represent information you'd need to track manually anyway, like knowing that $time is seconds rather than milliseconds. I'd be more worried about:
1. The risk of conflicts between different groups of people who both defined their own type called called "Newtons" and want their libraries to work together.
2. The sliding-scale of inconvenience if your implementation gets in the way of the mathematics, like having to call `foo.asFloat() * 2` rather than `foo * 2`. (You can relax things a bit by relying on best-effort annotations.)
That brings up an interesting issue -- is the "right way" to do that to have different types for these, or one type for distance with different factory function/methods for different units. E.g.: should "meters" and "feet" be different types, or should the "measurement" module have both a "meters" and "feet" method that return the same ("distance") type.
It seems to me the latter is more conceptually clean.
OTOH, math might be easier to express with the former (particularly in a language that allows you to define implicit type conversions, so that if you pass an "inch" value to a function expecting a "cm" parameter, it gets automatically converted to the appropriate "cm" value.)
Most built-in type-systems (sensibly!) only try to tackle the broadly-applicable computational-type, since the rest is context-specific.
However, in a contractual sense, all of those are important, if any are not as expected, you'll get bugs.