A polyglot's guide to multiple dispatch
eli.thegreenplace.net
eli.thegreenplace.net
This is why I spent a section on showing the "failed attempt". It's surprising how many folks don't understand the difference between overloaded functions (compile-time dispatch) and true multiple dispatch (at run time).
Visitor pattern always blows people's minds when they first see it in action hehe.
Would it be more surprising to Java newbies if it did work that way?
Multiple dispatch screams for the abolishment of methods as properties of objects, in favor of generic functions.
Youc an still have the syntactic sugar obj.foo(arg), but it basically just translates to foo(obj, arg) where foo is a generic function not associated with foo's class in any way, and obj and arg have equal status.
This generic function is what has methods.
Say we implemented "funky" numbers, and would like to implement multiplication of funky and real. When funky is the left argument we get funky.mult(real). If real is no the left, we get real.mult(funky). These functions do the same thing (the multiplication is commutative), yet ... they have to live in separate classes? It is nonsensical.
Under generic functions, we just have mult(real, funky) and mult(funky, real) which are methods of the generic mult(whatever, whatever). They both belong to the one generic function: that is the more reasonable organization.
One benefit is that the name `bar` is looked up in `foo`'s namespace, which means that it doesn't have to be in the current global namespace. This is a mixed blessing, but it's sometimes what you want.
These name spaces do not have to be conflated with program lexical scopes, whereby if you are a piece of code such and such a "blessed" function, you "see" the name space of the class as if it were lexical variable bindings.
You definitely don't want crud like different scopes in the program which are all working with exactly the same object, and want to access slot S, but they all refer to a different S because they are associated with a different place in the inheritance hierarchy, and that hierarchy contains unrelated slots named S at different levels. A given object must just have one S.
class Bar
def foo(baz :: Baz)
...
end
end
def foo(bar :: Bar, baz :: Baz)
...
end
could be in this hypothetical language that the former would be allowed to access private member variables on bar = Bar.new, while the latter wouldn't.The questions then are:
a) Is private state something we want?
b) Could private state access control be implemented in a different way (Lexically? Package level? Additional annotation?)
The idea of friend functions of a class could be adapted from C++.
It could work dynamically like this.
Firstly, given obj.slot or obj.fun(); if the type of obj is not known in the given scope, then these accesses are disallowed unless slot and fun are public.
Secondly, a class can be (dyamically) associated with a (dynamically extensible) list of friend generic functions.
When a new method is being compiled for a generic function, the compiler uses the class information about the argument (which specializes it) to check the friend list of that class. Then inside the friend method, obj.slot and obj.fun() are allowed even if slot and fun are private.
(Friendship of compiled code couldn't be revoked; once a method is processed by at least the expanding code walker, if not compiled, then it cannot lose access even if the generic function is removed from the friend list of that class.)
Firstly, obj.member could be disallowed if
Multiple dispatch screams for the abolishment of methods as properties of objects, in favor of generic functions.
Youc an still have the syntactic sugar obj.foo(arg), but it basically just translates to foo(obj, arg) where foo is a generic function not associated with foo's class in any way, and obj and arg have equal status.
This generic function is what has methods.
I think this comment sums it up well: https://news.ycombinator.com/item?id=11528588
P.S. I love Seibel's PCL - great book
Naturally, these approaches require intrusion (special field/static member in each class)
I'll have to re-read that section of Modern C++ to recall the runtime algorithm, but I'd be interested on any pointers to newer research or engineering on dispatch techniques.
It's too bad that (AFAIK) no 'mainstream' high performance language has a multiple dispatch abstraction that could motivate more work.
The cheating comes from the run-time compilation features being leveraged to inspect "already compiled" code and dodge around the dispatch. However this causes problems in larger systems (Julia has some serious shortcomings when it comes to large projects, and they ban people who talk about it too much) because if you add a method to the multiple dispatch already compiled functions won't make use of it without recompiling them too, this is basically the expression problem being displayed in a dynamic language (because they choose speed over true dynamicism; they leverage the dynamic features for a different purpose and loose the feature typically gained). This is a very bad form of cheating because multi-methods are often seen as a solution to the expression problem, but Julia still has the problem even though they use multi-methods.
Try the following:
# Library
abstract bar
ex(x) = println("dynamic")
ex(x::bar) = println("bar")
test(x::bar) = ex(x)
type foo <: bar
x :: Int
end
test(foo(5))
# Now our code
ex(x::foo) = println("foo")
test(foo(5))
It should print `bar\nfoo` it doesn't.https://github.com/JuliaLang/julia/issues/265
Look at the commits at the bottom.
The current behavior, although theoretically a problem, is not a big issue in practice. The reason is simple: unlike the example above, real-world programs tend to define a bunch of types and methods first and then the main program runs, using those types and methods.
Where it is actually a problem is in interactive development. At the REPL, people tend to redefine the same methods over and over, and it can be quite bothersome that those changes aren't always reflected. You can force an update by redefining the calling method as well, or you can just restart the REPL. Annoying, but not the end of the world.
Also real systems have many libraries. Which will each take turns in defining things. And some may run initialization code, causing certain things in their library to run. The fact of the matter is you don't have "real" multi-methods because of it, and it's a problem with making large systems.
Lets see, there is a logging library with a custom string formatter, in another library I specialize the string formatter, and I log that the library started up with a value based on that abstract tag. Now as a user I specialize a concrete type for logging, but it doesn't work, why? That's a weird error, and also, the expression problem via multi-methods, well done you guys implemented multi-methods that don't solve one of the core purposes of multi-methods.
How much you wanna bet I can find startup code in many python libraries which calls code that is later specialized by users. Edge cases are where it's important.
I'm wondering what this performance cost will be and if significant, can all that brainpower at MIT etc really not find a way to mitigate it?
This is a practical question for me because I'm considering Julia for a project and my personal data language.
This is the upcoming python alternative ecosystem, also with multi dispatch:
https://github.com/libdynd/libdynd https://github.com/numba/numba https://github.com/blaze/blaze
Just don't expect to be using those shiny numbers for anything but scientific computing (and most specifically when involving large matrices). Their namespacing and modularity is terrible, writing a complex gui, non-trivial web server, 3d renderer, or any sort of real time system is right out. And they seem dead set against changing that (did I mention they ban people for making too much noise about it). But if you want to simulate something over a cluster it's a great choice.
I should probably note that I have a poor opinion of their community from my observations (they are hilariously salty), I may be a bit biased.
There are many dozens of open issues discussing namespacing, interfaces, levels of type abstraction and extensibility, garbage collector improvements, i/o latency problems, task-switching performance, and more.
Prioritization of core effort in certain directions does not mean other issues are suppressed or even forgotten. Considerate proposals in other areas are welcome, and pull-requests implementing such proposals will be gratefully reviewed. They may not always be accepted -- immediately or sometimes ever. There is certainly some conservatism about adding features (such as interfaces), because features have costs and are near-impossible to remove once added. But other things just take time to get merged; for example, the generational GC took over a year, and there are other very large proposals that have been pending for longer than that because the implementation is lacking in some way or the implications are not fully worked-out to the point of consensus acceptance. That's ok, and IMHO healthy.
Other people seem to agree on that point: http://danluu.com/julialang/
As I said, the people actively working on the compiler have to prioritize because time and resources are finite. This is not the same as anyone being "dead set against changing", or a "ban people for making too much noise about it".
> when did any of them last move.
Here's a sampling of some improvements over the past 4 months.
- Debugging, everyone's priority by astronomic margins, has seen massive work and a recent preview-release (I've been using it since the announcement: rough edges, but very solid core): https://github.com/Keno/Gallium.jl/commits/master https://github.com/Keno/ASTInterpreter.jl/commits/master https://github.com/JuliaLang/julia/pull/15859 https://github.com/JuliaLang/julia/pull/15444
- Revamp of closures to make anonymous functions as fast as any other: https://github.com/JuliaLang/julia/pull/13412
- Significant compiler performance improvements: https://github.com/JuliaLang/julia/pull/15300 https://github.com/JuliaLang/julia/pull/14543 https://github.com/JuliaLang/julia/pull/15609
- Work towards static compilation (a number of PRs): http://juliacomputing.com/blog/2016/02/09/static-julia.html
- Multithreading improvements: https://github.com/JuliaLang/julia/pull/15917 https://github.com/JuliaLang/julia/pull/15106
- Someone else pointed you to recent commits toward fixing the recompilation issue.
I'm not sure why Julia itself should make it difficult to "do the right thing" when exceptions occur...
Real-time 3D rendering is a big area. Will we have virtual reality with ultra low latency in the next year? Definitely not! Can we have 3D CAD/scientific visualizations which hardly stutters? Definitely! Our library is already very smooth and we didn't even start optimizing yet. That's something that seems to be almost impossible in e.g Python (without just calling C for everything).
So while most of these arguments are currently spot on, I do think that there is nothing fundamentally wrong with Julia. And while the maturity is not there yet, you can already be very productive with Julia and build for the future.
Untyped, sometimes incorrect, exceptions, https://github.com/JuliaLang/julia/issues/12485 and related.
I want to keep my general code and scientific code in the same language.
So python it is.
Only one person has ever been banned on GitHub and it was not for talking about any shortcomings. That person is still allowed to and regularly does post on Julia mailing lists – no one has ever been banned from mailing lists.
https://groups.google.com/forum/#!forum/julia-users
Here's a recent (and typical) example of a lengthy series of technical and sometimes pointedly-critical questions:
https://groups.google.com/forum/#!searchin/julia-users/%22ve...
The naive implementation would be:
def process():
A()
B()
if [conditions for process 2]:
C'()
else:
C()
D()
You should implement it as: class Process1:
def A()...
def B()...
def C()...
def D()...
class Process2(Process1):
def C()...
[put C'() code here]
def process():
right_process = figure_out_right_process()
right_process.A()
right_process.B()
right_process.C()
right_process.D()
process() becomes a lot easier to test, since you only need to mock out/outsmart figure_out_right_process() instead of mimicking the right logic for [conditions for process 2]. You can test each process independently, since you don't need to fulfil the [conditions for process 2] in each test, and if those conditions should change, you only need to change figure_out_right_process() instead of at each branch in your code.It may look verbose with only one branch, but the second you have more than two processes, or nested branches, the class based way looks a lot cleaner.
2) For symmetry, I would make C abstract in the parent class, and create two subclasses, one implementing C, one implementing C'.
3) Having done that, the logic defining the order of the steps can be moved into the parent class.
4) If you want to mock things, it may help to make A, B, and D overridable in the parent class, too.