Building a Language and Compiler for Machine Learning
julialang.org
julialang.org
you mean given some data that's modeled by an ODE you want to fit the ODE to the data (and therefore discover parameters of the ODE that would have produced that data) ?
I thought I was having some kind of stroke or terrible deja-vu (didn't I read this comment earlier this morning?) until I realized you copy and pasted your comment from 20 hours ago on https://news.ycombinator.com/item?id=18594103
I'm also not entirely sure what is going on under the hood with Tracker types, and the documentation is not that great, which became a problem when I was trying to chase down errors in something really custom I was doing.
I much prefer Knet's way of autodifferentiating, which is more intuitive to me, but Knet's layering doesn't feel as nice as Flux's.
I really wish GPU computation in Julia had a different semantic - by making it a 'virtual computational node', accessible using Distributed module with the same semantics as a totally separate node. That would really make async distributed batch processing a thing, the system could profile all the nodes in use and if we really want to get fancy be able to use something like JUMP to make best use of the processing power available to it.
M = [1 0 0 0
0 1 0 0
0 0 1 0
0 0 0 1]
v1 = [1,2,3,4] (column vector)
v2 = [4,5,6,7] (column vector)
you can either do (M * v1, M * v2) on two cores as r1a = [1 0 0 0
0 1 0 0] * v1 (core 1)
r1b = [0 0 1 0
0 0 0 1] * v1 (core 2)
r2a = [1 0 0 0
0 1 0 0] * v2 (core 1)
r2b = [0 0 1 0
0 0 0 1] * v2 (core 2)
then r1 = flatten([r1a, r1b])
r2 = flatten([r2a, r2b])
OR, you could just have the whole model in each core: r1 = M * v1 (core 1)
r2 = M * v2 (core 2)There are also various autobatching packages being developed.
Regarding the GPU semantics, wouldn't that be solved by simply using a distributed array of GPU arrays?
Since flux is lightweight, generic, modular and pure Julia, these things can be developed in third party packages.
No, I don't want to necessarily have distributed GPUs, I want to treat a GPU as a distributed compute node. As in "the GPU is a remote machine that I can send julia code to" (this is how julia normally treats running on clusters, or even on multiple threads).
Also, what does an approach like bucketing even look like for the approach that Julia is taking? The idea there of course is to have 'slop': to combine many similar examples whose tensor sizes differ by small amounts, and to carefully define all your primitive operations such that they can can ignore the padding used to combine similar tensors into a uniform shape. Doing this requires awareness of the tensor sizes all the way back to the way you sample the training data, so I don't see how compiler magic can achieve the same performance as you get from bucketing.
Of course, bucketing becomes more complex for things like trees and graphs and other higher-level objects. And bucketing, theoretically, can bias into your gradients, if there is any correlation between the gradient of an example and its tensor shape.
"Automatic Batching To get the most from these accelerators – which can have significant overheads per kernel launch, but scale very well over input size – it is common to batch programs, applying the forwards and backwards passes to multiple training examples at once. In simple cases, such as with convolutional nets, it’s simple to handle this by concatenating, say, 10 images along an extra batch dimension. But this task becomes much harder when dealing with variably-structured inputs, such as trees or graphs.
Most researchers address this by taking on the significant burden of batching code by hand. Different solutions have been proposed for different frameworks (DyNet, TensorFlow Fold, which heuristically try to batch some high level operations together when possible, but these typically either have their own usability issues or do not achieve the performance of hand-written code.
We suggest that this problem is identical to that of Single Program Multiple Data (SPMD) programming, which has been well-studied by the language and compiler community for decades, and becomes visible in more recent approaches to batching like matchbox. Indeed, it is very similar to the model of parallelism used by GPUs internally, and has been implemented as a compiler transform for the SIMD units of CPUs. Taking inspiration from this work, we are implementing the same transform in Julia to provide SPMD programming both for scalar SIMD units and for model-level batching. This allows us to reach the ideal of writing simple code that operates on individual samples, while still getting the best performance on modern hardware."
where is logging, where is model storage and versioning, where is input data processing and normalizing, where is results processing?
the hard part of ML stacks is AD and GPU not all of those other things (i'm sure there has been zero cutting edge research done on better ways to log).