Yanni – An artificial neural network for Erlang
blog.ikura.co
blog.ikura.co
I am unsure why anyone would use Erlang for number crunching. Training neural nets is basically just multiplying big matrices. I was hoping this project would come up with an interesting approach (how about using SIMD on the binary comprehensions that can use it? now that would be cool) but performance / memory usage does not seem to be looked at here.
It is naive / uneducated to think that "Erlang’s multi-core support" + distributedness will enable many things for you. How does the VM scale on 32, 64 threads? Have you tried making a cluster of 50+ VMs? Unfortunately Erlang Solutions Ltd.'s marketing has hyped many.
I am not against projects like these, I am just looking for reasons behind the choices made.
Some people still seem to get very upset when someone proposes that some langauge is not suitable for some use, but there aren't any languages that are the best for everything. The languages in which I would want to write heavy-duty number crunching code will be nowhere near as good as Erlang at writing high-concurrency, high-reliability server code.
Also, to avoid a second post, contrary to apparently popular belief it is not possible to just make up for slow code by bringing in a lot of CPUs. Number crunching in Erlang is probably a bare minimum of 50x slower than an optimized implementation, it could easily be 500x slower if it can't use SIMD, and could be... well... more-or-less arbitrarily slower if you're trying to use Erlang to do something a GPU ought to be doing. 5-7 orders of magnitude are not out of the question there. You can't make up even the 50x delta very easily by just throwing more processors at the problem, let alone those bigger numbers.
You're right. Please never implement your number-crunching in Erlang. It will be slow.
That's exactly what dirty schedulers were made for - run longer blocking C code but without having to do the extra thread + queue bits yourself.
> is not possible to just make up for slow code by bringing in a lot of CPUs.
It entirely depends on what you are doing. So number crunching could be a small part amongst lots of protocol parsing and binary matching and sending to different backends and so on. Rarely it is just purely a simple executable that runs and multiplies a matrix and exits. In that context it could make sense to start with Erlang and then do dirty scheduler or an drivers or such for number crunching.
In the context of NIFs, long running could mean 100ms. The NIF API supports threads so you can always pass work to a thread pool and only block the scheduler thread for as long as it takes to acquire the lock on your queue. Or use dirty NIFs which I think are no longer experimental. There's also a new-ish function you can use to "yield" back to Erlang from a NIF but that kinda makes me nervous.
* Use dirty schedulers. Those are available since 17 and in 20 are enabled by default
* Use a linked in driver and communicate via ports.
* Use a pool of spawned drivers (not linked in) and send batches of operations to them and get back results.
* Use a C-node. So basically implement part of the dist protocol and the "node" interface in C and Erlang will talk to the new "node" as if it is a regular Erlang node in a cluster.
* Use a NIF but with a queue and thread backend. So run your code in the at thread and communicate via a queue. I think Basho'd eleveldb (level db's wrapper) does this
* Use a regular NIF but make sure yield every so many milliseconds and consume reductions so to a scheduler looks like works is being done.
I've talked to a few people from the financial world who to trading with Erlang. They use it because it makes to easy to take advantage of multiple cores.
And like others pointed out, if there is a need to optimize things, they can quickly build a C module and interface with it but they didn't need it so far.
Of course that's fine for many workloads.
I probably underestimate the likely cost by several times and then the cooling would be a great science fiction set properly to 1:12 scale, but I certainly know businesses who have a real desire for a product like that.
Am I missing a showstopper preventing the possibility? I'm not going to be persuaded that it couldn't be done by mere impracticalities. I'm quite prepared to take heatsinks the size of Cantelupes...
This is why typically you see them adding new cache levels, instead of drastically expanding the size of the cache, especially the lower level caches (L1 and L2).
I noticed a paper about a cache-aware work-stealing scheduler which I have not yet read[0].
[0]: https://www.researchgate.net/publication/260358432_Adaptive_...
I ended up rewriting key components in C# just to speed it up and make maintenance bearable.
Unless you are a non-programming researcher that learned labview and needs to work on fpga, you should just use something else.
IMO it's worth learning systemverilog even if you're in that position; labview has so many 'gotchas' and is so gross for anything large, I think it is never the right answer.
At the end of the day, anything that lets you bang away uninterruptedly on a processor (no context switches, no cache shenanigans) seems like a suitable implementation.
And of course you get to write in a fun language that is amazing for other use-cases.
[1]: https://github.com/mochi/mochiweb/blob/master/src/mochigloba...
Good point, I wouldn't use it in production. But I think it can be a great educational tool to learn about implementation of NNs and overall topology.
(Handbook of Neuroevolution Through Erlang) https://www.springer.com/us/book/9781461444626
You're like huh. There's people that's doing it. I don't know how but yeah. I can't even crunch prime number on it without giving up cause of how slow it is.. I guess there's more to it.
I'm in no way an expert, but I work in Erlang in my day job and just glancing at the repo, this solution can't possibly be performant. A) Erlang is slow at math. B) Arrays don't have O(1) access(ETS tables might be able to help with this). C) You can't scale this solution with more Erlang nodes(without some additional work).
I really like Erlang and want to evangelize it but I don't think this is a good way of doing it. I only see this as a neat toy but not a selling point for using Erlang..
As a side note: I noticed the repo has a feature note about adding NIF's for performance bottlenecks (native C code for Erlang to talk to). If you end up writing C code, then what are you gaining from Erlang?
src/yanni_trainer.erl:78: type array() undefined
It was introduced in revision 20. 3 files updated, 0 files merged, 1 files removed, 0 files unresolved
[bherman@archy yanni]$ hg update 19
[bherman@archy yanni]$ make
rm -f notp notp.boot
rm -fr ebin
mkdir ebin
erlc -o ebin src/*.erl deps/*/src/*.erl
src/yanni_lib.erl:77: Warning: random:uniform/0: the 'random' module is deprecated; use the 'rand' module instead
src/yanni_trainer.erl:112: Warning: random:uniform/0: the 'random' module is deprecated; use the 'rand' module instead
deps/*/src/*.erl: no such file or directory
make: *** [Makefile:6: default] Error 1
[bherman@archy yanni]$ hg update 20
3 files updated, 0 files merged, 0 files removed, 0 files unresolved
[bherman@archy yanni]$ make
rm -f notp notp.boot
rm -fr ebin
mkdir ebin
erlc -o ebin src/*.erl deps/*/src/*.erl
src/yanni_lib.erl:77: Warning: random:uniform/0: the 'random' module is deprecated; use the 'rand' module instead
src/yanni_trainer.erl:78: type array() undefined
src/yanni_trainer.erl:118: Warning: random:uniform/0: the 'random' module is deprecated; use the 'rand' module instead
make: *** [Makefile:6: default] Error 1
[bherman@archy yanni]$ make
edit: formattingFrom Erlang release notes:
The pre-defined types array/0, dict/0, digraph/0, gb_set/0, gb_tree/0, queue/0, set/0, and tid/0 have been deprecated. They will be removed in Erlang/OTP 18.0.
Instead the types array:array/0, dict:dict/0, digraph:graph/0, gb_set:set/0, gb_tree:tree/0, queue:queue/0, sets:set/0, and ets:tid/0 can be used. (Note: it has always been necessary to use ets:tid/0.)
It's funny that's what they did with disco (http://discoproject.org/).
Python + Erlang = Hadoop 1.0 like.
This library uses one "process" per neuron for concurrency. Processes are extremely lightweight and entirely unrelated to system processes or threads.
Running one process per neuron would actually be a very efficient way to do it.
The recent article from Discord [2] also mentioned "Sending messages between Erlang processes was not as cheap as we expected, and the reduction cost — Erlang unit of work used for process scheduling — was also quite high. We found that the wall clock time of a single send/2 call could range from 30μs to 70us due to Erlang de-scheduling the calling process."