Tinn: A tiny neural network library written in C99
github.com
github.com
kann is a stunning piece of engineering, and your code is magnificent, and that is precisely why I mentioned it. I apologize, I did not clearly express my intent.
Passing small structs by value isn't as expensive as it once was, modern ABIs pass the initial fields in registers. However, this struct does have ten pointer/integer fields, and I'm not aware of a mainstream ABI that allows that many. x86-64 allows up to six on Linux, IIRC. A quick check of the AArch64 ABI says it takes up to eight. So yeah, there will be stack traffic involved if these function calls are not inlined.
There may be cases where by-reference (yes, that's an absolutely acceptable thing to say, even in C) is more efficient: If the struct is a kind of "this pointer" (as in object-oriented programming) that is passed to a bunch of functions but relatively rarely used, then passing it by reference causes much less register pressure. Passing it by value would mean occupying a lot of registers with values that are probably not accessed.
On the other hand, if you expect the callee to do actual computation on all or most of the fields of the struct you pass in, passing it by value should probably be the way to go. This way, the values to be processed are already in registers and don't have to be stored to the stack by the caller and loaded back from the stack by the callee.
There is no general answer that is uniformly best, it depends on the characteristics of the actual program and their interactions with the compiler's optimizations. With inlining in particular, which should often allow the compiler to generate the same code for both versions.
I don't want to be nit-picking. But I would like to know why "by-reference" is an acceptable thing to say.
When I learnt C, C++, Java, etc. I was repeatedly told that pass by reference exists only in C++. In C, a pointer is passed by value. In Java, a reference is passed by value.
Doesn't calling "pass pointer by value" as "pass by reference" lead to confusion and distortion of the true meaning of "by reference"?
C++ and Java have their own specific language-lawyer definitions of the word reference, but outside of those languages, "reference" is a catch-all term that includes pointers, handles, and indexes.
You have no problem using the terms 'loop' or 'recurse' to describe C code, right? But these high-level ideas are implemented without using those terms in the code. Same with 'pass by reference'.
Structs on the stack I pass by value. Structs on the heap I pass by pointer. It's a mental thing that I find helps with large codebases.
- a student trying to get a better grasp on neural networks
Essentially this library is a demo toy that will need a lot to work to make useful in real conditions.
It's 200 lines of C, what were you expecting?
Let people enjoy things.
One of the core requirements of modern deep neural networks is performance. The high accuracy you get from very fast training over large amounts of data is the only reason to look at neural networks over very fast, very accurate systems like XGBoost.
If something doesn't deliver that, then it's worth pointing out that it is inadequate for anything but toy tasks.
[1] http://neuralnetworksanddeeplearning.com/chap1.html#implemen...
Checking the vectorization report of icc 2018 for fprop() shows that the loops got inlined, unrolled and vectorized (with AVX512 and AVX2 instructions using XMM/YMM registers). I doubt that doing it by hand would be faster, but it would make the code a lot more complex.
You have a point on the parallelization part though, it can be cleanly parallelized with a few OpenMP Pragmas.
godbolt link of the source + assembly -- all I did was combine the header, source and test files:
https://godbolt.org/#z:OYLghAFBqd5QCxAYwPYBMCmBRdBLAF1QCcAaP...
The second row is the probability of each given the input.
e.g. This output means it's 98.6% certain that the input represents the digit 7.
0.000000 0.000000 0.000000 0.000000 0.000000 0.000000 0.000000 1.000000 0.000000 0.000000
0.000009 0.000005 0.000000 0.000000 0.003899 0.012933 0.000436 0.986013 0.000005 0.000000Q. It seems that it only does 1D feature vectors?
A. Convolution and other inductive biases are only necessary when you have small data.
The third line of the readme says "We are finally at our April 1 Release (v4.1.2018)." The date may be relevant. ;)
Finally people have started speaking up and doing something about it! May we see more of such mentality in the future, and computer industry will be a better place to work!
Please make your applications modular. Optional dependencies are there for this reason.
Most good packaging tools support them.