Web Neural Network API
webmachinelearning.github.io
webmachinelearning.github.io
Please correct me if I'm mistaken (not a ML expert): On top of these existing APIs, we can build efficient neural networks. Of course, we wouldn't be able to use specialized hardware such as Google's TPU or Graphcore's IPU, but I doubt there's a use-case for those in web.
Is there any non-GPU/CPU based consumer hardware aimed at fast neural networks that I'm missing?
EDIT: Seems "NPU"s have been slowly arriving for the past 2 years to iOS/Android-based devices... Maybe there's a raison d'être for the WebNN API after all.
Decades ago, compiler developers solved dealing with multiple CPU architectures, by creating languages that succinctly expresses their common traits: Doing arithmetical-logical operations and accessing memory. Languages like C abstracted us from CPU registers, stacks. Optimization passes like vectorization abstracted us from different ISA extensions.
Now it feels it's the same issue all over again. Is it really hard to simply express: `x' = W*x + b` and let the compiler target the CPU/GPU/NPU as needed?
But that's more an expression of API design than any particular reason.
Once TPUs can run JS (which may be coming sooner than it seems), it might also be reasonable to write a "webpage" which TPUs then "execute".
I think neural networks in general need to learn to be more flexible. Right now it feels like you're carving a network out of marble. It should be more like clay.
More practically, tensorflow.js does run in the browser; you could port this API to a tf.js backend.
I know this is mostly an Apple problem but I’m proposing the working group do more to apply pressure if it all possible.
Controversial, but stapling in unnecessary machine learning as a web api when we already have enough performance issues with web apps as it is. Anyways, why go to the effort of this, when if you really wanted to run ML models on the front-end you could use ONNX (a format designed already for portable network serialisation) and WASM.
> The application is watching whether she is in front of her PC by using object detection
let shape = [1,-1,1,1];
return nn.add(
nn.mul(
nn.reshape(scale, shape),
nn.div(
nn.sub(input, reshape(mean, shape)),
nn.sqrt(nn.add(nn.reshape(variance, shape), nn.constant(epsilon)))
)
),
nn.reshape(bias, shape)
);
I find nested function calls like these very hard to comprehend.
I would like intermediate variables and comments explaining what is happening and why. Make it flat. let shape = [1,-1,1,1];
# line up shapes.
mean = nn.reshape(mean, shape)
variance = nn.reshape(variance, shape)
scale = nn.reshape(scale, shape)
bias = nn.reshape(bias, shape)
# calculate 1/sqrt(variance); add an epsilon to prevent dividing by 0
epsilon = nn.constant(epsilon)
inv_var = nn.rsqrt(nn.add(variance, epsilon))
# batchnorm, I think?
x = nn.sub(input, mean),
x = nn.mul(x, inv_var)
x = nn.mul(x, scale)
x = nn.add(x, bias);
return x
I don't know why they didn't use nn.rsqrt, since it's a common operation. In fact it's so common it's a builtin op: https://www.tensorflow.org/api_docs/python/tf/raw_ops/RsqrtI'm skeptical the reshaping is even necessary. Any reasonable framework should broadcast those shapes automatically.
This example uses the current polyfill and it's local.
Microsoft should've not been allowed to join the W3C.
It's a community group document, so not a standard and not on the way to become one.
Second, W3C function as a standard promoter is compromised when it allows its members to make use of draft/private proposals as ersatz standards.