6 karma · joined February 25, 2025
this preprint is not coming from a standpoint of optimizing the inference/compute, but from trying to create models that we can interpret in the future and control
one simple usecase for them is physics-informed neural networks and neural ODEs, where using activation functions is discouraged, mainly because they aren't infinitly differentiable, and they use the tanh or the sin most of the time, this kernel i introduced works better then the neurons followed with a tanh to solve different PDEs
what i did was rely on both the angular information and spatial information between the input x and the weight w to measure how "similar" they are.
the lower bound of the yat-product is 0, and it is achieved only when two vectors are orthogonal and away
I was able to create a new kernel that allows you to learn non-linearity without using activation functions, making the models whitebox, and without any information loss.
MiniGPT with huggingface datasets streaming: https://www.kaggle.com/code/skywolfmo/yat-nnx-minigpt-finewe...
It's a PWA that runs entirely on your local machine. No vendor lock-in, no data sent to the cloud overlords. WebLLM is supported out of the box.
Fork it, customize it with your own extensions, and deploy it for free on Netlify. You can also host your own hub on Firebase.