An SLM trained on $8 ESP32-S3
github.com
github.com
I wonder when we'll start seeing clusters of ESP32-S3s... not sure how interconnects would go though, but I guess the interconnect wouldn't be the bottleneck anyway.
> Backpropagation (gradients derived by hand)
What does by hand mean in this context?
Also how did you write the readme? It's a curious blend of human and AI writing.
Anyways, it looks like gradients derived by hand means they didn't use autograd. They have written out the expression for dL/dW themselves.
Specifically: I wrote it in Spanish first, which is my language, and then translated and condensed it into English with Claude. That's probably the blend you're picking up. The jokes, the structure and the decisions are mine; the English phrasing has fingerprints that aren't.
The code is more clear cut and it says so in the header: generated by AI under my direction. I wrote the architecture, the decisions and the validation, not most of the C.
I'd rather say that upfront than have someone find it later.
I think it is a small thing.
Number of parameters: 319K
(Disclaimer) The models is tiny and is not a pocket chatbot. It mistakes and is not capable to support conversation, but that's not the goal of the project
The goals of the project is to bring training to edge devices and it worked out
How can one use it? By using solar panels such device could be turned into autonomous meteorological station