rpdaiml··on Show HN: Bonsai 1.7B ternary model at 442T/s on M4 MaxNice work, that throughput is wild.
rpdaiml··on Show HN: Muesli – If Granola and Wisprflow had an open source on device babyHow are you handling the on device speech pipeline, especially around model size, latency, and accuracy tradeoffs on consumer hardware?
rpdaiml··on Show HN: I built a tiny LLM to demystify how language models workThis is a nice idea. A tiny implementation can be way more useful for learning than yet another wrapper around a big model, especially if it keeps the training loop and inference path small enough to read end to end.
rpdaiml··on Show HN: 1-Bit Bonsai, the First Commercially Viable 1-Bit LLMscurious and wondering its porting and usage on edge devices?