2,136 karma · joined February 9, 2021
http://github.com/bwasti
[all posted thoughts and comments are my own]
the absolute most impactful improvements for inference comes at architecture design time. I firmly believe everyone who cares about impacting model efficiency should look there
half an hour to process 10k tokens on an M5 seems... not great
I've used it with some raspberry pis to create hi-fidelity walkie talkies it's quite pleasant.
whats wrong with this? You may be over-indexing on the need for large quantities of examples. These days self-play through RL is far more effective and data (not compute) efficient.
[1] https://www.hyundai-n.com/en/models/rolling-lab/n-vision-74
``` 2x + y = \operatorname{eml}\Big(1,\; \operatorname{eml}\big(\operatorname{eml}(1,\; \operatorname{eml}(\operatorname{eml}(1,\; \operatorname{eml}(\operatorname{eml}(L_2 + L_x, 1), 1) \cdot \operatorname{eml}(y,1)),1)\big),1\big)\Big) ```
for me Gemini hallucinated EML to mean something else despite the paper link being provided: "elementary mathematical layers"
https://www.oldnyc.org/#707133f-a this is supposed to be here https://www.oldnyc.org/#702487f-a
also, if folks are interested in these old depictions of NYC, check out https://1940s.nyc/ as well!
new attention mechanisms also often need new kernels to run at any reasonable rate
theres definitely a breed of frontend-only ML dev that dominates the space, but a lot novel exploration needs new kernels
dropping down into the familiar or the simple or the dumb is so innately necessary in the building process. many things meant to be "pure" tend to also be restrictive in that regard.
what makes you say this? modern LLMs (the top players in this leaderboard) are typically equipped with the ability to execute arbitrary Python and regularly do math + random generations.
I agree it's not an efficient mechanism by any means, but I think a fine-tuned LLM could play near GTO for almost all hands in a small ring setting