There's also NoPE, I think SmolLM3 "uses NoPE" (aka doesn't use any positional stuff) every fourth layer.
2,808 karma · joined June 13, 2012
There's also NoPE, I think SmolLM3 "uses NoPE" (aka doesn't use any positional stuff) every fourth layer.
I understand that you can get highly power efficient XORs, for example. But if we go down this path, would they help with a matrix multiply? Or the bias term of a FFN? Would there be any improvement (i.e. is there anything to offload) in regular business logic? Should I think of it as a more efficient but highly limited DSP? Or a fixed function accelerator replacement (e.g. "we want to encrypt this segment of memory")
Couple of gotchas that I ran across. I found that on Linux, desktop PC fan control support is pretty abysmal. The sensor library that everyone relies on, lm_sensors, is semi-abandoned and didn't recognize sensors on my relatively popular, 7 year old ATX motherboard and GPU. It also requires having Perl installed.
About GPU cooling in particular - modern NVidia cards in particular seem to have a built-in minimum of 30% fan speed when controlling them manually. The connectors are also a different, smaller connector (perhaps a JST PH?).
There are also bunch of sellers who package samples (aka "decants" - buy a 100mL, split it into smaller bottles). I found that 1-2mL is plenty to get an idea. I've had great experience with LuckyScent (mentioned in the article), Surrender to Chance, as well as random reddit swaps and highly rated Ebay sellers.
The perfume scene is super wide and diverse, and I found that although there are general trends, it's hard to even know all the popular brands, and everyone's nose is unique. Skip stuff like Aventus and Sauvage and buy some discovery sets (surrender to chance puts together some good ones).
There is definitely a spectrum between "wearable crowd-pleaser" and "avant-garde storytelling" - Afrika-Olifan comes to mind - love it for the creativity and execution, but it would be rude to go outside wearing it. There's also some storytelling - Black March, for example, starts off with grassy fresh earth after a rain, then turns into flowers.
I also remember trying to fit a distribution so that I can generate synthetic data (not for a lack of data, but more for understanding the problem space better). The synthetic data quantized pretty differently - my guess is that it's because of random areas of density and sparsity.
I'm not quite following your exact rearranging idea though. Not sure if the above answers the question.
EDIT: groped -> grouped
However, I'm talking about the probability distribution of tokens.
Let's say that we have 15k unique tokens (going by modern open models). Let's also say that we have an embedding dimensionality of 1k. This implies that we have a maximum 1k degrees of freedom (or rank) on our output. The model is able to pick any single of the 15k tokens as the top token, but the expressivity of the _probability distribution_ is inherently limited to 1k unique linear components.
Is there a zero-copy interface for larger objects? How do object lifetimes work in that case? Especially if this is to be used for ML, you need to haul over huge matrices. And the GIL stuff is also a thing.
I wonder how Mojo handles all that.
I was trying to figure out the difference between the Stockfish approach (minimax, alpha-beta pruning) versus Alpha Zero / Leela Chess Zero (MCTS). My very crude understanding is that stockfish has a very light & fast neural net and goes for a very thorough search. Meanwhile, in MCTS (which I don't really understand at this point), you eval the neural net, sample some paths based on the neural net (similar to minimax), and then pick the path you sampled the most. There's also the training vs eval aspect to it. Would love a better explanation.
I noticed that LLMs need a very heavy hand in guiding the architecture, otherwise they'll add architectural tech debt. One easy example is that I noticed them breaking abstractions (putting things where they don't belong). Unfortunately, there's not that much self-retrospection on these aspects if you ask about the quality of the code or if there are any better ways of doing it. Of course, if you pick up that something is in the wrong spot and prompt better, they'll pick up on it immediately.
I also ended up blowing through $15 of LLM tokens in a single evening. (Previously, as a heavy LLM user including coding tasks, I was averaging maybe $20 a month.)
I love sci-fi, I love challenging ideas, and I really liked the concepts explored in Blindsight - except that I learned those concepts through summaries and selective reading.