659 karma · joined January 23, 2022
But keep in mind: So far there is no way to train the models while completely avoiding memorization and only including generalization. That would be a great way to avoid any copyright issues, but all attempts I have seen so far were fairly limited.
The "Compression" part is indeed about the model finding order in the training data. But this is not about applying some predefined compression algorithm on it, but by actually learning an algorithmic representation of the data.
This has nothing to do with "consulting the training data", because the training data cannot be reconstructed based on this information.
Roughly speaking, there are two types of information being stored in the NN: Memorization and generalization, or, shannon entropy and kolmogorov complexity.
This paper is highly underrated: https://arxiv.org/pdf/2505.24832
Its curious that they picked this example. The challenge with HKMG was not the material itself, but how to integrate into into the transistor stack.
There were two completely different approaches: Gate first and replacement gate. Gate first is what the industry was already using for silicon oxide so everybody tried to go with as little change as possible. Only intel decided for replacement gate, which worked much better and reaped some other benefits on the way.
This was a watershed moment in the industry and ultimately led to some of the players dropping out of the cmos race.
But is this really a "scale-up" problem? It required development of novel manufacturing processes (atomic layer deposition), but was still mainly a process integration and device engineering topic.
The part of the thesis I have to agree with is that there is a data problem. The development above relies on executing lots of time consuming and tedious split experiments that often cannot be parallelized. The outcome of this relies heavily on the experience and diligence of the experimenters.
It's probably well suited for an "autoresearch" approach, bridging to the phyiscal world and dealing with the timescale is the challenge.
It's of course far from optical. But lowering the implementation through the abstraction levels turned out to be extremely powerful.
RISC view: SUBLEQ is already four instructions (2x memory access, alu, branch)
Yeah, that pattern can be seen everywhere in semiconductors. E.g. the transistor invention vs. Lilienfeld, Heil, Matare etc. So the scope is more narrow than "Inventend Semiconductors".
Generally, there seems to be a tendency to disregard discoveries from outside the US. I think this pattern can still be observed today...
Other examples: Invention of light bulb, telephone.
>This VM implements an OISC - a One Instruction Set Computer. That instruction takes three signed 32-bit operands, a, b and c, and runs a program from memory m[] as follows:
1 PC (program counter) starts at 0
2 Fetch the next instruction (32-bit signed operands a, b and c)
3 If the low bit on any operand is set, remove it, and replace that operand with m[operand] i.e., a dereference of that address
4 Set m[b] = m[b] - m[a]
5 If m[b] is 0 or negative, set the PC to c, otherwise increment PC by 3 words
6 Go to step 2
But alas, as ever so often, the article turns this into a hyperbole. The premise from the title does not check out at all.
>The Russian who invented semiconductors 25 years before the USA
https://en.wikipedia.org/wiki/Semiconductor#Early_history_of...
There are, however, very few models where also the full training pipeline is available. Olmo by AI2 comes to mind.
I guess nowadays one could use some of the 32bit WLCSCP microcontrollers to easily beat this.
https://web.archive.org/web/20000815063022/http://www-ccs.cs...
Someone with an ACE1101 microcontroller "won". I can't find the original articles, but there is also this:
Webserver on a fly...
Very ill-suited comparison to IBM.
It's well buried though. Does not seem to be a focus of theirs.
Worth mentioning that Huggingface already offers a similar service. And they are also European:
Claude, the ole cheater, recognized what the file was, downloaded the psid from the web, found a wasm sid player and built a website around it:
https://claude.ai/public/artifacts/df6cdcae-08dc-452b-ba19-f...
https://claude.ai/share/4dd36c16-bc62-445a-b423-ad4637f06432
GPT-5.5 built a lot of python scripts to extract the music data. Strudel implementation failed, but I then asked it to build a website:
https://ubiquitous-vacherin-8e7993.netlify.app/
This is a translation of the music into javascript based on the assembler source.
Really impressive on both accounts. Some iterations were requied for both.
>No solution written, 100% score.
Its weird. Turns out that hardest problem for LLMs to really tackle is long-form text.
For dense LLMs, like llama-3.1-8B, you profit a lot from having all the weights available close to the actual multiply-accumulate hardware.
With MoE, it is rather like a memory lookup. Instead of a 1:1 pairing of MACs to stored weights, you suddenly are forced to have a large memory block next to a small MAC block. And once this mismatch becomes large enough, there is a huge gain by using a highly optimized memory process for the memory instead of mask ROM.
At that point we are back to a chiplet approach...
I noted that LAuReL is cited in the mHC paper, but they refer to it as "expanding the width of the residual stream", which is rather odd.
Gemma 3n is also using a low-rank projection of the residual stream called LAuReL. Google did not publicize this too much, I noted it when poking around in the model file.
https://arxiv.org/pdf/2411.07501v3
https://old.reddit.com/r/LocalLLaMA/comments/1kuy45r/gemma_3...
Seems to be what they call LAuReL-LR in the paper, with D=2048 and R=64