> Although they reduced the number of operations, the researchers were able to maintain the performance of the neural network by introducing time-based computation in the training of the model. This enables the network to have a “memory” of the important information it processes, enhancing performance. This technique paid off — the researchers compared their model to Meta’s state-of-the-art algorithm called Llama, and were able to achieve the same performance, even at a scale of billions of model parameters.
I disagree. The authors aren't conveniently omitting anything. They show all details in a comparison against LLama models.
Moreover, all evidence I've seen so far suggests that tritwise models can scale up to state-of-the-art sizes.
---
PS. I'm talking about the paper, not the fluffy press release.
Thanks for pointing it out!
PS. I added a PS to my comment above.