HNHacker News
TopNewBestAskShowJobs

rikimaru0345

9 karma · joined August 9, 2020

submissionscomments
rikimaru0345··on Transformers know more than they can tell: Learning the Collatz sequence
Ok, I've read the paper and now I wonder, why did they stop at the most interesting part?

They did all that work to figure out that learning "base conversion" is the difficult thing for transformers. Great! But then why not take that last remaining step to investigate why that specifically is hard for transformers? And how to modify the transformer architecture so that this becomes less hard / more natural / "intuitive" for the network to learn?

rikimaru0345··on Cerebras-GPT vs. LLaMA AI Model Performance Comparison
> The difference between them is only one digit (i.e., the last number). Therefore, it's not possible to tell if either value is greater or lesser by just looking at their values without knowing more information about what those numbers represent and how they were obtained in the first place.

That one is especially hilarious. But the part at the end "how they were obtained" is really strange. Where in its dataset would it possibly have learned such nonsense? Doesn't matter where numbers come from to compare them.

It implies that it doesn't understand what numbers even are in general, and that giving it a calculator (that it can use perfectly) only masks a much deeper problem.

I mean, I'm reading a ton of people say "its not just pattern recognition and token prediction, it has emergent properties!!!!" and from experimentation I believe it.

But if the models can pick up language and its intricacies, and even do simple logic tasks, shouldn't it also be able to pick up on what numbers are and how they work? At least knowing that where a number came from doesn't matter when its just about comparing their value in a pure mathematical sense?

What does that mean for concepts other than numbers? Do those models fake a LOT more than we already believe they do?

rikimaru0345··on Open Flamingo – open framework to train multimodal LLMs
> Parse the chess board:

Could it be that the actual issue has to do with it having trouble with small tokens (letters, numbers)?

Does it give a different result if you ask it to answer in a format like this?

> Please name what kind of piece is on each square of this board > A1: white rook > A2: white pawn > A3: empty > A4: empty > ...

Prompting can be so unintuitive sometimes. Maybe it just has an issue with the output representation or something...

rikimaru0345··on Keeping up with the overwhelming pace of AI innovation
Dismiss it at your own peril. You'll get blindsided like you won't believe.

I've heard this sentiment before from a few people over the last few days, and they changed their tune very quickly once they actually looked into it.

rikimaru0345··on The Age of AI has begun
what ground truth model? is there some paper i can read? can you post a link to it?