Using ctrl-f I was able to see that they were identical in one another.
Obviously this is a single sample but saying 90% seems unlikely. They were around ~80k tokens total.
Using ctrl-f I was able to see that they were identical in one another.
Obviously this is a single sample but saying 90% seems unlikely. They were around ~80k tokens total.
I would note that LLMs handle this task better if you slice the two documents into smaller sections and iterate section by section. They aren’t able to reason and have no memory so can’t structurally analyze two blobs of text beyond relatively small pieces. But incrementally walking through in much smaller pieces that are themselves semantically contained and related works very well.
The assumption that they are magic machines is a flawed one. They have limits and capabilities and like any tool you need to understand what works and doesn’t work and it helps to understand why. I’m not sure why the bar for what is still a generally new advance for 99.9% of developers is effectively infinitely high while every other technology before LLMs seemed to have a pretty reasonable “ok let’s figure out how to use this properly.” Maybe because they talk to us in a way that appears like it could have capabilities it doesn’t? Maybe it’s close enough sounding to a human that we fault it for not being one? The hype is both overstated and understated simultaneously but there have been similar hype cycles in my life (even things like XML were going to end world hunger at one point).
Needle-in-a-needlestack contrasts with needle-in-a-haystack by being about finding a piece of data among similar ones (e.g. one specific limeric among thousands of others), rather than among disimilar ones.
This is such an anti-intellectual comment to make, can't you see that?
You mention "sample" so you understand what statistics is, then in the same sentence claim 90% seems unlikely with a sample size of 1.
The article has done substantial research
He's a much simpler and correct description that almost everyone can understand: it fucks up constantly.
Getting something wrong even once can make it useless for most people. No amount of pedantry will change this reality.
This isn't pedantry, it's science.
Also: would you expect random people to fare any better?
It has done more complex things for me than this and, sometimes, gotten it right.
Yes, it’s supposed to be able to do this.
It's purported to be a major use case.