RAG is needed for the same reason you don't `SELECT *` all of your queries.
16 karma · joined September 21, 2015
RAG is needed for the same reason you don't `SELECT *` all of your queries.
> This Reddit / Stack Overflow API drama feels weird to me as a user. It's USER submitted data. R/SO didn't create it.
> So what ChatGPT scraped it? They're benefiting the same users who contributed the content on R/SO, and then some.
> It's not like ChatGPT is cloning them.
> It would be like if Printing Houses tried to block authors from publishing their content as an eBook or Audiobook.
> Users contributing to your ecosystem doesn't give you perpetual dominion over that content
My understanding is that there are 1000s of different compression algorithms, each with their own pros/cons dependent on the type and characteristics of the file. And yet we still try to pick the "generically best" codec for a given file (ex. PNG) and then use that everywhere.
Why don't we have context-dependent compression instead?
I'm imagining a system that scans objects before compression, selects the optimal algorithm, and then encodes the file. The selected compression algorithm could be prefixed for easy decompression.
Compare a single black image that's 1x1 to one that's 1000x1000. PNGs are 128bytes and 6KB, respectively. However, Run Length Encoding would compress the latter to a comparable size as the former.
That's an expensive boat.