Would love to know exactly what the latest process is to keep slop out of training data.
Would love to know exactly what the latest process is to keep slop out of training data.
:)
There's way more value, if seeking out answers, in following the links to external sources, scraping books, and other sources that aren't "unwashed masses saying whatever they want".
> ...
> scraping books, and other sources that aren't "unwashed masses saying whatever they want".
The problem is there's a lot of knowledge that only exists as reddit comments, blog posts, or social Q&A.
Kind of a steep slope to convince me that Reddit has a higher P/E ratio than Nvidia because of its authentic content… Reddit seems extremely overvalued because it is a network of people programmed to accept whatever they see and engage.
No doubt that there is good information there, discord too.
But Reddit is far more bots and “organic propaganda” than you are thinking.