I wonder how much third party apps are being caught up in this other issue.
I wonder how much third party apps are being caught up in this other issue.
Biggest parties have already mined the data, which is enough for models for long time.
Unless you want the model to find some specific comment yesterday.
By allowing LLMs low cost or free access to their users data companies like Reddit are essentially helping companies like OpenAI choke their traffic.
Over time as people realise it’s quicker to ask ChatGPT a question than it is to post a question on reddit they will start losing content also.
Also think about the difficulty of policing content generated by swarms of LLM bots with API access for PR campaigns.
I am just wondering, where is the limit, since in that case the model might not be trained anymore and instead it is used for similar purpose than search engine. I guess Bing is already doing this, without Reddit API.
I just asked Chat GPT this very question and it didn't mention the Fenrir. It mentioned the other contender (MODE) but it failed to differentiate it from all the other offerings i.e it has the ability to connect a SATA SSD and hold the entire Sega Saturn library on the device.
The cases LLM will need this data in is compiling together more modern useful human knowledge as our knowledge base grows. Information in existing LLMs could shift. This is sort of the issue even academic textbooks deal with when publishing what is considered foundational knowledge: sometimes we discover something new that makes it either not quite correct or invalid.
These are the sort of obvious failures and disconnects that should become apparent if training lags behind. LLM services interested in revenue without plans for continuously updating training data are somewhat betting that not too much will change from most end users perspectives for awhile and for some use cases that might be true but the limits of training data over time for public instances of GPT for example have already hindered some.
Much of prompting, from my anecdata, needs to take that into consideration as one of the base constraints (does this model even have up-to-date information it could query and dump something useful from). To some degree those training limits also help expose "hallucinations" or interpolation/extrapolation attempts of LLM models. If I know it doesn't have this information in the training set and test the system against it, I can observe how well it interpolates, extrapolates, and is transparent about when that's happening. For example if I ask existing models about new syntaxes and structures introduced in Java 21, most should return something back like it doesn't exist, it lacks that newer information, or something to that effect. If instead it starts producing code samples that it couldn't possibly have knowledge of, then I know it's passing back garbage. If it's being continually updated at some frequency, I'm no longer so sure and it may actually be providing new useful information.
The problem goes away if the apps charge for a subscription, which covers their API access.
And require something like a public key to access the API so you can track requests coming from multiple hosts.
Then any rate of access above that of a power user can be charged for appropriately. And make it so that activating API access is something that can't be automated so people can't create thousands of API dummy accounts.
The problem here is that many users are behind CGNAT, meaning many end users share a single IPv4. Unfortunately the days of counting distinct users by their IP(v4) are over.
So I think they are thinking that and I don't think they're wrong.
Yes, the data is valuable for product recommendations, but not at the price they are asking. And if they ever block traditional scraping then all those "append reddit to google" searches will be gone too. Google is not going to pay them either.
Unless they put 100% of content behind a login-gate, it's currently legal* to scrape and use for derivative works, as long as you have the money and the chutzpah to deal with lawsuits that may or may not come.
- https://blog.ericgoldman.org/archives/2022/12/hello-youve-be...