I originally assumed that this was due to the increase in demand. It never went back to being as sharp as it was during those first hours of usage
I originally assumed that this was due to the increase in demand. It never went back to being as sharp as it was during those first hours of usage
OpenAI have repeatedly stated the model hasn't changed so how could this happen otherwise?
If it gets out and you're right, I think it will cause major trust issues with the product.
I'm not sure what you mean by "conspiracy", but this sort of thing isn't unknown. All it takes is the employees being bound by an NDA and a marketing team can say anything it likes (within the bounds of legality, anyway) without fear that they will spill the beans.
This is the part of hype cycle named "Peak of Inflated Expectations", at least when it comes to the potential of the tool. I think we are still in the "Innovation Trigger" phase in terms of applying the technology.
From my point of view, it is sad that these sort of socio-political constructs (copyright) are hindering innovation. The funny thing is that in say, 10 years, the "pirate" version of LLMs will be way more powerful and useful than the "corporate" versions. Once the RIAA/MPAA has sued all of them preventing them of using their audio/video; after the editorials have sued them to prevent them from using their books and articles, and after every internet site has sued them to prevent them from using their text. LLM models trained on SciHub, Library Genesis and Torrents will be amazing in comparison.
I sincerely wish that some country would apply to information copyright a similar approach to what India does for medicine patents.
[1] https://news.ycombinator.com/item?id=31852138 [2] https://news.ycombinator.com/item?id=36138930
Why not pay authors of the data the LLM has ingested?
So should they be paid a few pennies every time the LLM spits out a response that "used" that training data?
And I am pretty sure it's not even possible to really link the output back to training data anyways.
You're thinking of broadcast radio. Streaming is a fraction of a penny!
> So should they be paid a few pennies every time the LLM spits out a response that "used" that training data?
That doesn't sound unreasonable! I understand that current LLMs have no way to report "this token came from this data" but that doesn't mean it's impossible to build. (Ack: this is a full-on proper dunning-kruger, having not looked into it & having zero knowledge of the field.)
But I think it's probably more reasonable to simply split a % of the service revenue across everyone whose data was used. Or pay an up-fee for ingesting the data in the first place.
Generally, it's crazy to name that some people think it's reasonable for these companies to pay for GPUs and CPUs and electricity to run them but not for the data that's the actual core of their service.
Didn't Sam want to create UBI? That's one way!
HN I agree with you, but I feel like I'm getting enough back from this community to make my time worth it.
Care to take this opportunity to explain India's approach to medicine patents, and why you think it's good?
I remember seeing interviews and reading many comments saying that that the copyright data shouldn't be an issue because once the model is trained, it kind of "forgets" the copyrighted material and we're just left with pure, unfettered intelligence...it's not a fuzzy jpeg of the web etc...
I think the copyright angle is on the money though, this is why Bard isn't as good, just as Adobe Firefly isn't as good as Stable Diffusion. The inputs aren't as good so the outputs aren't either.
Google can't come out in public and accuse OpenAI of blatant copyright infringement for legal reasons, but I bet internally, they know what's up.