Stanford’s Alpaca shows that OpenAI may have a problem
the-decoder.com
the-decoder.com
and yet i bet openai was trained on data that they did not bother negotiating and honouring "terms of use" for. what's the difference?
Companies can take away a dizzying variety of your rights if they simply say so. It's how they won most of the anti-scraping cases.
AI output is strictly not copyrightable under current law, so standard tricks to limit what people can do generally don't work [2].
[1]: https://github.com/pointnetwork/point-alpaca [2]: https://www.federalregister.gov/documents/2023/03/16/2023-05...
Now, in case of scraping _all_ of public useful information. Let's take Stackoverflow as an example of a community-produced content. It gets 90% of its traffic from Google, people show up, read, ask and answer questions (click on jobs and ads to keep SO servers running). If a scraper comes by that inhales all of that data, and regurgitates it under Microsoft's name, giving no credit, no references and absolutely nothing back to the authors of those questions and answers. What is the reason for SO to exist anymore? What is the motivation for people to keep answering and refining the questions? 90% of Stackoverflow will evaporate. The way it's currently set up it's an existential threat to any meaningful content on the web. Maybe there have been some legal precedents to repeal anti-scraping, but this is a very big problem now.
They would not have a moat from this strategy.
At the end of the day, they still make available a model with high performance. The whole point is that anyone, for very little money, can fine-tune another model to match the performance of OpenAI's models.
I've seen awful amount of fishy urls, magnet links, and diff files and I guess it works well enough for open source communities. This buzz and flow of amazing innovations building around LLaMA will later be repurposed when a truly open LLM + weights is released, not much is lost there. But it could be maybe x10 times the impact if startups could build their new products with the finetuned LLaMA models.
It just sucks that the price was already paid - 2.6 million KWh hours of electricity and 1,000 tons of CO2 emitted into the air. Now it needs to be done again to get to same result, why?
For that matter, I don't see how "OpenAI" could even try to legally enforce its terms against competitors training their models using the output of "OpenAI" models… at least not without being laughed out of the court room at best or ending up having to pay enormously much more themselves at worst, given how "OpenAI" blatantly disregard any licenses on the content and data they use to train their models.
Doesn’t mean it’s enforceable though. That being said, given the priority this is given in their ToS, I suspect they’ll be hearing from lawyers.
Edit: to be clear I don’t support this clause in their ToS, this is just something I noticed having had to study their ToS and privacy policy within the past week.
I don't see how a company could license the usage of something that they don't have the legal rights to - the output text in particular. Obviously OpenAI can terminate people's accounts for whatever reason they want, but that's largely meaningless. They have 0 chances of deterring anything unless they can secure substantial damages.
[1] - https://www.smithsonianmag.com/smart-news/us-copyright-offic...
I mean they didn't care about where all that training data came from but that doesn't mean anybody else should (in their minds).
That's them pulling up the ladder behind OpenAI
The genie escapes: Stanford copies the ChatGPT AI for less than $600 - https://news.ycombinator.com/item?id=35238338 - March 2023 (145 comments)
Stanford Alpaca web demo suspended “until further notice” - https://news.ycombinator.com/item?id=35200557 - March 2023 (77 comments)
Stanford Alpaca, and the acceleration of on-device LLM development - https://news.ycombinator.com/item?id=35141531 - March 2023 (66 comments)
Alpaca: An Instruct Tuned LLaMA 7B – Responses on par with txt-DaVinci-3 - https://news.ycombinator.com/item?id=35139450 - March 2023 (11 comments)
Alpaca: A strong open-source instruction-following model - https://news.ycombinator.com/item?id=35136624 - March 2023 (296 comments)
can you elaborate on what that means?
Their newer chat style endpoint for the GPT-3.5-turbo and GPT-4 models no longer supports this. https://platform.openai.com/docs/api-reference/chat
It can also be a free binary only model as well, but I would prefer a transparent one. Either way seems like Stanford Alpaca and LLaMa has taken off in cloning ChatGPT and making it good enough to compete against it at a affordable cost.
Might as well make it cheap and easily portable and fine tunable
People may loose jobs... but do companies want this? Can you imagine where google/fb/twitter/insert company name here no longer being able to hide behind "No one can reach a human"....
Be patient.
I haven't seen any but the news is going so quickly now lol
The chatbot user benefits from having one great answer pulled from the best sources (and no ads), but the websites that underpin the chatbot’s usefulness will no longer have monetizable traffic. In the long-term this disincentivizes people from publishing online, which would reduce the quality of not only the chatbot output but the web as a whole.
Imagine local news being totally unavailable online by any means because the rise of chatbots means that nobody can make any money writing about local news.
Edit: A first reaction to this might be “have the chatbot show ads and share revenue with its sources.” This probably wouldn’t solve the problem. Journalism (and many kinds of writing) would be a less attractive career if your readership consists mostly of people getting second-hand summaries via chatbot. If chatbots do become popular, I worry about a bleak future where journalism and other writing is replaced by an anonymous blob of underpaid foreign laborers whose only job is to shovel up-to-date facts into chatbot databases.
That's a valid perspective, but remember that the quality degraded when we started having paid memberships and ads to websites.
What's a news outlet worth paying for? who actually pays for content online? who knows how to block all ads and didn't do that?
That business model really proved to be worthless, it dragged the quality down with more desperate pay-to-read prompts. Nobody will miss this and we will have again people with real world experience writing knowledge or opinions in their free time. So it goes back to that, quoting scientific papers, books and actual reputable and knowledgeable people writing blog posts.
With the news being mostly propaganda, I don't know what to quote there and how many outlets still have a reputation.
I personally write during my free time in my personal ad-free minimalistic website and have list of blogs I track for contents.
Things changes, waiting for full time "authors" to come up with a profitable plan to just progress is also not an option.
I don’t know how far we’re going to get if we not only expect volunteers to supply everything the model needs in order to be up-to-date and useful, but also expect access to the model to be free of charge and ad-free.
That’s a pretty massive amount of man-hours, compute, and R&D effort to expend for literally no return on investment.
> we will have again people with real world experience writing knowledge or opinions in their free time
We already have this now in the form of a bored knowledge-worker blogger class. It’s not nearly enough to provide up-to-date info for a chatbot, and I don’t see how the implosion of journalism as an industry will lead to more people spending their free time writing for no pay. If anything, it will drive more knowledge behind pay walls like Substack, which will be inaccessible to the chatbot anyway.
Either Microsoft is successful charging people for it or it's going to become just as infested as Google is right now.
Maybe even then. Double dipping is not exactly unheard of.
Early Bing AI had a tendency to be rather unhinged.
i.e they can't adequately defend their business with trade secrets.
Patents probably wouldn't work either because the structures are too easily recombined to bypass any conceivable patent that would be enforceable.
That said I think all of this is actually emblematic of a deeper problem with the space which is that none of the LLM stuff recently has been groundbreaking but rather just just continual refinement of a given branch. We aren't seeing evolution, just increasing either number of parameters or quality thereof + additional context. Which is why it was so easy for other folks to make the same progress in similar time periods.
Time will tell if we are about to slam into a local maxima or if someone finds a significant evolution or better yet stumbles on a way to properly combine LLM for context + NLP with traditional AI/logic/expert systems to engineer something that actually thinks and learns rather than regurgitating statistics.
For example, imagine that engines were regulated earlier, because steam engines could blow up and kill people. That's a good reason to regulate steam engines. But then we probably would have never invented other engine types like gasoline engines and jet engines. With that, we'd never have invented planes or flight, because a regulation steam engine would have been too heavy.
The only thing AI regulation would do is hand AI supremacy to China. This is the case with almost every other technological development that we kneecap ourselves with, nuclear power, high speed rail, etc. We waste endless energy on bureaucracy while China is building.
The only reason why propaganda is so effective is because life is so terrible. No one bought USSR propaganda in the 60s that the US was terrible because people remembered growing up without electricity. A majority believe Russian propaganda about the US today because life expectancy in Thailand is higher.
In which case, it wouldn’t even need to be censored, would it?
But, in this case, as already discussed in other threads, why not simply transform everything people write or say? They still get to see what they wrote. Others only see a cleaned up and supportive of the gouvernement version. And to that on texts, chat, social networks,… everywhere you can.
And suddenly, you actually are in 1984. Except you don’t even have to send the police, or beat up people: they’ll all be deeply convinced they are part of a very small minority. If not alone.
If you want a generative AI without any decency then make it yourself, but don't act like this AI is censored just because it communicates in the way that everyone in society does when they are not protected by the veil of online anonymity