Look at the startup OpenAI generations 1-2 years back - they have largely sank at comparable or worse rates than any other startup from 2020/2021.
The GPT-3 first wave companies, around translations, basic quizzes, summarization tools, language learning apps, and ofcourse the notorious paraphrase tools (almost entirely obsolete since ChatGPT) can't be found in that form anymore, they've all been forced to shut down or move functionality significantly. Early on, OpenAI limited output to 300 tokens max - and less for most usecases, often 50-150. Chatbots were not allowed IIRC for over a year. If ChatGPT hadn't came along, much of what langchain enables wouldn't have either, nor so many big companies willing to now risk.
I can count none over 2 years old which have not since been made obsolete by raw ChatGPT access or are now default dead due to competition from existing unicorn (e.g. Duolingo, Quizlet Q chat) who are now crushing them.
It must be painful to have spent $xx,xxx on GPT-3 at $0.06 a token to obtain users, and now have your market ripped from you by a $B+ company paying $0.006...
So I doubt the template UI startups will sustain retention or stay about long term, unless they really find traditional startup ways to nuggle into niches and use cases/vendor lock in. This isn't an innovators dilemma for most companies, it's just an obvious sensible thing to try at this point, so startups don't have much to balance with on risk. That being the case, the market surely should seem less appealing than the open ended "blue ocean new value" of 2021 GPT tech, but - I guess not.
I thought the whole Python 3 thing was a huge problem. Lately I've been doing JS/Typescript dev and breaking changes like this happen continually and no one blinks.
Wasn't that hard for quite a while to find situations where dependency A was Py3 compatible and B was not (and the incompatibility went both ways, especially in early 3.x releases, you could NOT have one codebase that worked with both).
Sometimes A dropped Py2 support before B gained Py3.
Pain pain pain.
Then add the increasing level of insanity as the answer to "python packaging sucks" was repeatedly to add yet another layer.
But this is by 'design', I have very unreasonable suspicions that indeed, the whole VC 'world of entrepreneurs' is just the way the USA government does R & D on an industrial-corporate scale. The 'brilliance' behind this way of doing R&D is that they only pick up the winners after they won, so they don't "waste" money on R&D death ends nor moonshots.
on the other hand, this is a good way to 'explode' for cheap the technological applications of already developed scientific innovations. meaning none of those VC-backed startups are doing innovative research, but in fact are devleoping commercial applications for corporate overlords who having seen who won, step in to buy them out.
even music industry is shifting to that model, they are now only signing bands/artists/influencers who already build their audience.
At least that's what's going on in Hungary, the most corrupt government in EU, I hope other parts are a bit better.
(Additionally, newer LLMs like Perplexity.AI's correctly cite content sources, so that is even more similar to search engines)
It talks at length about this specific problem and migration techniques for it:
Existing foundation models are trained on copyrighted material. Deploying these models can pose both legal and ethical risks when data creators fail to receive appropriate attribution or compensation. In the United States and several other countries, copyrighted content may be used to build foundation models without incurring liability due to the fair use doctrine. However, there is a caveat: If the model produces output that is similar to copyrighted data, particularly in scenarios that affect the market of that data, fair use may no longer apply to the output of the model. In this work, we emphasize that fair use is not guaranteed, and additional work may be necessary to keep model development and deployment squarely in the realm of fair use. First, we survey the potential risks of developing and deploying foundation models based on copyrighted content. We review relevant U.S. case law, drawing parallels to existing and potential applications for generating text, source code, and visual art. Experiments confirm that popular foundation models can generate content considerably similar to copyrighted material. Second, we discuss technical mitigations that can help foundation models stay in line with fair use. We argue that more research is needed to align mitigation strategies with the current state of the law.
Further, new laws get made in reaction to new things whenever they push an existing doctrine beyond the original ruling, and these are certainly in that territory.
Of course. As I said originally "This is clearly not a given". It's very unclear how this will be decided, but anyone who thinks that just because models contain copyrighted data they don't have a leg to stand on is very wrong. There are multiple good arguments and precedents to show that they do, depending on the circumstances.
I think they contain massive amounts of copyrighted data, and reproduce them exactly, and that’s why they don’t have a leg to stand on. It’s a personal opinion, and I think backed by your citation. But thanks for the reference there, and glad to chat.