It appears to have made some people billionaires. Examples:
Spotify was started by an employee uploading their MP3 collection.
The FB scraped the Harvard student directory.
It appears to have made some people billionaires. Examples:
Spotify was started by an employee uploading their MP3 collection.
The FB scraped the Harvard student directory.
It’s a slightly more interesting exercise to estimate the multiple on how much OpenAI can’t afford what they’re doing with any conceivable investor if they had to pay for the training data, but I’ll posit it can be demonstrated to the satisfaction of a reasonable person if not to a business-friendly court. I’ll bet the rent it can be demonstrated to a European court.
The only question left unanswered is how useful OpenAI will be. If it becomes as useful as Google (many claim they already did), then the reality (read: legislation) will bend for them.
Selling human organs will never be the same as selling beef liver, while the difference between a human's liver and a cow's is much smaller than the difference between your brain and a A100 tensor core.
Gen-AI is an empowering technology for the masses and easy to run locally. And the providers are in a race to the bottom on pricing.
(Whether they should be able to hike drug prices by 100% overnight or sell drugs with insane markups is a separate question.)
OpenAI is charging a _really_ high monthly fee for ChatGPT, and it’s quite popular. It’s very limited - there’s no way the cost incurred for usage is near that. Obviously their costs include the R&D that has already happened, but I still think it’s priced way over that.
Drug manufacturers are famously making money hand over fist. Patients aren’t their customers. Insurance companies are. They are surely gouging insurance companies to their fullest capacity.
And big LLM providers will certainly try, but they have competition now, all 3 tied up at roughly the same level. And then there is competition from open models that can handle 50% of what big models do.
So they can only set a high price for very advanced/critical tasks, where usage will be much lower.
With every new technology, there needs to be legislation and judicial rulings to determine the exact definition of things so that the best people can offer you is speculation or their interpretation. The quality of those interpretations will vary by person; I suspect most will be relatively uninformed.
One important factor to consider is the terms of use of the websites from which the data is obtained. For example, YouTube's terms clearly state that content may only be used for personal, non-commercial purposes.
Additionally, the current generation of models does not derive core aspects of knowledge or build world models; instead, they compress vast amounts of data and use that information to generate content. This process likely involves storing copyrighted information, which may violate copyright law.
People produce videos hoping or expecting to be monetized. People who monetize videos (by advertising or sponsoring) do so hoping for the audience to buy some product or service. People who host videos do so expecting to be paid by the monetizing people. That's how the current economic model for ad-supported video works.
A machine learning model is not an economic actor[1]. It's not going to go out and buy a flat wallet or a new set of headphones or whatever else is being advertised on stream. So the current economic model under which that content produced (which is predicated on the assumption that the audience is human) totally falls apart if the audience is an ML model.
[1] yet. And even when they routinely can take part in economic transactions they aren't the audience the advertisers are paying to reach.
So I’ll do it for you, because it’s squarely within site guidelines and norms to defend high-signal comments from getting gang-tackled.
The parent is right, you all know it, and as pernicious and ugly as this side of our business has always been, the stakes are some huge discontinuous amount higher this time.
So downvoters, remember what that button means: it means this is a low-signal or otherwise net-negative value comment on the thread. It explicitly doesn’t mean I disagree.
There’s a disagree button, it’s labeled “reply”. Use it, or get lost.
Regarding the argument, I have seen both sides argued well. The best argument that I have seen for OpenAI being at fault is in cases where gen AI outputs exact copies of copyrighted work, be they visual or text. Dall-E and GPT both need to have serious guardrails added to prevent this, as is shown by the leaked system prompts. If you can jailbrean an LLM, it happily provides the user with exact copies of copyrighted work, in some cases.
The ethical debate had settled into a soft consensus that this sort of thing was icky at best on topics ranging from FB leveraging the Harvard facebooks to Uber ignoring medallion laws to Google scanning all those books to AMP cutting the publisher out of the loop to Amazon eating the OSHA fines rather than give the workers bathroom breaks to FB/Google paying the fines and laughing over EU regulations (which in fairness they seemed to have backed down on), to currently Apple trying to malicious compliance the DMA gatekeeper regulations.
Some of this stuff has been litigated, some not, all of it is cringe-worthy, this is a new low.
The parent cited germane examples from the list. That adds to the discussion.
Folks are free to do the moderate effort thing and write a dissenting reply, they’re not supposed to do the lazy thing and grey out something they spinal-reflex don’t like.
You took the time to write a reply, so even though I think you’re wrong, I applaud you taking the time to state a position.
That's simply incorrect.
And besides, even if garage and dorm room projects sometime do start by illicit means, I don't really think it's responsible to encourage anyone to normaliE that and think it's grossly irresponsible to use that idea to excuse well-connected multi-million dollar organizations who can't hide behind ignorance and should be held to the highest standards lest the playing field for everyone else is ruined.
On your substantial point about garage bands needing to sometimes move fast with limited legal infrastructure but arbitrarily resourced actors being held to a higher standard we’re in violent agreement.
The strongest argument (and it’s not one I’d like to have to defend) for law-breaking tech corporations netting out useful is probably Uber?
The situation pre-Uber was a capture cartel nightmare where basic transportation was a non-starter in certain situations/municipalities.
The (admittedly weak) case that they should still be involved in some, “civil disobedience”, is that the medallion people seem to have decided they’re just never going to give up.
If anyone has a “less bad” example of a big powerful company serving the public good I’d be keen to be reminded of it, but this “fair use for me, copyright for thee” seems to go hand-in-glove with stuff that belongs in both civil and criminal court.
Unfortunately, this is your personal interpretation of the downvote buttons. I agree with your interpretation, in fact, but it is not the norm. And confidently telling people to "get lost" while being misguided is very much not appreciated here.
Here is some comment from the horse's mouth more than a decade ago on this very same topic:
I meant that in the “karma bombing” sense as far as “explicitly” goes.
And I should be more clear in both this subthread (as well as one other comment) that the guidelines I’m referring to aren’t the succinct ones on the literal guidelines page, but rather the evolved apparent consensus under “the @dang administration”.
This is a pretty natural and healthy seeming extrapolation of the idea that rigor should go up as contentiousness goes up, something I personally fail in too often even these days, but will always apologize for failing in.
Threads about OpenAI have set a new record for karma bombing proportional to front-page exposure, blowing away classics like crypto and social media and patent law, I’ve got at least one or two tuples of adjacent comments where the high-quality one got bombed and was on a #1 submission and the far worse form of the argument got too many upvotes on a less-engaged submission.
I’ll decline (for now at least) to speculate on the mechanism/explanation around this (apparent) trend, but it’s inching into “concerning”.
Another take: is real true value creation only possible outside of capitalism? Hence those communist pockets like academia and practices like treating everything as public property are allowed to exist, otherwise there would be no innovation etc just a bunch of companies suing each other?
After all, while companies didn’t have to pay to pollute the environment, humanity has to eventually suffer; and in the same way the current crop of generative has people’s their work assimilated into models that results in profits for AI companies while the author of the content is denied their fraction of the economic value generated from this process, as well as any recognition and other second order effects they may have benefited from.
On the other hand, AI models directly derive their value from the training data, and to a large extent, produces works that are a substitute for the original works.
By the time Google came on the scene, web crawling was well established. AltaVista and Lycos, among others, were already entrenched search engines.
robots.txt was created by the guy that created the first search engine.