GPTBot – OpenAI’s Web Crawler
platform.openai.com
platform.openai.com
Courts are more deliberate than you would like — no denying that. But this is a feature not a flaw. It may be that damage will be done by then. Perhaps irreversible. But I would like to think if there is a will there is a way and that if things are terrible enough the governments will be bold in their responses.
Let's say the French government decides that OpenAI must change something about their business practices if they want to continue operating in France. OpenAI says "nope", and blocks access to French users.
Suddenly French companies aren't able to use GPT-X anymore – while their competitors in other countries can. How long do you think it will take before a storm of corporate outrage forces the government to relent?
Any individual government (except, perhaps, the combined US and EU governments) is powerless against today's technology megacorporations, because they can take much more away from a country than that country can take from them. If push ever comes to shove, it will become obvious where the true power lies. So far, the corporations have barely even tried to throw their weight around.
That's one possible outcome. (ETA: You DO have a point here, but...)
The other is, you know, something like every website explicitly telling me, via an annoying popup, how much they value my privacy. Also, me not being able to access half of US news sites to this day.
The last time EU raised their finger, every technology company (FAANG included) shat their pants.
And that was simpler times, times when a cookie stored in your temp folder without websites shouting they're about to do so, was somehow the biggest concern of an EU netizen. It almost seems ridiculous, compared to the damage AI could do (the extent of which which nobody really knows).
Bof, les alternatives à ChatGPT ne sont pas si mal.
And even if the open source alternatives were far behind rather than just a bit — all this talk about corporate moats and their absence may be blind to the strengths of OpenAI's offerings, but even so it can be replaced if it must — the storms of protest in France are normally by the people, not by the corporations.
But that's not true, and people know it.
> the storms of protest in France are normally by the people, not by the corporations
Correct. CEOs of big corporations just call the ministers directly and tell them to get in line, or else.
Based on what I've seen? They're good enough to be interesting, more so than GPT-2.
They don't need to be amazing from day one to be a foundation for replacing the status-quo.
> CEOs of big corporations just call the ministers directly and tell them to get in line, or else.
I roll to disbelieve (that it works, not that CEOs attempt it); that sounds like conspiracy theory to me.
You mean corporations that wield more power than most governments, and have revenues equivalent to the GDP of entire countries?
If Universal or 20th Century Fox were to ever become a serious obstacle, Google and Microsoft are simply going to buy them. This isn't the early 2000s anymore. The power balance has shifted dramatically.
US$26.2 billion globally in 2022 according to IFPI, and US$31.2 billion according to Statista.
Other than Netflix, I think FAANG just doesn't care that much about such a small market (the market being "actually producing it", given they're already part of the previous numbers for selling and streaming it).
And of course, both A's and the N of FAANG have their own commissioned TV/film content.
Yeah, doesn't remember. Mhm...
Oh, it just can't remember the license terms of the code it "reads", so it can't comply with these licenses or help people to comply with these licenses.
Convenient.
I suspect the answer to the question "is it, though?" is one for the lawyers and lawmakers rather than for the software developers, and it may well vary wildly by jurisdiction.
People dont want to acknowledge that the LLM structure reflects rather closely what it is being trained on, but the incredibly large number of parameters suggests it is closer to a photographic fit than a true abstraction. larger models being more likely to memorize training data (Carlini et al., 2021, 2022)
The fact that the information gets mangled and somewhat compressed doesnt change this close relationship.
But that's essentially what LLMs are doing, lossy compression of the entire web
If they had announced this sooner hardly anyone on the internet would have noticed. Props to them for adding it now.
There's also copyright reversion, which is a related new provision that applied to older copyrighted works. Quoting from an article I just pulled up
"...the 1976 Act created a new right allowing authors and their heirs to terminate a prior grant of copyright, the Act also set forth specific steps concerning the timing and contents of the termination notice that must be served in order to effectuate termination. The termination of a grant may be effective “at any time during a period of five years beginning of the end of 56 years from the date the copyright was originally secured”..."
But this is a red herring because the fact a model has been trained in the past doesn't mean a copyright lawsuit is "retroactive". The infringement would presumably be occuring anew every day you make it available on your web site.
Meta has been lobbying hard around that for years.
Only copyright can see through all of that, you would have to gut fair use in order to have an effective anti-scraping law.
I don’t think stack overflow is all that valuable once your model has access to github due to their good friends at MS.
The money in proprietary AI is on the top end now, open source / edge is destroying monetisation on the lower end. Top end means high quality domain specific data.
As a heavy ChatGPT user I disagree. Lack of up to date data is one of the biggest issues I face every day - technology changes fast, libraries change APIs, new tech comes out, etc.
As of my last knowledge update in September 2021, Quickwit is an open-source search engine infrastructure that is designed for building and deploying search solutions quickly and efficiently. It focuses on providing fast and scalable full-text search capabilities for applications and websites. Quickwit is built on top of the Rust programming language and leverages technologies like the tantivy search engine library.
[1] In my case, it relies on third-party APIs that are heavily rate limited. Any bot ignoring rate limitation measures will effectively (D)DOS my service.
Also if you're dealing with an actual malicious adversary real or automated rate limiting can be more effective than blocking. (logic to detect and overcome an even very significant rate limit is much more complex than to detect dropping, ignoring, or 4xx 5xx response blocking methods)
For example, a method to rate limit based on IP with nginx
http://nginx.org/en/docs/http/ngx_http_limit_req_module.html
Well, you're helping improving the model.
Physical property is stolen. Information is copied.
Physical property is either borrowed, owned, sold, and so on.
If your spouse takes your car to work without your knowledge it's borrowed. If they take it and sell it without consent it's theft.
Same applies to data. But data is electrons and as such it can't be moved, it is "copied". So technically speaking you are right, but practically you are not. If you steal NBC's prerelease movie then that's theft. As is copying it without constent. Once you pay for it you can copy it from their servers to your device. But you can't copy it to someone else's machine.
No. Advocates of expanded IP law have attempted to spread the idea that copyright infringement is "theft" as it adds emotional weight to their arguments. "You wouldn't download a car" etc. Same for the use of the word "piracy" - borrow an emotionally laden term from another context and hope nobody notices the sleight of hand.
And it's important that we reject this definition because it distorts the reality of the situation.
Depends who's reality. A content creator's reality is that their content is indeed stolen and monetised by someone without permission.
"Advocates of expanded IP law" do appear to be in the right, at least by law. Copying and distributing digital products is treated more or less as theft, particularly when done at scale.
AI and current training practices are even worse than stealing someone's work. It steals someone's identity. AI can copy unique characteristics, not just individual content to reproduce identical content. It can replicate a person's unique style without consent, and that's uniquely dangerous.
On a trivial level this is correct as words mean what we collectively decide they mean.
However I am making the point that a) the meaning has been changed and b) it has changed in a way that is deceptive and masks a useful fact about the world
Correct, and collectively we decided that reselling digital work without permission is indeed theft, just as we rightfully decided that digital goods for the most part are like physical goods.
> a) the meaning has been changed
It hasn't really, digital theft still has the same meaning as any form of theft. Some did try to change the meaning and non trivialise the act based on the fact that digital goods are not like phyisical goods. But that's a techincallity based on the nature of digital goods.
Similarly, AI folks wish to change the meaning of theft based on the false assumption that an AI system "learns just like a human". But that's a false assumption. The software does mimic human behaviour, but we all know that it is neither human nor intelligent (if it were intelligent you'd show it a set of multiplications, and from that point onwards it would figure it out on its own. same with writing stories). Yet some are trying to change the meaning of words to accommodate their view of the world in which software that can ingest people's IP at massive scale, mix it in, and output something that looks novel is somehow similar to human learning.
Therefore the matter is trivial. Software ingestsing digital content without permission, and outputting content made of even tiny bits of the original, is theft. Simple as that. However, that does not mean that AI should be banned. It's how the AI software is fed its data that must be brought in line.
You are a douchebag lol
But the street is one-way, and it's the company that has the upper hand. The company can (and does) retract access to the AI, but they themselves keep what they took. If in the meantime people became attached to what the company gave, the company even does damage to them, not just by taking away the access, but because of severing the supply for a dependency.
So the people are taken advantage of because the company took the assets, they are taken advantage of because they help to further train the AI by using it, and then they get, at most, the privilege to pay for something that grew out of them.
That's why it's not a win-win. It's a win for the company, and a questionable outcome, and a risk for the people.
Many good things don’t last forever. If they go away that doesn’t invalidate the experiences you had.
I do want to stress something from your conclusion though. That people do better if they anticipate change, and can adapt to it.
I don't expect compensation from anyone who uses them, whether it's public or private use, so I don't feel like I've lost anything. Sometimes people "pay it forward." If I actually get something back, that's a win.
There are web search engines and AI chatbots that might be very slightly better (unmeasurably so) due to having been trained on stuff I published over the years. Meanwhile I get a lot of benefit from using free stuff on the Internet. I think that's a one-sided deal in my favor.
(I also pay for GPT4 access. Whether it's worth $20 a month is more questionable, but it's fun to play with and so far I'm interested enough that I haven't cancelled.)
I completely agree. At the end of the day, winning and losing in this situation cannot be measured, especially the "losing" part wrt/ people, so it all boils down to how the individuals perceive it. (Which is of course why powerful entities put so much effort into PR.)
I personally feel better if there are some safeguards around usage, and so I like licenses like the GPL family, where regulations are in place so that the effort is not completely trivially closed up.
But really, at the end of the day what we can control best is our perception of thing. Life is what we make of it.
Sure a large org is receiving some ancillary benefit, but do you feel the same hostility for people working at [large corp] using what you worked on to help them at work?
I honestly don't understand the hostility towards llms using public data
The large org is receiving the greatest benefit.
But at the end of the day, I understand that altruism doesn't work this way. But this just means that while I have some tendencies, I'm not altruistic after all. I attach a lot of feelings to where my work ends up and how it affects things, which is, for example, why I like "sticky" licenses like the GPL, and tend toward efforts like the Effective Altruism, however ineffective I think they end up being.
>I honestly don't understand the hostility towards llms using public data
So, getting back to the topic, feelings are attached to where the publications end up and how it affects things. Because of the unintended consequence of companies training AI on publicly available data, people harboring these feelings feel like their thing has been taken from them without their consent. And that is a bad feeling, powerless, inability, and one of the ways of coping with that is coping with it on the outside, directing the feeling outward, whereby it becomes active defense, or hostility.
Well, because a lot of corporations couldn't care less about the public good and are happy to cause harm if it makes them more money. OpenAI doesn't care about your welfare or mine any more than a sleezy ad company or spyware product does.
If OpenAI were actually an open source company working to benefit the broader ecosystem I would agree with you, but that's about as far as possible from the current state.
Generally people need some kind of incentive to produce content. This could be just the thought of somebody, an actual human, having consumed your content. Or a like, a comment exchange that further enriches the topic. Perhaps it leads to a new follower or even a new (online) friend. A job opportunity. Even a date. Or maybe just plain ad impressions to make your effort worthwhile.
The picture of content production was already bleak. Google gets to take it all for free and is the traffic controller deciding who gets the crumbs, and even then is also the sole advertiser. But at least they might throw you some traffic, leading to all the interactions I just mentioned.
OpenAI just steals your shit without permission, credit or payment and completely cuts of any direct human interaction with the original content or its maker.
How can you not "understand" the hostility? This is existential not just for the open web, also the closed web. Have you missed the developments at Twitter, StackOverflow, Reddit?
If you don't want others to learn from what you have to say, just talk to a brick wall.
What do we want to teach it?
When you realize LLM models are very broad statistical models with nearly 0 sense at all they become easy to manipulate with wrong information.
The annoying thing is going to be LLMs teaching people things they publish and feed back into the next training of LLMs which will become pervasive to the extent that verifiable information will be much more difficult to come by and highly prized. Will drive even further nostalia or just real valuation of analog methods and artifacts and glitch/lofi/noise which are the kinds of abberation which analog systems make especially those that ML has difficulty emulating.
The PR industrial complex has been trying so hard to convince us that the all-knowing all-seeing almighty AI is going to take our jobs and turn us into Soylent or whatever. Now let’s feed it some garbage and see if in all its glory it can tell sense from nonsense.
# all nginx virtual sites
if ($server_protocol != HTTP/2.0) { return 302 https://auth.domain.tld$request_uri; }
# in auth.domain.tld virtual site
auth_delay 4s;
location / { auth_basic "Wamp Wamp"; auth_basic_user_file /dev/shm/.p; }Presumably, the incognito crawlers are only used on sites that have already granted the regular crawler access. That's content that ends up in their index which they want to vet.
https://www.theguardian.com/technology/2020/feb/03/berlin-ar...
Also, I would've sworn that happened circa 2015 and not 2020. The passing of time for the last few years has such a muddled feeling.
Feed them data created by LLMs.
EDIT: scuba diving NOT scooba diving
openai is worth $29,000,000.00 you contributed 0.00000000001
punches numbers in calculator
thus the value of your free credits is 0.001 cents. minus any accounting fees.
You might have missed a few zeros.
It is an interesting question. I would have no qualms paying for a textbook or university course for curated learning (worth noting OpenAI has paid datasets too), but paying for (or being paid for) relatively diffuse and low quality content through hobby blogs seems at odds with my expectations as an individual, and as a society we were never (en masse) concerned about things like Google's search excerpt answers...
If humans could perfectly remember information, I’m sure copyright would be very different.
With a gardened proprietary paywalled model, what I wrote ends up as some constituent of giant arrays of floating point numbers which I must pay to use.
If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me?
I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff is put out into the world for other humans to learn from and use to make themselves better, and that they don’t owe the original authors anything other than the price of admission.
I guess it comes down to this: do we think that training a model is:
- like storing and later reproducing a version of some collected data, or
- like learning from collected data, and synthesizing new info?
Is there even a meaningful distinction, for a computer?
(Is there even a meaningful distinction for a human…?)
You can go to a library to borrow a book, but you can't go to the library and copy all the books for your own use.
It's just impractical to photocopy every page of every book in a library.
Otherwise, intellectual property laws can perhaps apply.
It'd be a hard push to claim it's fair use, a wholesale copying of other's works.
Why? What's stopping me from doing that? The only limitation is time.
You can actually copy all the book but the things is you can't publish it as your own book after you copied. Because obviously it is not your work.
Things get interesting at corporate scale. There are fat VC funds, executives, board of directors and what not - making far more and far more comfortable than an individual trying to get better at their craft to put food on the table. And on top of that, you don't give me access to the product that was refined on my input.
It is like someone learning photography from my website but later taking a really masterpiece shot but asking me for money each time I want to view the photo in their studio.
There are no easy answers, I concur.
Thanks for your comment though, really. :)
> do I owe you 1% of what my clients pay me?
I would still derive some immaterial gain or satisfaction from you reading my website specifically and using what you learnt to improve yourself. As I expect most people would, so it's still a give and take relationship. LLMs sever that link.
It is doubtful many people will be as willing to continue "putting stuff out into the world" if they know that they are only contributing to some sort of (arguably semi-dystopian) hive-mind.
IMHO whether what they are doing or not is justifiable from a legalistic perspective is tangential and not that relevant if we're talking about free/non-commercial content.
Do they though? I mean, do you personally have a link to the people that are consuming the content you post publicly?
I find all the vitriol around LLMs being trained on public data to be a bit weird. If you don't want that data being used then don't publish it for the world to see? Why get mad when you are the one freely publishing the data in the first place? That's like posting your content on a bulleting board in the dorm common room and telling the trust-fund kids they can't read it because they are rich and you don't want them learning anything from you that might make them richer. Maybe a bad analogy, but I feel like it's a fair approximation of the vitriol I see.
Far in the future - if ever - where we have biological grade artificial beings which you can't program, control and limit in the classical software development sense, this could be rethought.
Until then, we don't need to humanize machines.
If I watch a youtube video my browser is also in a way scraping youtube and storing a (temporary) copy of the video. Does it make sense to protect the protect the owner's right's at this point? Absolutely not. Instead we wait to see if I share that downloaded video or content from it again, or somehow reuse it in my own products. Only then does the law step in.
But you don't get any of that from an LLM.
The reason copyright law exists in the first place is due to the difference of scale between copying books by hand and using a machine to do it, so I think "it's different because a machine is doing it" is a completely rational stance to take.
I believe no. Most people would make a distinction between “normal” and “rich”. They would give normal people free access, but the rich should pay for it.
It’s like a billionaire asking for a free hot dog. It’s like “come on, you can easily pay $100, which could even sponsor it for the next 100 people”.
Here it’s not the AI itself that’s exploiting you. It’s the rich people that make the AI that get even richer - partly thanks to your free work.
AI training is basically only extractive and has the potential to severely disrupt the actual field that made the AI systems possible at all. It's a much more mechanical process that the human interaction of studying a master. It doesn't develop any human skills.
Even if the processes were the same (and I don't think they are, as someone who has actually done computational psychology research), I would still think the AI companies are doing something they know is harmful to actual creative people that generate real value.
But you do get paid in kind - you "gave" information for the AI to train on, the aI gives you information back, contextualised to your needs. Sometimes those 1000 tokens are worth much more than $0.06
You still need to be able to pay for inference costs, it's crowded and expensive on GPUs nowadays.
If so, then pay.
You can't be serious. Thank god our world doesn't work like that.
What do you think why writers and actors have included AI in the reasons of their strike?
Because they are about to become obsolete, and they believe that screaming as loudly as they can is going to stop that.
Their chances of success are roughly the same as if they were protesting against the law of gravity.
We make our own rules. We decide what to allow and what to value. If technology changes something, it's because we let it.
AI will not make writers “obsolete”, that is utterly absurd. Would you say reality TV made tv writers obsolete? No? Oh well.
You get what you pay for. That includes what you pay for as a producer…
Of course. And those so-called "computers" won't make human calculators obsolete. After all, they are as large as an entire room, and by the time they are ready to receive input, a human with his slide rule has already computed three and a half entire logarithms!
Human creative professions have 5-10 years left, if they are very lucky.
So in that sense do developers have ~2 years left? Code is much more rigid than acting or creative writing and AI seems to be getting there first. I mean if the all-powerful AI can make modern movies than clearly it can handle writing all code right?
I think they're _very_ worried and rightfully so. I assume it would be very difficult to cancel an AI.
Yes, I might run a course or something. You are still not entitled to pay.
The moral thing to do would be to use opt-in training data.
A contrarian take to support the original commenter is that if the site owner had ads, i probably got him or her some increment in site visits and helped in some small way with monetization, site ranking and boosted his / her public persona, credibility.
When GPT bot visits, none of that happens. Much worse - people who might have visited the hobby site and contributed to traffic and ad revenue will now start getting their answers from the OpenAI chatbot and never visit this hobby site.
That's exploitation and I think that's what most of the responses on this thread miss.
When a private company takes the sum of human knowledge without permission, attribution or payment and then monetizes it via the back door whilst cutting of any connection between the intended consumer and publisher, then we're dealing with a system I'd describe as criminal. It cannot be morally defended as "fair" in any major economical or political system.
The fact that they call it "Open" AI shows the level of trolling involved.
This is a bit pedantic but the term is "scuba diving". Scuba is an acronym that's short for "self contained underwater breathing apparatus". It doesn't work if you don't spell it right.
If you can train your bot on my blog post about scuba diving without my permission and then people can ask your bot for scuba diving advice instead of reading my blog, that doesn't seem very fair.
[1]: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
No, you don’t.
That’s a factor weighing in favor of fair use, but the fair use factors are not defined in such a way that that is a necessary factor.
I think you'll find if you do try to push Google Search too far, its not quite "limitless" either.
Hasn't happened in a long time.
Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too.
How would contributing to the weights of this LLM help content creators?
On a second thought, I guess a lot of marketing content would also love to be crawled by anything that crawls…
But the book writer who wrote a detailed, expert book on how to deal with the software ("Photoshop for Dummies"?). OpenAI might be seen as a competitor.
A government would be easier to say all their data isn't allowed to be crawled, so they can sue later or just say no later on when they figure something classified was in there, or simply when they change their mind.
I believe the default response should be 'no, we'll look into it' for anyone, and then carefully let legal take a look at it (gonna be expensive). For the software vendor, too. Although their crown jewels are likely the source code to their product(s).
This is particularly weird since the EU Datamining directive that got us into the mess inside the EU seems to suggest that robots.txt seems to be a valid means to retain copyright for data mining (there is no 'fair use' otherwise inside the EU). Are there other machine-readable standards? I further don't quite understand, how EU copyright relates to training a model outside the EU and using it within again (probably this is the biggest enforcement gap)
It's true that after you use the bot to fetch you the papers, you do still need to read them... but given what a dramatic difference there is between GPT3 and 4 I'd say this is a problem that will be utterly annihilated before most people even hear it exists.
So that LLMs can learn from it? Profit is not the only thing that motivates people. I’ve spent years contributing to Stack Overflow to help people solve their problems, with the understanding that they had an open data policy and anybody could access the data dump easily to build things with it. It pisses me off that they are now trying to lock that information away where LLMs can’t access it. The whole reason to contribute is to help people. Locking that information away instead of exploiting this new channel to help people more effectively is antithetical to the reason I contributed in the first place.
With Stack Overflow, everybody contributed to their data set. This data set is centrally managed by Stack Overflow and access is whatever they choose to allow. When they block access to that data set, it effectively takes it away from the public.
With OpenAI, they aren’t locking anything away. They are analysing the data and adjusting the weights in their model. They haven’t stopped people from accessing the data they are training upon.
What Stack Overflow are doing is stopping the free flow of information. What OpenAI are doing is providing an additional channel for it to flow through.
Well, "more accurate" means roughly: "So that your content can be absorbed and used as output by our generator."
Google at least linked to your website, while ChatGPT hides your website and only uses your content.
And now they also put many links of SEO spam with ads firsts.
Same with this crawling by bots (Google, Bing, Meta, OpenAI; doesn't matter). Jurisprudence on Google News and Google Cache seems to show citing is OK, if done in moderation. Remember: just because you can access (download) something on the internet (WWW or otherwise) does not mean you're allowed to watch, use, save it. That argument was lost during the battles of copyright infringement in the years of 2000s.
OpenAI isn't even citing in moderation. Its making a derivative work without citing (hence obscuring) it does.
The bottom line is this: ML which doesn't cite sources should be regarded as hostile: a blackbox, and a copyright infringement paradise.
Does it? Last time I checked, “cloud” in “Creative Cloud” meant “now you have to pay a monthly subscription”.
And reverse engineering Photoshop to make a competitor might be a legal practice, if done properly – for example, see the ReactOS project.
Aside from that, I think your point still stands though.
I doubt it's copyright infringement in this case, at most it's just against the ToS.
>Its making a derivative work
A derivative work includes major copyrightable elements of a first, previously created original work, and that's how it's treated in court. Most AIs will not generate derivative works (unless you ask them to).
I'm considering whether to add it to my own site, but given that the future is already here and, while it's shitty to steal and regurgitate content without attribution at minimum, it's also not a big deal for my hobby site. It may serve my interests better to not include crawling restrictions for ClosedAI specifically
> To disallow GPTBot to access your site you can add the GPTBot to your site’s robots.txt
Too late - they already grabbed content from my personal website.
Regarding noai tags - is this respected or just wishful?
<meta name="robots" content="noai, noimageai">However, it's trivial to know whether the bot crawled your site or stopped at robots.txt.
I'm not sure this first reasonably working chat bot is going to be the last version we ever need, and afaik this sort of thing is as hard to port as it is to untrain, the problem in both cases being that it's a big black box
I can see a bots.txt entry in the near future that discern the site's data-usage for bots vs humans
User-agent-class: AI
Data-Policy-Allow: /news/* /articles/*
Data-Policy-Deny: */commentsAlthough no human is going to read robots.txt or bots.txt
It'll end up as a small section in the EULA of the website which nobody reads:
> Before you click the 'reply' button please be aware we are allowing AI to crawl our comment section for training. Thank you for your consideration.
There's a little problem though:
1) Websites don't have an incentive to inform their users about this, and websites don't have an incentive to allow AI to crawl their content unless they get something back from it (e.g. payment). From this PoV, its time for OpenAI to start paying.
2) The competition (China, Russia) doesn't care about bots.txt or robots.txt and will just crawl whatever the hell they can.
My reasons are:
- I don't recall seeing any evidence that OpenAI has included new data in pretraining beyond the previous limit (Sept. 2021?) for GPT-3.5 or GPT-4
- Maybe they did finetuning or RLHF on new data but this is likely to be highly curated data
- AI generated content should be absolutely tiny in comparison to the data they are already working with.
[1]: LLM Engine Optimization
So, I doubt this is the plan.
If (generally) more data is better. As a site owner I might be happy to give free access to most of my public data and urls. You be at the mercy of my sites unique formatting and hiccups
But for some fee, I might be happy to provide api-like data access to some of my historic data that are more rich with information and promised some sort of format-encoding (AI-JSON ?)
The economics I think are still currently be discovered.
> For robots.txt, we do follow the same restrictions applied to googlebot, otherwise Google benefits from its dominant position.
https://community.brave.com/t/stop-website-being-shown-in-br...
would it fall under the cfaa?
I'm confused.. if it's being trained on data up to a certain date, than why would the web crawler matter?
I believe the chatbot is prompted to not answer for things after Sept 2021, rather than the data itself being limited.
I could be wrong though.
In most cases I do think so. It could mention you in a conversation, the analog to you appearing in Google Search results. And maybe even better, provide the necessary context to generate more real customers for your business. You don’t want traffic to your website, you want customers to your business. If you currently convert 10% of your traffic to customers, you’d be happy with 10% of the traffic of which you convert all to customers, because they are already converted before they even clicked your link.
I'm working on something like that as well.
Allow: /directory-1/
Let's all make this "directory-1" and fill it with absolute garbage.
Also, where is this data going? The existing ChatGPT says it has nothing past 2021.
To the next ChatGPT
"To allow GPTBot to access your only parts of your site"
What will be the point of sharing knowledge or content if it will no longer be associated to an individual or organization?
Some people publish content freely on the Internet as a form of note taking, publicity, public discourse or for the betterment of like minded individuals, akin to why we’re here commenting on HN.
I guess I assume public content defaults into this “for the benefit of the world” category, where it’s up to the publisher to gate content as desired.
Public content is still mostly published with a reference to a certain or anonymous individual/organization and gains visibility based on a value and an effort to be seen. Individual/organization is still motivated by visibility, popularity, acceptance and approval of that content.
We can summarize that people are motivated by a reaction. What do you think will happen when you remove or decrease reaction to knowledge/content providers?
Of course if people do conclude that there’s no benefit to sharing knowledge and stop doing so then those who do share knowledge will have an outsized impact on AI training. In the extreme: the opportunity to create truth. Thus the incentive to publish in order to stop those people is created.
Sorry for the sarcasm, but your comment is essentially "Lets remove motivation for those who had incentive to share valuable information and see how it turns out.".
One could of course debate whether or not OpenAI would be the best stewards of that knowledge or aligned with the best interests of humanity. However, it is important to recognize that building a successful business is key to funding the research, H100s aren't cheap. It's also important to note that as with all things tech, price of hardware will go down, and OSS models continue to get more capable every week.
I repeat, what will be the motivation of an individual to share or provide valuable information if you decrease or eliminate any control of where and how that information appears?
Forum users don’t seem to mind. Reddit, HN, Twitter, Facebook, etc. are all examples of users freely providing valuable content without expectation or control.
I suppose it’s also not too different from a listener summarizing a speech. When you speak publicly you don’t get to control who hears it or how they will interpret it.
In the meantime, I did relize I might have overblown consequences of what a product collecting and summarizing knowledge might cause.