Deliberately putting garbage data in your llms.txt could be funny though.
Deliberately putting garbage data in your llms.txt could be funny though.
I think you should think about it as: I want the LLM to recognize my site as a high quality resource and direct traffic to me.
Imagine user asks ChatGPT a question. LLM has scrapped your website and answers the question. User wants some kind of follow up - read more, what's the source, how can I buy this, whatever - so the LLM links the page it got the data from.
LLMs seem like they're supplanting search. Being early to work with them is an advantage. Working to make your pages look low quality seems like an odd choice.
Whether you like it or not LLMs are going to be how people explore the web. They simply work better than search engines - not least because they can quickly scan numerous sites simultaneously, consume and synthesize the content.
You can choose to sabotage your own content in a likely futile effort to make things worse for LLM users if you want - my point is just that it serves no purpose and misses out on the opportunities in front of you.
Obviously, they would not make it just for an AI company to scrape
Here's an example. Let's say I run a dev tools company, and I want users to be able to find info about me as easily as possible. Maybe a user's preferred way of searching the web is through a chatbot. If that chatbot also uses llms.txt, it's easy for me to deliver the info, and easy for them to consume. Win-win
Of course adoption is not very widespread, but such is the case for every new standard.
https://arxiv.org/abs/2410.13722v1
I have noticed some popular copied but incorrect leetcode examples leaking into the dataset.
I suspect it depends on domain specificity, but that seems within the ability of an SEO spammer or decentralized group of individuals.