llms.txt directory
directory.llmstxt.cloud
directory.llmstxt.cloud
No one wins in the long run by creating technical solutions to human incentive problems. It is just a prolonged arms race until
* the incentives are removed
* the process is made so technically complex or expensive that only a few players can profit from them
* it is regulated such that people can make money doing other things which have better risk/reward
* most people just avoid the whole ecosystem because it becomes a cesspool
Step 1: Punish ranking and visibility of sites whose llms.txt differs from a random sampling of actual web (HTML) content.
Step 2: There is no step 2.
1. Figure out how to embed content that only LLMs see which affect their output
2. Wait for that to stop working
3. Innovate another way to get past new technical problem
But maybe this time the naive technical solution will work
You, me, LLMs, Google, humanity, Earth, Sol...
We can choose to carry on and perform [what we think are] improvements and make the best of it, or we can choose to cash it in early and simply give up.
It looks like its text meant to be fed into the llm as a system prompt specific to the site.
The most simple ones just look like a sitemap restricted to documentation: https://www.activepieces.com/docs/llms.txt
Some interesting stuff is in some of them. Like this one that prompts the LLM to explain to the user the ethical issues of using AI agents along with a disclaimer:
That site appears to be someone's blog and they don't seem like big fans of LLMs.
https://boehs.org/node/llms-destroying-internet
A pretty clever use of llms.txt.
It’s not related to model training. Nearly all the responses so far are about model training, just like last time this came up on HN.
For instance, I provide llms.txt for my FastHTML lib so that more people can get help from AI to use it, even although it’s too new to be in the training data.
Without this, I’ve seen a lot of folks avoid newer tools and libs that AI can’t help them with. So llms.txt helps avoid lock-in of older tech.
(I wrote the llms.txt proposal and web site.)
In this case, we might need a versioning scheme. Libraries have multiple versions and not everyone is on the latest. I still need a way to point my LLM to the ver I'm actually using.
Is there any evidence that the presence of the llms.txt files will lead to increased inclusion in LLM responses?
I also don’t understand the problem it purports to solve.
Deliberately putting garbage data in your llms.txt could be funny though.
https://arxiv.org/abs/2410.13722v1
I have noticed some popular copied but incorrect leetcode examples leaking into the dataset.
I suspect it depends on domain specificity, but that seems within the ability of an SEO spammer or decentralized group of individuals.
I think you should think about it as: I want the LLM to recognize my site as a high quality resource and direct traffic to me.
Imagine user asks ChatGPT a question. LLM has scrapped your website and answers the question. User wants some kind of follow up - read more, what's the source, how can I buy this, whatever - so the LLM links the page it got the data from.
LLMs seem like they're supplanting search. Being early to work with them is an advantage. Working to make your pages look low quality seems like an odd choice.
Whether you like it or not LLMs are going to be how people explore the web. They simply work better than search engines - not least because they can quickly scan numerous sites simultaneously, consume and synthesize the content.
You can choose to sabotage your own content in a likely futile effort to make things worse for LLM users if you want - my point is just that it serves no purpose and misses out on the opportunities in front of you.
Obviously, they would not make it just for an AI company to scrape
Here's an example. Let's say I run a dev tools company, and I want users to be able to find info about me as easily as possible. Maybe a user's preferred way of searching the web is through a chatbot. If that chatbot also uses llms.txt, it's easy for me to deliver the info, and easy for them to consume. Win-win
Of course adoption is not very widespread, but such is the case for every new standard.
Plus, why wouldn’t llms be able to crawl websites as well?
Are llms going to improve the web or fragment it further as ads did?
(but we'll just ignore the obvious irony in that end bit about detection of bots getting smarter... wonder where all this "intelligence" will come from? probably not some natural source, but possibly some sort of... Antinatural Intelligence?)