Every so often I get lost in the docs trying to do something that actually isn't supported (the library has some glaring oversights) and I'll search on Google to see if anyone else came up with a similar problem and solution on a forum or something.
Instead of telling me "that isn't supported" the AI overview instead says "here's roughly how you would do it with libraries of this sort" and then it would provide a fictional code sample with actual method names from the documentation, except the comments say the method could do one thing, but when you check the documentation to be sure, it actually does something different.
It's a total crapshoot on any given search whether I'll be saving time or losing it using the AI overview, and I'm cynically assuming that we are entering a new round of the Dark Ages.
It's also obnoxious on mobile where it takes up the whole first result space.
In theory, an AI should be able to fetch the llms.txt for every library and have an actual authoritative source of documentation for the given library.
This doesn't work that great right now, because not everyone is on board, but if we had llms.txt actually embedded in software libraries...it could be a game changer.
I noticed Claude Code semi regularly will start parsing actual library code in node_modules when it gets stuck. It will start by inventing methods it thinks should exist, then the typescript check step fails, and it searches the web for docs, if that fails it will actually go into the type definition for the library in node_modules and start looking in there. If we had node_modules/<package_name>/llms.txt (or the equivalent for other package managers in other languages) as a standard it could be pretty powerful I think. It could also be handled at the registry level, but I kind of like the idea of it being shipped (and thus easily versioned) in the library itself.
But isn't the entire selling point of the LLM than you can communicate with it in natural language and it can learn your API by reading the human docs?
Good for humans example: https://docs.expo.dev/llms-full.txt
Bad for humans example: https://www.unistyl.es/llms-small.txt
Would you say "pah why are you shipping a search engine that only sometimes finds what I'm looking for?"?
If there's nothing answering what I was looking for, I might try again with synonyms, or the think documents aren't indexed, or they don't exist.
That's a very different failure mode than blatantly lying to me. By lying to me, I'm not blaming myself, I'm blaming the AI.
For troubleshooting an issue my prompt is usually “I am trying to do debug an issue. I’m going to give you the error message. Ask me questions one by one to help me troubleshoot. Prefer asking clarifying questions to making assumptions”.
Once I started doing that, it’s gotten a lot better.
“There's a library I use with extensive documentation- every method, parameter, event, configuration option conceivable is documented.”
This is the perfect use case for ChatGPT with web search. Besides aside from Google News, Google has been worthless to find any useful information for years because of SEO.
It’s the classic XYProblem.
So don’t worry about writing that documentation- the helpful AI will still cite what you haven’t written.
this rhymes a lot with gangsterism.
if you don’t pay our protection fee it would be a shame if your building caught on fire.
In many, if not most cases, the producers of this information never asked for LLMs to ingest it.
This seems more like a model-specific issue, where it’s consistently generating flawed output every time the cache gets invalid. If that’s the case, there’s not much Google can do on a case-by-case level, but we should see improvements over time as the model gets incrementally better / it becomes more financially viable to run better models at this scale.
If you don't want hallucinations, you can't use LLM, at the moment. People are using LLM, so having giving it data, to hallucinate less, is the only practical answer to the problem they have.
If you see another, that will work within the current system of search engines using AI, please propose it.
Don't take this as me defending anything. It's the reality of the current state of the tech, and the current state of search engines, which is the context of this thread. Pretending that search engines don't use LLM that hallucinate data doesn't help anyone.
As always, we work within the playground that google and bing give us, because that's the reality of the web.
1. Start by drawing some circles.
2. Erase everything that isn't an owl, until your drawing resembles an owl.