I'd even question that. The pre-LLM solutions were in most cases better. Searching a maintained database of curated and checked information is far better than LLM output (which is possibly bullshit).
Ditto to software engineering. In software, we have things call libraries: you write the code once, test it, then you trust it and can use it as many times as you want forever for free. Why use LLM generated code when you have a library? And if you're asking for anything complex, you're probably just getting a plagiarized and bastardized version of some library anyway.
The only thing where LLMs shine is a kind of simple, lazy "mash this up so I don't have to think about it" cases. And sometimes it might be better to just do it yourself and develop your own skills instead of use an LLM.
It's better to take an existing, already curated and tested library. Which, yes, may have been generated by an LLM, but has been curated beyond the skill of the LLM.
If you really can one shot it and it’s simple(left-pad). Great. But most things aren’t my that simple, the third time you have to think about it, it’s probably a net loss.
If you ask for a link, it may hallucinate the link.
And unlike a search engine where someone had to previously think of, and then make some page with the fake content on it, it will happily make it up on the fly so you'll end up with a new/unique bit of fake documentation/url!
At that point, you would have been way better off just... using a search engine?
Which would certainly explain things like hallucinated references in legal docs and papers!
The reality is that for a human to make up that much bullshit requires a decent amount of work, so most humans don’t do it - or can’t do it as convincingly. LLMs can generate nigh infinite amounts of bullshit for cheap (and often more convincing sounding bullshit than a human can do on their own without a lot of work!), making them perfect for fooling people.
Unless someone is really good at double checking things, it’s a recipe for disaster. Even worse, doing the right amount of double checking makings them often even more exhausting than just doing the work yourself in the first place.
One time I tried to use Gemini to figure out 1950s construction techniques so I could understand how my house was built. It made a dubious sounding claim about the foundation, so I had it give me links and keywords so I could find some primary sources myself. I was unable to find anything to back up what it told me, and then it doubled down and told me that either I was googling wrong or that what it told me was a historical “hack” that wouldn’t have been documented.
These were both recent and with the latest models, so maybe they don’t fully fabricate links, but they do hallucinate the contents frequently.
Grok certainly will (at least as of a couple months ago). And they weren't just stale links either.
Eventually it gave up and commented out all the code it was trying to make work. Took me less than two minutes to figure out the solution using only my IDE's autocomplete.
It did save me time overall, but it's definitely not the panacea that people seem to think it is and it definitely has hiccups that will derail your productivity if you trust it too much.
"Tell me how to do X" (where X was, for one recent example, creating a Salt stanza to install and configure a service).
I do as it tells me, which seems reasonable on the face of it. But it generates an error.
"When creating X as you described, I get error: Z. Why?"
"You're absolutely correct and you should expect this error because X won't work this way. Do Y instead."
Gah... "I told you to do X, and then I'm going to act like it's not a surprise that X doesn't work and you should do something else."
Now instead of the wikipedia article you are reading the exact same thing from google's home page and you don't click on anything.
It’s for queries that are unlikely to be satisfied in a single search. I don’t think it would be a negligible amount of time if you did it yourself.
I let Claude and ChatGPT type out code for me, while I focus on my research
On the other hand, where I think llms are going to excel, is you roll the dice, trust the output, and don't validate it. If it works out yayy you're ahead of everyone else that did bother to validate it.
I think this is how vibe coded apps are going to go. If the app blows up, shut down the company and start a new one.
wondering how is it going to work when they "search the web" to get the information, are they essentially going to take ad revenue away from the source website?
I think we all understand that at this point, so I question deeply why anyone acts like they don’t.
More convenient than traditional search? Maybe. Quicker than traditional search? Maybe not.
Asking random questions is exactly where you run into time-wasting hallucinations since the models don't seem to be very good at deciding when to use a search tool and when just to rely on their training data.
For example, just now I was asking Gemini how to fix a bunch of Ubuntu/Xfce annoyances after a major upgrade, and it was a very mixed bag. One example: the default date and time display is in an unreadably small "date stacked over time" format (using a few pixel high font so this fits into the menu bar), and Gemini's advice was to enable the "Display date and time on single line" option ... but there is no such option (it just hallucinated it), and it also hallucinated a bunch of other suggestions until I finally figured out what you need to do is to configure it to display "Time only" rather than "Data and Time", then change the "Time" format to display both data and time! Just to experiment, I then told Gemini about this fix and amusingly the response was basically "Good to know - this'll be useful for anyone reading this later"!
More examples, from yesterday (these are not rare exceptions):
1) I asked Gemini (generally considered one of the smartest models - better than ChatGPT, and rapidly taking away market share from it - 20% shift in last month or so) to look at the GitHub codebase for an Anthropic optimization challenge, to summarize and discuss etc, and it appeared to have looked at the codebase until I got more into the weeds and was questioning it where it got certain details from (what file), and it became apparent it had some (search based?) knowledge of the problem, but seemingly hadn't actually looked at it (wasn't able to?).
2) I was asking Gemini about chemically fingerprinting (via impurities, isotopes) roman silver coins to the mines that produced the silver, and it confidently (as always) comes up with a bunch of academic references that it claimed made the connection, but none or references (which did at least exist) actually contained what it claimed (just partial information), and when I pointed this out it just kept throwing out different references.
So, it's convenient to be able to chat with your "search engine" to drill down and clarify, etc, but a big time waste if a lot of it is hallucination.
Search vs Chat has anyways really become a difference without a difference since Google now gives you the "AI Overview" (a diving off point into "AI Mode"), or you can just click on "AI Mode" in the first place - which is Gemini.
Everyone is entitled to their own opinion, but I asked ChatGPT and Claude your XFCE question, and they both gave better answers than Gemini did (imo). Why would you blindly believe what someone else tells you over what you observe with your own eyes?
If you're using ChatGPT like you use Google then I agree with you. But IMO comparing ChatGPT to Google means you haven't had the "aha" moment yet.
As a concrete example, a lot of my work these days involves asking ChatGPT to produce me an obscure micro-app to process my custom data. Which it usually does and renders in one shot. This app could not exist before I asked for it. The productivity gains over coding this myself are immense. And the experience is nothing like using Google.
It might seem quaint today but one example might be fact checking a piece of text.
Google effectively has a pretty good internal representation of whether any particular document concords with other documents on the internet, on account of massive crawling and indexing over decades. But LLMs let you run the same process nearly instantly on your own data, and that's the difference.