No idea how much openai's computational cost is per query. Unless it's an order of magnitude higher than google's, we can assume the next thing after yahoo -> altavista -> google is here.
No idea how much openai's computational cost is per query. Unless it's an order of magnitude higher than google's, we can assume the next thing after yahoo -> altavista -> google is here.
But it's an incredible tool for brainstorming or generating content. I think that soon a large percentage of all online text content will be GPT-generated, and that comes with a lot of new issues that we're not prepared for. It's going to be really difficult to trust anything online and tell fact from fiction.
However after googling the actual API, turns out ChatGPT's answer, while convincing, was utter rubbish.
It will however index hallucinated results generated with GPT and published somewhere, so once we're at that point it really doesn't matter anymore.
Some kind of hybrid of this and search would be great.
Google just gives you associations provided by random other people on the internet. It's largely garbage, most often deliberately disingenuous (to make you look at an ad). Ad revenue models for the internet encourage the generation of this type of false material.
A better criticism would be that the same thing will happen to something like chatgpt -- and the question is whether the model for analysis can better handle it at scale.
No, it absolutely does not. Yes, there is SEO spam in the index, but no - it is not Google hallucating it. It really exists on the internet, see also the second point of my comment.
> the same thing will happen to something like chatgpt
This isn't something that "happens to" GPT, GPT is doing it. There's probably even already GPT -> SEO spam pipelines out there generating websites.
The exact same thing is true for chatGPT, or any other computer system. It is providing information and associations based on the input dataset.
> "GPT is doing it."
And google is "doing it," when google decides there is an association between my query and a bad response. Both systems are analyzing a corpus, drawing associations, and returning parts of that corpus. The output is deterministic based on the input.
The types of associations differ in their depth, but there is no fundamental difference in terms of agency or outcome.
No, it's not, for example if you ask google to show you papers about some topic with words in quotes you think you remember from the paper it will show you the proper link IF it exists and language model will just generate you a result that doesn't exist.
If I search something on Google that doesn't exist or that it have no answer I can see looking at list of search results that probably either what I look for doesn't exist or my assumption is false but language model will generate you plausible explanation/answer that can be 100% false and it doesn't know or understand that it's false and you will have no way to know if it's false or true and no point of reference because ALL the results you will receive could be hallucinated.
Not really.
What will actually happen is Google may link me the actual research paper, along with thousands of other associated pages which may or may not have all kinds of fake, false information. For example, if I search for "study essential oils treat cancer" I get an estimated 190 millon matching documents. A huge percentage of these have false and misleading information about using oils to treat cancer.
> "language model will generate you plausible explanation/answer that can be 100% false and it doesn't know or understand that it's false and you will have no way to know if it's false or true and no point of reference because ALL the results you will receive could be hallucinated"
With google it is third parties "hallucinating" the wrong answers (or worse: intentionally answering wrong in order to exploit and profit). The overall dynamic is not different. Google is providing these wrong answers, written by others.
The overall dynamic of providing information of questionable veracity is generally the same - because the question of who creates the associations and incorrect content is not particularly germane.
No, what I'm doing is considering the web content (incentivized by the ad systems Google provides) with the web search systems.
> "and not only it can return answer that was NOT in the dataset but it can make it plausible looking."
This is exactly what SEO spammers do. It's the same.
I wouldn't categorise that as "hallucinating fictitious results" - the algorithms still only returns existing results. If you follow the link, you will find key words embedded in the HTML or visible text in the browser.
Different kettle of fish entirely.
But scientific facts are a different story. Nobody has any incentive to claim that 1+2=4 or that some function in Python does X when it really does Y. So when you search for these kind of facts on Google you can pretty sure that you get correct answers, or at least someone trying to give you the best answer they can. But not so with GPT. It may give you incorrect answers even for these kind of facts if they are not within the reasoning ability / training data.
Incentive is irrelevant. What mattes is whether these things do happen, irrespective of intent -- and they do! I very, very frequently find incorrect answers to math questions, tech function questions, etc.
Incentive is an important part of the dynamic, but it's not important to consider if we're looking empirically at the integrity of the results.
> "So when you search for these kind of facts on Google you can pretty sure that you get correct answers, or at least someone trying to give you the best answer they can. But not so with GPT."
It is so with GPT. Both systems are "trying to give you the best answer."
I think what you're observing is that the Google search engine has two decades and billions of dollars behind it and ChatGPT is a research preview - not even a finished product.
I remember using search engines in the late 90s (in fact, I worked on one of the leading ones). I think you are extending far too much credit.
No, based on your responses you do not understand how language model works. Google is searching in index using keywords and rankings, ChatGPT is predicting plausible words without searching anything anywhere.
What you argue is like saying there is this two guys in library and you ask them to find you something that exists or maybe doesn't exists, both have read all the books, one (Google) have created index of all the words from the books and is going through it to answer you and the other (ChatGPT) do not use any index but he uses his memory with compressed knowledge of statistics between words and will answer by trying to predict any answer that fits statistics between words and in many cases it will basically lie to you and you will have no clue that you were lied to.
There is distinction between indexing human knowledge about some topic where most of the top results are correct (Google) and creating statistics model between words and making things up that never existed and are wrong (ChatGPT).
Expand your scope to both Google, and the creation of an ecosystem of SEO pages which Google incentives. They are the same, in totality. Google doesn't just index -- it also funds the creation of landing pages.
> "There is distinction between indexing human knowledge ... and creating statistics model between words and making things up that never existed and are wrong "
It's a false distinction. Google is more than a search engine; it is also an advertising company that incentivizes original content creation with the express intent of providing answers to queries.
Obvious straw man argument. Replace word google with search engine.
> Expand your scope to both Google, and the creation of an ecosystem of SEO pages which Google incentives. They are the same, in totality. Google doesn't just index -- it also funds the creation of landing pages.
This doesn't matter, you are mistaking dataset with the model. Search engine will not return to you things that were not in dataset, it will give you many results that you can judge with many points of reference. Language model will return you one answer, answer that could be a correct result that is inside the dataset or could be totally false and incorrect but plausible and you will have no point of reference to check that unless you use a real search engine.
No, I don't think I will. We are talking about the system as a whole.
> " Search engine will not return to you things that were not in dataset"
Yes it will, because the Google/SEO dynamic is in large part about incentivizing the generation of new content.
I understand you want to narrowly define away this fact to make a point, but the fact stands.
ChatGPT doesn't, so you don't know where it got its information, and you'd have to do web searches anyway to get some possible sources.
The trick is to feed relevant contextual information instead of using it closed-book. This can be automated with a search engine, or can be a deliberate manual process. But closed-book mode is not the right way to assess people or AIs.
What are your counter arguments?
I’m thinking ChatGPT is best used for generating ideas, not factual information.
This was my prompt: What's the relative distance between Sun and its planets compared to the size of Sun?
ChatGPT answer:
> The average distance between the Sun and its planets is approximately 5% of the Sun's diameter.
If you run into bunch of such answers, you will soon stop trusting it for anything.
As for Google - it doesn't even try to answer the question. But no answer is often better than a wrong one. Google at least gives me link to websites that have enough data for me to calculate it myself.
In the end convenience wins out at cost of correctness. The proportion matters, convenience::correctness.
It is not the start. It is at least the third iteration (if not more) of this tech.
> In the end convenience wins out at cost of correctness.
There's nothing convenient about a wrong answer. Wrong answers have no value. I could very easily build you a convenient service that is rarely correct but nobody would care.
> The distances between the Sun and its planets vary greatly. For example, the distance from the Sun to Mercury is about 36 million miles, while the distance from the Sun to Pluto is about 3.6 billion miles. The size of the Sun is about 864,000 miles in diameter, which means that the distance from the Sun to its planets is much greater than the size of the Sun itself.
Then it says something that doesn't follow, probably answering the more frequent question of relative sizes of the planets compared to the Sun
> To put it into perspective, if the Sun were the size of a basketball, the nearest planet, Mercury, would be about the size of a grain of sand, and Pluto would be about the size of a golf ball.
>The distance between the Sun and its planets is relatively small compared to the size of the Sun. The Sun is a massive object, with a diameter of about 1.4 million kilometers, while the distance between the Sun and its planets can vary greatly. For example, the distance between the Sun and Earth is about 150 million kilometers, which is about 100 times the diameter of the Sun. Similarly, the distance between the Sun and the outermost planet in the solar system, Neptune, is about 4.5 billion kilometers, which is more than 3,000 times the diameter of the Sun. So while the distances between the Sun and its planets are not insignificant, they are still relatively small compared to the size of the Sun itself.
I wonder if it chooses units based on where you are located, or if it's just arbitrary. Also, it seems clear from this example that is has no context for the answer because it doesn't see that its first statement is contradicting its next statements.
I also changed the Temperature from 0 to 0.5, and it gave the right answer:
Q: What's the relative distance between Sun and its planets compared to the size of Sun? Show your math.
A: The relative distance between the Sun and its planets is approximately 1/100th the size of the Sun. This can be shown mathematically by calculating the ratio of the radius of the Sun (6.96x10^8 m) to the average distance of the planets from the Sun (1.5x10^11 m), which gives a ratio of 1/100th.
It seems that the developers have placed guardrails around web-search-like queries not because ChatGPT can’t answer them, but because they want to discourage using it that way for—I’d guess because they want to direct usage towards the conversational / contextual aspects they’re trying to improve.
There's also WebGPT[0] already with such capabilities, which could've been merged to ChatGPT.