Microsoft is preparing to add ChatGPT to Bing
bloomberg.com
bloomberg.com
- Why would I ever let bing crawl my site if they aren't going to send any visitors to me?
- fair use of snippets has relied on them being brief and linking to the source. Lawsuits will be immediate.
I do love these imaginary scenarios where ChatGPT is going to find me the best air fryer, though. Where is that information going to come from, exactly? Barely anyone is making money writing reviews today, most sites are farmed content. What happens when even the ok sites' reviews are quickly scraped and put into the next model iteration? Bing is going to have to come up with some kind of radical revenue sharing too if they want anything written after 2023.
Websites allow SE crawlers because (a) whatever traffic they get is better than not traffic (b) because allowing crawlers is default and doesn't cost anything and (c) google/bing don't negotiate. They are one, sites are many.
This has already played out in news. News outlets wanted Google to pay for content. Google (initially) responded by allowing them to opt out of Google. Over the years, they have negotiated a little bit. Courts, in some places, forced Google to negotiate... It's news and politicians care about news specifically. Overall though, there have not been meaningful moments where people got pissed off with Google and blocked crawlers. Not newspapers and not anyone else. Site owners being mad doesn't affect google or Bing.
What does matter to search engines is walled gardens. Facebook pioneered this, and this does matter to Google. There is, in a lot of cases, a lot less content to index and serve users. All those old forums, for example.
These are search problems, and GPT-based search will inherit them. ChatGPT will have the same problem recommending the best air fryer as normal search does. GPT is a different way of presenting information... it's not presenting different information.
RE: Lawsuits. Again, history. Youtube, for example, started off with rampant copyright infringement. But, legal systems were primitive. Lawyers and legislatures didn't know what to do. Claimants were extremely dispersed, and would have had to pioneer case law. Ultimately, copyright took >10 years to really apply online and by that point youtube and other social media was entrenched.
The law lags. In practice, early movers are free to operate flawlessly and they get to shut the door after them. Now that Google is firmly entrenched, copyright law serves as one of their trenches.
No. The weird thing is this idea that because you put ads on your site, you deserve money. Your ads are making the Internet worse. You probably don't realize this, because you most-likely use an ad blocker, which means you want people too dumb to use ad blockers to subsidize the web that you can use for free, but the current web is working well for approximately no one.
Would I pay $5 a month for StackOverflow if it didn't show up for everything I Google? Most likely. Would this be a better world? almost certainly. We tried the thing with ads. It sucks. I welcome our new AI search overlords.
ChatGPT doesn't include anything written after 2021. I certainly wouldn't use it to find an air fryer. The results will be from over a year ago. I would want to see what the newest air fryer options are and it would be really important to have to up to date pricing.
AFAIK there is not a way to update a large language model in real time. You have to train on the entire dataset to do a meaningful update, just like with most forms of neural networks. For ChatGPT that takes days and costs hundreds of thousands of dollars. Every time.
It's great for explanations of concepts, and programming, and a few other things. But with the huge caveat that all of the information you're looking at is from one year ago and may have changed in that time. This really limits the utility of ChatGPT for me.
Manufacturers, with quality ranging from excellent to trash.
Consider trying to buy a 1K resistor at Digikey using their parametric search. Possible, but tedious and time consuming because you need a lot of domain knowledge to know what you want, and the technological range of "things with 1K of resistance" is extremely vast. At least its possible because the mfgrs are honest when Digikey imports their data.
Consider the opposite, consumer goods. 500 watt PC power supplies with random marketing number stickers on the same chassis ranging from 500 to 1200 watts. Consumer level air compressors and consumer level vacuum cleaners than plug into household wall outlets claiming "8 horsepower" or whatever insane marketing nonsense. Clothes with vanity sizing so a "medium" tag fits like a real world XXL. Every processed food in a store with a "keto" label is high carb sugar added garbage, much like happened with "organic" label in the old days (the employees at the farm, distributor, warehouse, and/or retail store level take the same produce out of one bin and put it in two places with different prices)
I think it will help when purchasing technical engineering type products but be an epic fail at inherently misleading consumer goods.
Making money from search traffic to your (presumably useful) site is going to get harder in a bunch of ways, due to generative models.
I'm sure ChatGPT will be able to write a bunch of terrible SEO prose that precedes the actual air fryer review (or worse, recipe) about how the author's grandma had an air fryer when she was young and remembered the great times with her grandma (etc), for roughly 95% of the text!
In all seriousness, being able to swerve all that terrible SEO content on reviews will always be welcome!
I don't think it's up to you, legally speaking: https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn
I mean, they could be nice and respect your robots.txt, but they certainly don't have to.
> fair use of snippets has relied on them being brief and linking to the source. Lawsuits will be immediate.
It's possible that fair use law will be expanded to cover this case, but as constructed the output of these models is generally fairly derivative of any specific original, and so probably protected under fair use. If it were spitting out exact copies of things it had read, it would probably be pretty easy to train that behavior out of it.
> I do love these imaginary scenarios where ChatGPT is going to find me the best air fryer, though. Where is that information going to come from, exactly? Barely anyone is making money writing reviews today, it's mostly farmed content. What happens when even those sites' reviews are quickly scraped and put into the next model iteration? Bing is going to have to come up with some kind of radical revenue sharing too if they want anything fresh.
I do agree with this, though. The LLMification of search is going to squeeze revenue for content creators of all kinds to literally nothing, at least if that content isn't paywalled. Which probably means that that's exactly where we're headed.
Google on the whole historically did not want to alienate publishers (and the advertisers that hang out on publisher content) and has avoided being in the content production business for this reason.
1) Everyone cryptographically signs their work for identity confirmation.
2) there exists a blockchain whose sole purpose is to allow content creators to establish copyright date on a digital piece of work.
3) a public that uses the two items above what evaluating the reputation of an artist.
I'd be super interested in a language model that was able to synthesize knowledge drawn from a large corpus of books and then cite relevant sections from various titles.
They will send at least some visitors, which is better than the zero visitors you will get from bing if you block it.
>- fair use of snippets has relied on them being brief and linking to the source. Lawsuits will be immediate.
Yes and microsoft has lawyers, who have presumably determined that the cost of fighting these frivolous lawsuits is not overwhelming.
You tell me! It's your site. If you want money maybe you should charge for your content? And honestly, the web that Google presents is just so terrible that I don't want to visit your site, unfortunately. And, maybe it's a price worth paying.
Google already does this with featured snippets
How can you refuse? The only way I know would be to require an account, but even then they could bypass it.
There are many air fryers on the market and the best one for you will depend on your needs and preferences. Some factors to consider when selecting an air fryer include size, price, features, and overall performance. Some popular air fryers to consider include the Philips Airfryer, the Ninja Foodi, and the Cosori Air Fryer. It might be helpful to read online reviews and compare the features of different models to find the one that works best for you.
For example, you could describe a probabilistic process to it and ask it what kind of distribution / process it is. Then, based on the extensive words you get back, you can continue your research on Google.
As such I think search engine integration is a really great idea, looking something like the follows
-> user: Hey searh engine, I have a thing that can happen with certain probability of success, and it runs repeatedly every 15 minutes. Could you tell me what kind of process this is and how to calculate the probability of 5 consecutive events in 24 hours?
-> engine: It sounds like you are describing a Bernoulli process. In a Bernoulli process, there are only two possible outcomes for each trial: success or failure. The probability of success is constant from trial to trial, and the trials are independent, meaning that the outcome of one trial does not affect the outcome of any other trial.
Here are some results on how to calculate probability of consecutive successes in a bernouli trial (result list follows)
(Note: if you try to ask this from ChatGPT it will not actually give you a correct answer for the calculation itself as there are some subtleties in the problem. But search results of "bernoulli process" will tend to contain very reliable information on the topic)
Edit: You could even just say "could you give me good search queries to use for the following problem" and use the results of that.
I know the results I'm going to get back are basically the same as if I went to Google, ran a query, it returns me their top 3-5 scraped "blog articles" based on relevancy, and then I ran it through one of those condensing/summarizing bots.
I'm not sure why it's as therapeutic as it is basically interacting with a search engine.
I wonder if this kind of technology will remain free for the foreseeable future. Google has to be coming up with something shortly, right? It's interactive search engine results "on steroids" (I think? I can't tell me if brain is tricking me to be biased that it's cooler/more useful than it is. Everybody I tell about it non-tech isn't that impressed/feels it's spammy/crufty/formulaic).
For me getting started is typically the most difficult part (thanks adhd) so this is a huge help.
I do agree with the overall point though. If you understand when to use it and when it's more likely to give you nonsensical answers, it can save a huge amount of time. But when I ask it about a topic that I don't know enough about to immediately verify the answer myself I'm forced to double check the answers for validity, which kind of defeats the purpose.
The best queries to ChatGPT are cases where I know what the answer should look like, I just forgot the syntax or some details. Bash scripts or Kubernetes manifests are examples here, I know them, I just keep forgetting the keywords because I only touch them every few weeks.
And don't get me started about asking ChatGPT about more general topics in e.g. economics or finance. What you get is a well-written summary of popular news and reddit opinions, which is dangerous if it's presented as "the truth" - The big mistake here is that the training procedure assumes that the amount of data correlates with correctness, which isn't true for many topics that involve politics or similar kinds of incentives where people and news spread what conveniently benefits them and gets clicks.
I’m not sure that fabricated nonsense would actually make Bing’s results any worse than they are today.
“It’s okay I don’t mind verifying all these answers myself” is an odd sort of sentiment, and also inevitably going to prove untrue in one sense or another.
I could imagine losing many hours from a ChatGPT answer. And if you have to go through the trouble to verify everything it says to make sure it's not just making crap up, then imo it loses much value as a tool.
It gives you answers with 100% confidence and believable explanations. But sometimes the answers are still completely wrong.
Add documentation to this method : [paste a method in any language]
For me the results have been impressive. It’s even more impressive if you are not English speaking because it explains what the code does but also translates your domain terms in your own language.
More than code generation I see a really concrete application in having autogenerated and up to date documentation of public methods. It could be generated directly in your code or only by your IDE to help you in absence of human written documentation.
Other interesting things it can does is basic code review by proposing a « better » code and explaining what and why it changed something.
It can also try to rewrite a given code in another language. I haven’t tried a lot of things due to the limitations in response size but for what I tested, it looks like it is able to convert the bulk of the work.
While I’m not really convinced by code generation itself (a la copilot) I truly think that GPT can be a really powerful tool for IDE editors if used cleverly, especially to add meaning to unclear, decade old codebases from which original contributors are long gone.
And knowing that what is hard is not writing but reading code, I see GPT to be a lot more useful here than helping writing 10 lines in a keystroke.
One pretty consistent way to see this is to ask for various very simple designs like a n-bit adder, it will almost always do something logically incorrect or syntactically incorrect with the carry in or carry out
For example recently I asked it for the best way to search in an mbox file on arch Linux. It proceeded to recommend a number of tools including mboxgrep. When I asked how to install it on arch it gave me a standard response using the package manager, but mboxgrep is not an arch package. It isn't even an aur package. It requires fetching the source and building it by yourself(if I remember correctly one has to use an older version of gcc too). None if it was mentioned by chatgpt.
This is not the first time BTW, there was another software it recommended that Debian doesn't know about, when I asked it another time.
I think a new implementation of ChatGPT is worth exploring though, one that cites sources and gives links to further information, and also one that has the ability to somehow validate it's responses for accuracy.
It could generate a python script that counted the days of the week with the letter T, but still insisted that Sunday had a T
Edit: Scratch that. I just tried again and now it says that saturday doesn't have a T in it
The function does not exist, it's entirely made up.
What's weirder is that when I told it that the answer was wrong it provided a different solution that was correct.
God forbid!
Google will probably build the same thing. Stackoverflow can suffer though...
Where do you think it got the information?
But the holy grail would be if it could write all my unit tests...
- Want a book summary quickly, without going through dozens of google search results? Try "Summarize Gilgamesh Epos" or "Summarize Carl Sagans Daemon Haunted World". => You can even ask a summary for each chapter of a book, if it has a record of that book.
- Want a quick overview about what DAO in crypto space means? "Explain DAO in the context of crypto"
- Want to get a quick overview of the old testament? "Give me an overview of the parts of the old testament, and describe in detail what the books are about"
- Want to compare Buddhism to Christianity? "Compare buddhism to christianity"
- Want a hypothetical conversation between Socrates and Napoleon, where Socrates tries to convice Napoleon that war is futile? "Create a fictional dialog between Socrates and Napoleon where Socrates tries to convice Napoleon that war is meaningless."
- Want a quick bio of Ataturk, in French?
- Want to have a quick overview about French Revolution, and how it effected the world? What other events were relevant during that time?
Even better: Want it to produce results in the Q&A style so that you can copy/paste it in your Ank app? "Create Questions and Answers about French Revolution so that I can use it in my Anki App. End each Question and Answer by triple X and a newline."
There are endless use-cases where I don't have the time to go ask google, collect the results, review and read them, to answer my question. This already replaces Google for me for these kind of quick Q&A sessions.
This is extremely exciting stuff, and I can't even imagine what we'll have in ten years from now.
An example of this is music:
ChatGPT responds to "What is the meaning of The Postal Service's The District Sleeps Alone Tonight?" With: "The meaning of "The District Sleeps Alone Tonight" by The Postal Service is open to interpretation, as the lyrics are somewhat vague and abstract. However, the song appears to be about the feeling of isolation and disconnection that can come from living in a big city. The lyrics reference the city as a "church" that "wraps [the narrator] in its twilight," and speaks of the planes flying overhead and their passengers "waving goodbye/ To all of their yesterdays." This could be seen as a metaphor for the transitory and fleeting nature of city life, with people constantly coming and going and not forming deep connections. The song may also be about the narrator's own sense of loneliness and longing for connection in the midst of the bustle and noise of the city."
There's two issues with this response: 1. The song is about the weird feeling of how places change with context, in this case it is about the context of a breakup and feeling out of place where you used to feel at home.
But there is a bit of a bigger issue.... 2. Those lines ARE NOT IN THE SONG. lol.
What this shows is clean readable responses that are oftentimes extrapolation of a thread or idea to levels not reasonably verifiable. Like how Facebook is/was tuned for giving more and more isolated extreme content, ChatGPT is tuned to follow threads and bits of information to extreme levels and confidence.
Then you can send that to another service; pushing math to wolfram alpha (etc) makes chatgpt suddenly give perfect math answers and it works (it takes a bunch of hacks to ignore it’s confident lying part of answering equations). When we it names a person or location, it shows Wikipedia or maps or both etc.
In this way you use it’s bluffing powers without assuming (or even reading) it’s content; it returns articles etc (which can be wrong as well of course but at least you might know if it’s reputable or not).
I can easily see how bing or google would use something like this to search and return very relevant content, the queries might not be better than some people can create themselves with effort. However many people cannot create good queries and chatgpt creates good queries from extremely little input.
More importantly; it’s a conversation; you can fine tune the results. You can tell it that you meant the programming language, not oxidation when it shows completely bonkers results.
This will work and will work very well as far as I have been able to try. I cannot go further as I have no easy access to search engines without getting blocked immediately for automated querying.
Not sure how this is looking cost wise: chatgpt at scale will even hurt a company like MS for now. Having a million techies use it for some experimentation is a bit different than a billion+ users searching whatever.
First type is info widgets which allow it access information outside of itself, like using a calculator by making it output something like {calc 1+1}, search wikipedia by {wiki something}, see the results of such a query and use that info in its output to the user.
Second type is UI widgets that allow it to display different types of things to the user. There is already a kind of widget in ChatGPT, which is display of Markdown. But it's easy to imagine widgets for say showing coordinates on a map, or visualizing something else.
Since there is a large number of such useful widgets, I'm guessing there will be an app store for them where developers can write their own widgets to extend it and make revenue based on how much they are used.
[0] https://blogs.microsoft.com/blog/2020/09/22/microsoft-teams-...
Remember, they are the ones that didn't publish GPT-2 because it was "too dangerous" for anyone else to have while numerous groups replicated it.
That is no longer the case. OpenAI should change their misleading name but it's worth remembering that their original mission statement has been fulfilled by OpenAI, academia, Chinese/European companies alike.
You can sign up for an OpenAI API key and build products with it. I'd be really surprised if they closed that down to only Microsoft.
What I picture happening here is social media going wild on day 1 with prompt injection fun, showing the Bing assistant as a racist paperclip maniac. I think we have seen this story before...
> Talk about the human colony that arrived on mars last month
>> The human colony that arrived on Mars last month is a significant milestone in the history of space exploration. It marks the first time that humans have established a permanent settlement on another planet.
>> The colony is made up of six people who were selected to make the journey to Mars after a rigorous selection process. These individuals were chosen for their skills and expertise in areas such as engineering, science, and medicine, as well as their physical and mental fitness to withstand the challenges of living on another planet.
The technology is cool, but until we can trust the answers it provides, it's just a fun toy. Nobody is seriously asking their Alexa today how to perform open heart surgery and accepting the response as gospel, but that's kinda where we are at with ChatGPT! The confidence it exudes is incompatible with learning, unless you like to be taught complete bullshit 50% of the time and then carry on with your life none the wiser.
Being "confidently wrong" to me would mean just giving me broken code, but the facts it solves the problem by itself and then acts like the bad code it gave me did it... That's really actively deceptive.
It works by getting Google results, finding the most relevant text chunks in the pages, then passing them to GPT-3 to generate a summary. The summary includes sources, so you can verify the info.
It works well for research questions like "what is the best smartphone with a small screen?", or "how do large language models work?". It gives better answers than ChatGPT alone for these types of questions.
I'm finding that the biggest time savings is from not having to wade through SEO-spam pages to find the one relevant paragraph.
Imagine a world where every child alive had access to unlimited 24/7 pseudo tutoring offered by robots who were almost as smart/correct as actual tutors? How many millions students just "give up" every day because while they can find thousands of pages of facts/guides/lessons about a subject, there are zero results that are custom tailored to their exact gap in understanding? As a student,
-->Google can give you endless facts about a subject that you can ram your head against until a subject makes sense.
-->Chat-GPT can give you intuitions and explain the exact part that confused you with as much detail/complexity as you want!
You wouldn't even need to call it AI or a tutor (so students don't wrongly give 100% trust). It could be a peer. You could call it "My Study Group's Best Guess". How much further could students get in high school with a 24/7 tutor beside them?
~~"Hi study group, I don't understand how my teacher jumps between steps 3 and 4 for this proof in Geometry class. Can you explain it in simple terms?"
~~"Hi study group, why did my flask turn blue when I did xyz in chemistry class? Everyone else had theirs turn pink."
~~"Hi study group, I wrote "cosin - sin = co" on my math test but my teacher put a big x on my problem without telling me why."
~~"Hi study group, why is the x "number of cars" on this homework problem XYZ"
Millions of students give up trying to understand challenging subjects every day. LLMs can dramatically lower the difficulty of absorbing new ideas. In 10 years, LLMs can likely raise SAT/ACT scores 20-30% equivalent and help more students finish the key math/science prerequisites for STEM careers in high school instead of college, lowering the barrier of entry for everybody!
We have to think broader and get out of the confines of how things work now and just think that technology only makes what we have now better. Some technology just completely changes the landscape.
LLMs will make education as we know it obsolete. There won't be need for tutoring if LLM is good enough to do anything. Anyone can ask anything on the fly. There won't be need to study anything except for fun.
A pathetic scenario but somehow consistent with the rules (or lack thereof) of the game
Search engines will have to rely more on signals outside the content, such as links from other authoritative sources, but it does not look like a qualitatively different world.
I can imagine good enough AI being able to spot truth even better than what humans do - by veryfing sites and commenters with sources of real information to estimate their credibility.
E.g. in a theme similar to Page Rank, you could have an AI that has some sites as a source of objective truth (Wikipedia, science journals, reputable sources of news etc), and then use that as a basis of estimating trustworthiness of a material.
Also, AI could find, for a given subject, opposing opinions, and estimate which ones are possibly fake, and which ones are real.
In essence - do what current fact-checkers do, but for every single website and comment in existence.
It is a book everyone of us should read so we may have an active role in choosing an expansion rather than contraction of wealth.
The disruption these AI systems will cause along with increasing concentrations of wealth away from the middle classes is worrying, but as humans we can engineer a humanist outcome.
Sadly I think greed will win, as is human.
Also the moral system ChatGPT has actually is the worst, most false-positive-laden firewall I've ever seen, period. It's actually impractical to use ChatGPT just because of how much that gets in the way. It obviously shouldn't even have that (and it's trivial to bypass), but companies gotta look good. It's actually funny because ChatGPT is like an American ideologue who gets triggered the very moment it conflates something you said as something against its ideology.
I now believe the future of web search is curated lists of non-commercial sites with a text index over them.
Personally, 95% of my Google searches are looking for an answer, not a specific document. If there were an all knowing AI that could answer all my questions accurately I'd almost never use classical document search.
You get a lot more bang for the buck by splitting up queries into very specific short ones.
Could you give some examples?
And it works for generating text that would be tedious to do manually.
But for searching information, I can’t bring myself to trust it.
Don't ignore rate of change, recognise it can only improve with time, you're looking at very early system that will quickly be orders of magnitude better.
With chatGPT I wrote a few sentences about our warehouse and asked it how to optimize some part of our process. Then it spat out 5 suggestions that were quite tailored. The same as I found when googling, but instantly in the correct context without me having to skim loads of google results and try various queries to avoid spam.
"How can I code a mock for a boto3 dynamodb client, using unittest.mock, using a patching context manager"
And it will give me a highly relevant example, often right or close enough.
I will occasionally ask to come up with a function based on requirements but I don't find that nearly as useful as using it as super advanced search engine.
If the endpoint is not the user clicking in a link then why should these companies give away all that value?
I live from the content I write. It's not fluff. Some of it comes from weeks-long email conversations with government officials. It takes a lot of research and help from experts I have long-standing relationships with.
If search engines serve that information but deny me the traffic, the website dies, as does the source of the information.
I can deal with lazy copywriters just rephrasing my work because the original still outranks them, and I have legal options to deal with them.
I can't do anything if Google - over 80% of of my traffic - decides to proxy my content and starve me of my income.
Stackoverflow has value for these models, the others will just make the nonsense responses worse.
Imagine a search engine that summarises the world’s knowledge in comprehensible terms. That doesn’t just return a ranked list of websites, but answers that are synthesised from a multitude of sources. Google has technically made human knowledge discoverable. But it has never helped mere mortals all that much to navigate the maze that is academic publishing, reputable journalism, SEO SPAM and outright manipulation. By contrast, imagine how ChatGPT in Bing can not only improve productivity, but for the first time give everyone access to the world’s knowledge. What I love about ChatGPT is that you can ask it to dumb down the answer it has previously given. And to summarise complex topics.
For the first time since Google killed the likes of Altavista, their moat has become leaky. About time! We need true competition in search and hopefully Microsoft and others will also innovate on the business model. Heaven knows I would pay for a search engine that doesn’t spy on me and gives me better results. Not to mention that this will also kill the perverse incentive for websites to do SEO. Hopefully, the web will become readable again, as a result.
You’re far more confident than I am that a language model is going to solve this. I’m concerned that all it will do is provide an air of authenticity to these wrong answers.
If a colleague tells me once "I'm certain it works like this" and it turns out to be wrong, probably I won't trust them again.
I already cannot trust ChatGPT, I have seen it bullshitting big time.
It’s amazing watching (a possible) paradigm shifts occur in plain sight, and how history (may) repeat itself:
Google is not willing to add this feature to their main search product, and looks like they wouldn’t make a smaller “toy” search engine, that can fail. They are too afraid. It’s the innovator dilemma at play.
For example, I asked chatGPT what percentage of electricity in Germany was produced by renewables, and it said 85.6% and quoted the federal environmental agency. When I went to the FEA website, it clearly said that only 41% are generated by renewables.
Still I would imagine that you could build a dedicated postprocessor that fact checks the text and another that fixes any math errors.
prompt: Please give some recent citations for scientific work examining the nature of consciousness
ChatGPT:
> Here are some recent papers that explore the nature of consciousness:
> 1. "The neural basis of human consciousness: a focus on the thalamic reticular nucleus" by Christopher D. Schmitt, Leonid L. Moroz, and Laurence J. C. Mailly (2019)
> 2. "Consciousness and the brain: an overview of the experimental data" by Anil K. Seth (2018)
> 3. "Towards a theory of consciousness" by Bernard J. Baars (2017)
> 4. "The hard problem of consciousness: a problem for physics?" by David J. Chalmers (2016)
[As far as I can tell (or actually, as far as google search can tell) none of these papers exist.]
> I hope these papers are helpful! Let me know if you have any other questions.
Hmm.
However, your point is still a good one. If Microsoft places ChatGPT somewhere alongside Bing, it'd be easy to mistake it for fact.
In my opinion, this new tool is going to be learned by the public just like Google and generic search was learned by the public. There's a certain way to prompt the machine and we all have certain expectations of what the machine should return.
With ChatGPT, those input/output expectations are vastly different than web search, but people will learn it.
I know we all dislike ChatGPT output posted here on HN, but since this is about ChatGPT I hope I'm forgiven.
Currently ChatGPT Q/A:
> Q: What are the arguments, backed up with facts, for Microsoft being a bad company? Are the Halloween documents relevant?
A:
There are several arguments that can be made for Microsoft being a bad company:
1. Monopolistic practices: Microsoft has been accused of using its dominant market position to stifle competition and maintain its dominance in the market. In the 1990s, the company was found guilty of violating antitrust laws by using its monopoly power in the operating system market to gain an unfair advantage in other markets.
2. Poor treatment of employees: Microsoft has been criticized for its treatment of employees, including low pay and long hours. In 2014, Microsoft was ranked as one of the worst companies to work for in the tech industry by a Glassdoor survey.
3. Poor working conditions in factories: Microsoft has been criticized for its use of sweatshops and poor working conditions in its overseas factories. In 2010, the company was ranked as one of the worst companies in terms of labor practices by the International Labor Rights Forum.
4. Invasion of privacy: Microsoft has been accused of invading users' privacy through its collection of personal data and use of tracking technologies. In 2013, the company was fined by the European Union for failing to adequately protect users' privacy.
4. Misleading marketing practices: Microsoft has been criticized for its misleading marketing practices, including making false or exaggerated claims about its products. In 2018, the company was fined by the Australian Competition and Consumer Commission for making false or misleading representations about its Surface tablets.
As for the Halloween documents, these were internal memos that were leaked in 1998 and revealed Microsoft's aggressive strategies to maintain its dominance in the software market. These documents are relevant as they provide evidence of the company's monopolistic practices and show that these practices were a conscious strategy rather than an unintentional consequence of success.
This isn't true [1] and I don't think the 2013 privacy fine is either. They were fined for not complying with browser choice requirements in 2013, but not for privacy violations afaik. This would be an entertaining libel suit if presented as fact.
[1] https://laborrights.org/publications/2010-sweatshop-hall-sha... (no mention of Microsoft)
It's like saying the worst thing about Al Capone was that he made a mistake on his tax forms.
Billy Joel's best known song is probably "Piano Man." Here are the lyrics:
Verse 1:
It's nine o'clock on a Monday
The regular crowd shuffles in
There's an old man sitting next to me
Making love to his tonic and gin
He says, "Son, can you play me a melody?
I'm not really sure how it goes
But it's loud and it's funky
And it's rhythm is right on time"
Chorus:
Oh, the piano has been drinking, not me
Verse 2:
He's talking to me, how he's got no time
For his only son, who's just turned twenty-one
And he says, "The bill's on me, son, just be my guest"
And I think it's the best of the three
Chorus:
Oh, the piano has been drinking, not me
Verse 3:
It's five o'clock on a Friday
I'm turning off the telephones
I'm getting ready for the weekend
Bridge:
I'm painting my room in a psychedelic style
But my friends say it's just a waste of time
"Cause all the madmen, they seem to have gone mad
And all the best guys, they're already gone
Chorus:
Oh, the piano has been drinking, not me
Outro:
Yes, they're sharing a drink they call loneliness
But it's better than drinking alone
I hope that helps! Do you have any other questions about this song or anything else?ChatGPT (or a similar product) focuses on solving users problem interactively. No ads, no going to another website etc. How would you make money from a search engine?
I was hoping a simple, paid model to start with. Over time, as the LLMs become commodity (200GB, runnable on Intel/AMD), ship it as part of the OS and other devices.
I think there is no technical barrier to add Ads to a chat bot. It's even more deceptive when promotion campaigns are embbeded into text contexts implicitly. It would be much more dangerous and harder to block than, let's say a DIV on a webpage.
Microsoft makes money from services so it doesn't have to make money from a search engine.
Google, on the other hand, makes money from ads so it doesn't have to make money from services.
If Microsoft can take Google's ad revenue away they no longer have to compete against Google's free services.
The converse also holds but it may be easier for Microsoft to weaken Google.
As your conversation evolves with the bot, targeted ads could be shown with the same (or better) level of intent data available based on the human’s input.
<chatgpt summary of cards>
Links to cards mentioned in the summary.
… now you can make money on the links directly, perhaps have sponsorships influence the recommendations, and have a strong signal of intent to purchase you can try to monetize later.
“Hey bing what’s a good hotel in downtown Montreal?” -> same
I’m joking of course but somebody is going to this
Granted they haven't been perfect but many times what I find on stack overflow isn't perfect either.
This week though after a back and forth with ChatGPT I was able to solve a pretty complex issue with some pixi.js code after no relevant help from Google. Likely saving hours of work.
Sorry, I was joking
I think the real challenge is internationalization - it would be a challenge to build GPT-N like models for all the other languages, that work as well as the original one.
Interesting if Google will roll it's own language model for that purpose. Is it possible, that we might get several language models, each one for a specific category of users, or would that approach lead to a loss in generality/quality?
Interestingly, it has no problem with throwing in words from my language when I forget english ones.
The only two issues I found: - doesn't do good rhymes when I ask it to write poetry - when I asked it to generate content that has mixed polish and english words (as if written by a pole who spent the last 20 years in US, and replaces some words with their english counterparts), it was unable to do so. It could only write either clear english or clear polish.
As long as I ALWAYS know when a GPT derived match is shown to me.
Preferably with a diagnostic akin to sql EXPLAIN.
Useful to be able to flag incorrect info although this is open to trivial subversion.
Accuracy issues will improve. Mostly likely the improvement necessary for LLMs to answer a vast majority of queries as reliably as necessary will happen quickly. Not for all cases. But for a sufficiently large set of cases to make Google alternatives very compelling.
And content rights issues won't stop the march of chat interfaces anymore than they stopped the march of streaming services. What will happen is the evolution of new dynamics and rules in the means by which information is produced and information producers are incentivized and compensated.
The key with respect to content is that LLMs require source material to be produced in the first place. And, for the foreseeable future, AI will still rely on humans to produce that source content. Human labor forces will need to evolve and reorganize to optimize the production of original information in that setting.
On what basis do you say this? Is it anything more than your gut feeling?
How much additional training data will be needed in order to achieve these improvements and do you have a concrete reason for thinking that amount will be sufficient? Do you have any expectations as to how this training data will be obtained?
2023 is going to be an interesting year.
PaLM (540B parameters): https://ai.googleblog.com/2022/04/pathways-language-model-pa...
Med-PaLM: https://archive.is/oC5Gj (https://twitter.com/vivnat/status/1607609299894947841)
Google very likely has the absolute best AI on Earth, and might even be a few steps ahead of Open.AI. However they are extremely coy about it and so far only use AI (public facing) to lightly augment their services rather than be the services. For instance, Google purposely makes assistant act more like a computer taking commands than a human having a conversation.
But we know that Google has at least parity with Open.AI, and it would be a fairly safe bet that they are ahead. We'll see if they have a "mic drop" moment when Bing comes out with this.
I think there's proven to constantly be a huge market for companies that do the dirty work that everyone knows they could do, but don't have the time or expertise to do.
If they can stay ahead of the curve on having best in class api-ification of the underlying tech, even if they develop competitors with similar quality, they have the momentum of being the brand name now.
I'd love to be able to query "what's the general consensus about XYZ, with lists of pros and cons. Exclude obvious affiliate sites, and sites that sell this thing."
For example, "write me an html file that contains a simple Preact app that calculates a tip" will get you back something that looks pretty correct, but it will just have tons of small basic misunderstandings like trying to use jsx syntax, or not actually importing things from the library even though it does include a script tag, etc.
It IS good at something like "give me ten rows of fake data that match this interface" though. That is useful.
If I made my blog GPL licensed would this be a decent way of "protecting" myself? I write programming articles for the sake of it, I'm not trying to turn them into a book to sell just basically collecting my thoughts. I don't even run analytics on it, basically no will to sell this info at all.
Would such an action be a good way to prevent companies from monetizing my writing without complying with GPL?
Whether they are right that this is fair use is above my pay grade, but if they're right, then GPL won't help at all. You can't stop fair use of your copyrighted works; that's not a technicality, it's the whole point of fair use doctrine. The best you can do is try to show that a particular usage isn't covered by fair use.
Setting up robots.txt is probably your best choice. In the EU commercial text and data mining must respect a machine-readable opt-out. In the US they could legally ignore it, but I reckon they'll follow it anyway to stay on the safe side and avoid having their crawler blocked for misbehaving.
The lawsuit website made a showing here some 60 days ago: https://news.ycombinator.com/item?id=33457063
I think probably not.
There aren't really laws/cases that apply to LMs creating derivatives yet. Maybe there will be in the future. Currently, in practice, GPT will digest your blog and incorporate it into itself.
I'm very interested in the UX decisions they make here. Is it just going to be ChatGPT with a Bing logo or will it be able to intelligently decide when a search engine is better? Will it give results in natural language?
If they sometimes do normal search instead at least that answers how they'll make money
The problem I see is the web being taken over by bots phishing, catfishing, trolling, stalking, doxing, etc. humans. For example, if I "offend" a bot or its PRNG decides, will it be able to construct a strategy such as deep fake you or your friend in deep fake porn and make it appear you were the culprit? We're a bit far off from general AI being let loose, but we're getting closer to losing control and the feedback cycle of tech amplifying itself into the technological singularity.
Surely this is Microsoft's end game after becoming a "preferred partner" to OpenAI?
Many of us has seen how it kills Google or at least vastly outperforms it in many respects, and Microsoft of course knew this all along, including when they started seeing the potential and investing in it? We're just reacting late to what Microsoft saw many moons ago. And we aren't even seeing the finalized product that will surely be based on GPT-4 rather than GPT-3?
I honestly wonder how far back this stretches. Nadella has steered this ship so well compared to Ballmer. His focus on Azure and later AI aligns so well with where computing at large is moving, and meanwhile Google seems stuck.
Also OpenAI already published an interesting paper a year ago about their WebGPT model which imitates human web browsing to collect references to improve factual accuracy:
How come HN doesn't have nearly the number of bots or lame AI commenters as Reddit or other places? Is it an effect of the voting system? The user filtering and security steps in place?
Or there's just no money in it?
However, I do think this means Google finally has real competition in search for the first time in it's existence which is potentially disruptive enough for them, even if that means losing a few basis points of market share. It's possible this might force Google to start charging for their other services if competition for search threatens their ad revenue.
Search has been compromised so much by SEO games that the thing that ranks is essentially the king of the spam hill. By inserting the output of the apex spam machine into the SERP Microsoft is saving themselves the headache of scraping all that content and figuring out who has the best spam themselves.
Might as well make it yourself. They already know, more or less, what should do the best. So why not supply your own spam to a market you utterly control? The outcome is the same. Naturally this is depressing, but it makes sense.
The only real alternative is human-curated aggregations with trust networks.
Currently there is a lot of power with the big corporations to put their angle/slant/bias on search results. But in general, as a user you still are presented with pages of results you can scan through and read.
With something like chatGPT, what is the likely hood result sets will be home more specific and echo chambered. And furthermore, what is the chance the big corporations will get even more control on applying their bias/slant/angle than we currently already get.
And, as a thought experiment, what might the answer be in 5 years and 10 years.
AI badly needs better explainability to encourage trust and as a means to better accuracy.
As others here have said, many of us have spent years curating online information (blogs, books shared online, general web sites). We have done this work with some expectation that we get benefits: people identify us with topics, we get ‘known and hired’, we make money off of our content, we make new friends, etc. ChatGPT obviously has no attribution in its current form.
Just like how Dall-E can paint a human face pretty well but then mess up on their fingers and other extremities, ChatGPT can give you some really logical-sounding statements but totally bomb the actual facts of things.
Am curious how they'll execute - I def see a lot of potential usefulness.
It’s simply a better UI for users.
I assume that bloomberg was just confused and incorrectly stated that Microsoft is directly using ChatGPT, when really Microsoft is using one of the language models which OpenAI offers as a service...?
If a language model is truly a better way to do search than current approaches then that alone is hugely disruptive... Like multiple industry ending kind of disruptive, imo.
ChatGPT in search could be the first paradigm shift in 20 years, so someone had to do it. I'm grateful it's Microsoft, since at least it's a company with liability and everyone will be ready to jump on them the moment this new era of search gets out of line, which it likely will.
Only automated and scaled.
Pretty please!
How to change start-site to google.com in edge?
Chatgpt gives us answers, but so do the top linked articles in a search engine. Why bother adjusting to something different if we already have something that solves the problem?
For most people, the marginal amount of time saved by using chatgpt instead of a search engine is borderline meaningless. It might be useful for lawyers to quickly tl;dr dry legal documents, but general use seems far fetched to me
1) My first thought would be that such a conversational search UI would never hit the original website, thus drying up monetization for the content creator. Kind of like snippets in current search. Then I realized that there is no "original website", no singular source for the AI answer given. That's even worse. Search is the main traffic controller of the web and now has zero reasons to send any traffic your way. Not only does it not have any reason to send traffic, it simply can't. You can't back trace an AI answer to a (single) source.
2) The implication of the above point may be that for human creators, producing new content for the open web becomes (somewhat) pointless, even in the case where you do not even care to monetize it. An exception is video, for now. You should not expect this change to lead to a drying up of content for AI though. Google/Microsoft will simply resort to making deals with major content platforms. That's where most content is produced anyway.
3) An interesting new problem for Google/Microsoft is the increased accountability regarding correctness. Right now, search results do not have to be correct. You're just showing results produced by other entities. This changes when you confidentially give an answer directly in the search UI though. Now you own the answer, even if this expectation is technically inaccurate. Since GPT3 is often confidentially incorrect, that's a major challenge. And we should hope that this problem remains for a long time, because the moment GPT3 is extremely accurate and correct, we're in a different world entirely.
4) Safety and political correctness is another concern. Google/Microsoft will likely deliver AI that is a little too safe instead of not safe enough. The implication is that instead of being a controller of traffic, it becomes a controller of opinion. Of right think and wrong think. Which may be ever changing and subject to institutional pressure. Combined with point 3, you're no longer a proxy. You're the source of truth and the source of what's right.
5) Some say that the open web is already close to dead thus the already ongoing trend of people resorting into private spaces (Discord, Mastodon, etc) might accelerate. An open web and a closed human web. Even in the closed human web AI-generated content can and will penetrate so there will be attempts to detect it. Which is likely futile, because one can use AI-generated content and then tweak it to make it more human-like.
6) We need legislation ASAP. None of our regulations and laws are designed for what is coming. Do you really want a congressional hearing 10 years from now where a 75-year old asks "So this AI thing, is it like a smart calculator?", after all damage has already been done?
7) We should unite in at least some basic pushback. Stop thinking in individual use cases like AI art, Copilot and GPT3. Think about the broader problem where all of human creation is sucked up for free and without permission, after which all benefits are centralized up to the point of even replacing you. It's a shared, human problem. We're unlikely to stop it but at least don't make it this easy.
I've been playing with you.com's GPT web search (YouChat). It's an excellent bullshitter. It always gives you an answer that sound plausible, but whenever you're asking for something beyond what plain google could do, the answers are wrong. Some examples (make sure you check the math -- a lot of 1+1=5 going on here):
Q: Which city has the highest ratio of houses to gas stations in Texas?
A: The city of Austin has the highest ratio of houses to gas stations in Texas. According to data from the U.S. Census Bureau, Austin has an estimate 1,340,742 housing units and only 748 gas stations, giving it a ratio of 1.79 housing units per gas station. This is much higher than the state average, which is estimated to be 0.47 housing units per gas station.
Q: What led to the bankruptcy of Facebook?
A: The bankruptcy of Facebook was caused by a number of factors, including an inability to effectively monetize its platform, competition from other social media companies, and a lack of innovation. Additionally, Facebook had a number of legal issues, including a lawsuit over its alleged misuse of user data. All of these factors, combined with a declining stock price, ultimately led to Facebook filing for bankruptcy in 2020.
Q: Why did Tom Hanks eat an apple?
A: Tom Hanks did not eat an apple. This is a false accusation and it is not true. Tom Hanks has never been accused of eating an apple.
Some other examples from an old article:
https://searchengineland.com/googles-one-true-answer-problem...
Teaching people not to trust what they read on the internet seems like a good idea to me.
> what led to the bankruptcy of jp morgan chase?
> The bankruptcy of JP Morgan Chase was likely due to a combination of factors, including a drop in liquidity, a decrease in market confidence, and the rising cost of debt. Additionally, JP Morgan Chase was heavily exposed to the risky subprime mortgage sector, and its increased risk-taking and lack of transparency about its financial activities likely contributed to its eventual downfall. Additionally, JP Morgan Chase also faced numerous legal issues, including investigations by the U.S. Department of Justice and the U.S. Securities and Exchange Commission, which may have contributed to its bankruptcy.
edit: it might be a special case with banks due to the difficulties in the late 2000's. It does get companies e.g. CNN or Fox correct.
There are so many jobs, roles, and tasks that require this very well honed skill. Given the number of semi-plausable BS I have to sift through daily from humans, increasing adoption of systems like GPT should make things... interesting.
I bake my own bread for my son due to certain food allergies. We have some leftover almond flour, so I googled up some bread recipes using almond flour this morning. Web searching in 2023 is almost impossible for food now. My results were the shovelware of autogenerated content farm recipes.
One example I found was a quickbread (In my country that's a baking powder based cake, sometimes with yeast as a flavoring rather than leavening agent, being marketed as a bread). This is a legit looking recipe and would probably work as an almond-flour-flavored quickbread. However, it was provided with ChatGPT flavor text surrounding it. Statistically, most bread recipes have instructions and answer questions about yeast and yeast problems. Therefore ChatGPT decided this recipe also needs an extensive wrapper of yeast-issue flavor text. Of course any actual sentient being that's not the output of mere statistical randomness would have enough cognitive function to instantly identify this recipe as a yeastless quickbread and would not include the flavor text about yeast issues. In a similar manner, ChatGPT "knows" that bread recipes always start with preheating the oven. However, this clickbait recipe was ostensibly for a bread machine. So the recipe and flavor text was a bizarre randomized non-cognitive mixture of oven vs bread machine instructions. "Preheat your bread machine to 375" and so forth.
I've had similar experiences with ChatGPT generated "technical" content. ChatGPT is great at imitating hot take formulaic blogpost clickbait type content, but with zero cognitive ability backing it up the content tends toward meaningless word salad. I think as per the Sokal Affair from the last century, some people get VERY mad when its pointed out their favorite, most emotionally triggering content sources have all the cognitive intellectual basis of a random number generator.
The sad part, is literacy is so weak in most countries that ChatGPT can "outwrite" a large segment of the mostly illiterate population.
I find it pretty easy to identify autogenerated filler content, both online and social media, because there tends to be zero cognitive function behind the well written article. Its like the article has a three dozen well written and stylish lines from Pulitzer Prize winning articles, but because they were selected based on randomness plus being found nearby each other, the higher level cognitive functioning is completely absent, which is jarring when juxtaposed with well written individual sentences and phrases.
ChatGPT is, essentially, a pollution source, and should be regulated and limited as such. I never would have guessed what kills the internet is "Super Eliza" flooding out real content with infinite random filler garbage content.
Now ChatGPT is 1 further step of intermediation. No longer are you competing with everyone else to get noticed by Google search, hoping Google actually has an incentive to bring people to your valuable page. No, now ChatGPT will just serve your stuff to the customer, having stolen your stuff wholesale with no chance of ever actually referencing you. Given that new situation why would anyone ever produce any text of value for the internet? Why? To pump MSFT share price? Well you won't. So what will be produced? Massive amounts (even more) of trash information will be pumped out, raising the noise floor and in turn bringing down the quality of ChatGPT, and who's to say what's right! The whole point of Google was to figure out what was right from the organic structure of the internet, but if you replace the organic structure with Google and ChatGPT there's no one left providing the original information that these services will use.
My proposed solution: People sign their articles with a key, and I specify for which topics I trust certain people.
I trust Paul Graham on the topic 'startups', so anything that he writes on the topic is considered a high pagerank for me. Paul can also trust other people on the 'startup' topic, and so they will automatically be ranked high on my pagerank (for that topic!). When I do a search, it takes into account the people that I trust on certain topics.
Please steal my idea and make millions!!!
Ultimately, human progress happens when they save time.(e.g invention of wheels, cars, airplane, telephone, messaging, internet and now AI)
Eventually no one can prevent AI progress. And any wise person shouldn’t. It is the progress of human specie. And If no one can prevent it, let’s embrace it with open arms and make it happen as quick as possible and advance the world further to solve even bigger problems using these advancements in tech.
I discovered the ecobee thermostat in the comment section of a post about home automation yesterday. It’s nice to know (based on my assumed wisdom of HNers on these subjects) that it’s a solid product. That kind of thing has gotten to be next to impossible to produce organically with Google.
>I really worry about the meta effects that Search is having on the internet. Google has slowly degraded as the internet has morphed into being almost entirely an SEO driven contraption.
Do people at Google know and acknowledge this? I mean that in a non-Upton Sinclair kind of way. "Search" for me doesn't really take place on google anymore, but myself and anyone else here who is in the same boat is an outlier. So has the idea at Google become that "search" has just evolved to giving "the masses" (i.e. the non-technical people) SEO garbage?
The problem is that Google is not incentivised to provide you with good results for multiple reasons:
* good results means not only you spend less time on the search results page, but you are less likely to click on an ad. Spam results make the ads more attractive in comparison and may solicit more clicks.
* spam sites often contain Google Ads and/or Analytics, so spam results also indirectly contribute to Google's revenue.
* websites (whether spam or not) are relatively more likely to pay Google (for ads, analytics, GTM, etc) than end-users, so even a 1% false positive rate on the spam detection would be bad for business, where as feeding spam search results to end-users doesn't affect business at all as long as the search engine is free and there's no good competition
Spam results are trivial to detect should there be an incentive to do so: use the presence of ads, analytics or affiliate links as a negative ranking signal. Spam sites can't easily conceal those either because the same techniques they'd use are also used for ad fraud and would be forbidden by their advertisers/affiliates/etc.
I gave up.
I actually started informing myself about Magisk via ChatGPT and that was very helpful. But if ChatGPT is going to start learning off of these blogs, as if it were using Google to teach itself, then we'll have an even harder time in the future, as ChatGPT will wrap that possibly wrong or even dangerous content in trustworthy sentences.
Amazon now is in a good position to grow their books catalogue a lot. They do protect the copyright of their authors religiously. You can't publish books on Amazon with text that is already on the Internet. Neither can you publish text that is already available in another book on Amazon.
What could cause a major problem for search engines is the balkanization of the internet from Web 1.0 personal sites to Web 2.0 shared (but still indexable) platforms to mostly unindexable private groups. Discord, and Whatsapp are models for how this might happen.
What might cause an acceleration in that direction? Bots. ChatGPT can create content that is authentic-looking, and if open forums become Markets for Lemons [1] then people will move to trusted groups like smaller chat groups. (The other alternative is verifying identities which will happen but has privacy issues.)
[1] https://www.fortressofdoors.com/ai-markets-for-lemons-and-th...
Google really has it coming. They bred the SEO industry that has essentially turned websites into SPAM. There isn’t a single website that doesn’t have you scroll through lengthy SEO-optimized bullish.. content.
As for ChatGPT “stealing” content: I’d rather take the view of aggregating human knowledge as a positive. Imagine ChatGPT aggregating, summarising, weighing and dumbing down academic content. And if as a result all that knowledge became accessible by a broad audience. What could this do for the average person’s scientific literacy? Or fake news for that matter?
Language models on top of search are an opportunity to filter out the same junk and spam you complain about. I think in a few years the quality of generated content will surpass most of what humans create today.
And we will have verification models that check against datasets of facts, do web searches, follow links and check for consistency. They can also extract licensing information so we can avoid copyright blunders.
Training a language model to solve tasks with a search engine, especially with current day Google, with all its issues, would be an interesting project. Just define thousands of tasks, solve them with humans, then let the model learn how to use search and filter the junk in order to produce the same result.
The LM will be calibrated by training on real Google to solve real tasks. It is a reinforcement learning task to invert Google's disease. Problem solving will calibrate the model.
This is the world before ChatGPT. Eventually they will have to update it. How do you separate pristine Information from ChatGPT Information during training? If not addressed properly this will raise the noise level even further.
In fact, one of Google's advantages of altavista and such was that pagerank diluted the effects of keyword stuffing, the primitive SEO that had degraded primitive websearch. Search, media and advertising is dynamic, cat and mouse stuff... always. I wouldn't be surprised if "Language Model Optimization (TM)" is currently being invented by a drop shipper living under his basement.
Measure something, act on it, and you change the thing. You have to remember that these things have symmetry. The seo is trying to help users get to his site. From their perspective, Google's anti-spam is hostile to that goal. Google is trying to serve search results. To them, SEO is a hostile attempt to prevent that.
Before Google the web was a lot less searchable. You had to go to Yahoo, Altavista, or similar websites and wade through pages and pages of nonsense before accidentally stumbling on a good link. Very hard to find stuff back then. The effect search has had on the internet is that we can now actually find stuff pretty easily.
Of course the new thing here is that we are now mixing search and content generation. Before you had to manually create content, put it somewhere and then hope people could find their way to that content. Now people ask a question and the AI generates the content that contains the answer. AI generated content has the potential to vastly outnumber any human generated content.
Trash content, is what we have already. E.g. reddit is full of it. Lots of people are addicted to spreading their nonsense there and trying to be as offensive as they can possibly be. There's also some good content there of course. With AI there will just be more of it.
This is an issue both for training and for users. A model trained on nonsense will produce nonsensical content. Not good. So there's an incentive for separating the low value garbage from the high value content. Another issue is that you get feedback loops once we start training models on AI generated content.
In other words, there's a need to tell what is what for effective training of models as well. A good strategy is to simply figure out reputability and nature of sources of content. You could use an AI for that. Digital signatures would help too. Trusting certain websites more than others is helpful. Some websites helpfully score and moderate their content or have people liking things to indicate it is important. So, plenty of signals to work with here. That's why Google is still pretty effective. And that's why chatgpt is mostly pretty impressive and not rehashing a lot of alternate reality nonsense from some substack/reddit. But it's definitely an arms race.
The next level in this is going to be people getting more protective of their content. IMHO it's a problem that most content on the web is not signed by their authors. We have no way of doing that in a sane way currently. We have SSL certificates for domains but there's no such thing that guarantees a blob of text comes from a specific person or entity (an AI model could be signing it's generated content as well). As soon as you have this, you can start figuring out reputability and relevance to a particular question much more effectively.
People are rushing around trying to find practical uses for ChatGPT. I think this time they'll finally start finding them, probably a lot of them. I think this is one of those dam bursting points. The difference between GPT3 and ChatGPT, to me, suggests a lot.
Writing code, writing documents, reading them... search... lots of things are about to change fast. The conversation leaked by this microsoft mole is happening in a lot of places. If you can't think of things worth trying to build with GPT... you need to take a bath and try again.^ Bonus points if you can get GPT to code it for you.
There are, absolutely, spectres hanging over this stuff. Monopoly, centralisation, nontrasparency is expected. Career anxiety. Skynet is not out of the question. Lots of reasons for trepidation. That said, ATM, I'm mostly excited. This technology is amazing, and it is about to spill out into the world.
^I want a reference book that works as a cascade. A table of contents (or title even) at the top end. Click in for a chapter abstract, then summary, etc. Concise at the top, detailed at the bottom.
meta question but is it fair to report someone else’s scoop but bury that fact in the last sentence 4 paragraphs into the article? as a result of this burying Bloomberg gets to the top of HN rather than The Information which got the actual story. If I were that reporter i’d be pissed.
link to real source https://archive.ph/o/1ChFk/https://www.theinformation.com/ar...
"The cheapest flights from Prague can be found to the United States, with a one-way ticket costing $209 and a round-trip ticket costing $368[1]. The most popular route is from Prague to Los Angeles, with the cheapest round-trip airline ticket found on this route in the last 72 hours being $315[1]."
Completely ignoring reality of Wizzair, Ryanair, Easyjet, Eurowings, where you can buy one way tickets for like 30 dollars, so while it may be somehow usable, it's not really reliable.
Hey Microsoft, how about instead of adding chatgpt, filter the results so Bing actually works as a search engine? That is, select only the search results that contain the search terms, rather than returning the whole internet in an order somewhat influenced by the search terms.
Once search engines work like ChatGPT, the incentive to click through to the source will be extremely low. That destroys any incentive to create content and will kill many blogs and sites with high quality content. The likes of Reddit and Twitter can survive this, most blogs have no chance to exist if ChatGPT uses their content but users don't click through.
It surprises me to read criticism on HN about how this can be "misleading". Are you kidding me? This is the internet. Even with regular search results do you blindly trust them or conduct your due diligence?