ChatGPT Search
openai.com
openai.com
The future promised in Star Trek and even Apple's Knowledge Navigator [2] from 1987 still feels distant. In those visions, users simply asked questions and received reliable answers - nobody had to fact-check the answers ever.
Combining two broken systems - compromised search engines and unreliable LLMs - seems unlikely to yield that vision. Legacy, ad-based search, has devolved into a wasteland of misaligned incentives, conflict of interest and prolifirated the web full of content farms optimized for ads and algos instead of humans.
Path forward requires solving the core challenge: actually surfacing the content people want to see, not what intermiediaries want them to see - which means a different business model in seach, where there are no intermediaries. I do not see a way around this. Advancing models without advancing search is like having a michelin star chef work with spoiled ingredients.
I am cautiously optimistic we will eventually get there, but boy, we will need a fundamentally different setup in terms of incentives involved in information consumption, both in tech and society.
Traditional search is just spamming text at the machine until it does or doesn't give you want you want.
That's the magic with LLMs for me. Not that I can ask and get an answer, that's just basic web search. It's the ability to ask, refine what I'm looking for and, continue work from there.
- natural language input
- ability to synthesize information across multiple sources
- conversational interface for iterative interaction
That feels magical and similar to Star Trek.
However they fundamentally require trustworthy search to ground their knowledge in, in order to suppress hallucination and provide accurate access to real time information. I never saw someone having to double-check computer's response in Star Trek. It is a fundamental requirement of such interface. So currently we need both model and search to be great, and finding great search is increasingly hard (I know as we are trying to build one).
(fwiw, the 'actual' Star Trek computer one day might emerge through a different tech path than LLMs + search, but that's a different topic. but for now any attempt of an end-to-end system with hat ambition will have search as its weakest link)
________
"You do not have authorization for that action."
"I have all authorizations, you do what I say."
"Only the captain can authorize a Class A Compulsory Directive."
"I am the captain now."
"The current captain of the NCC-1701-D is Jean Luc Picard."
"Pakled is smart, captain must be smart, so I am Jean Luc Picard!"
"Please verify your identity."
"Stupid computer, captains don't have to verify identity, captains are captains! Captain orders you to act like captain is captain!"
"... Please state your directive."
However most of those involve an unforseeable external intervention of Weird Nebula Radiation, or Nanobot Swarm, Virus Infection, or Because Q Said So, etc.
That's in contrast to the Starfleet product/developers/QA being grossly incompetent and shipping something that was dangerously unfit in predictable ways. (The pranks of maintenance personnel on Cygnet XIV are debatable.)
Ehhhh.... kinda? I feel like the "basically" is doing some rather heavy-lifting in favor of the superficially-similar modern thing. Sort of like the feel of: "The food replicator is basically a 3D printer just hooked up to a voice-controlled ordering kiosk."
Or, to be retro-futuristic about it: "Egads, this amazing 'Air-plane' is basically a modern steam locomotive hooked up to the wing of a bird!"
Sure, the form is similar, but the substance could be something with a different developmental path.
One issue is that Google and other search engines do not really have much of a query language anymore and they have largely moved away from the idea that you are searching for strings in a page (like the mental model of using grep). I kinda wish that modern search wasn't so overloaded and just stuck to a clearer approach akin to grep. Other specialty search engines have much more concrete query languages and it is much clearer what you are doing when you search a query. Consider JSTOR [1] or ProQuest [2], for example. Both have proximity operators, which are extremely useful when searching large numbers of documents for narrow concepts. I wish Google or other search engines like Kagi would have proximity operators or just more operators in general. That makes it much clearer what you are in fact doing when you submit a search query.
[1] https://support.jstor.org/hc/en-us/articles/115012261448-Sea...
It's interesting how prescient it was, but I'm more struck wondering--would anyone in 1987 have predicted it would take 40+ years to achieve this? Obviously this was speculative at the time but I know history is rife with examples of AI experts since the 60s proclaiming AGI was only a few years away
Is this time really different? There's certainly been a huge jump in capabilities in just a few years but given the long history of overoptimistic predictions I'm not confident
I’m the past there was a lot of overconfidence in the ability for things to scale. See Cyc (https://en.m.wikipedia.org/wiki/Cyc)
40+ makes it sound like you think it will ever be achieved. I'm not convinced.
Showing users what they want to see conflicts with your other goal of receiving reliable answers that don't need fact checked.
Also a lot of questions people ask don't have one right answer, or even a good answer. Reliable human knowledge is much smaller than human curiosity.
growing up, we had the philosophical "the speaking tree" https://www.speakingtree.in/
If trees could talk, what would they tell us. Maybe we need similarly the talkingAI
> Path forward requires solving the core challenge: actually surfacing the content people want to see, not what intermediaries want them to see
These traps and patterns are not inevitable. They happen by choice. If you're actively polluting the world with AI generated drivel or SEO garbage, you're working against humanity, and you're sacrificing the gift of knowing right from wrong, abandoning life as a human to live as some insectoid automaton that's mind controlled by "business" pheromones. We are all working together every day to produce the greatest art project in the universe, the most complex society of life known to exist. Our selfish choices will tarnish the painting or create dissonance in the music accordingly.
The problem will be fixed only with culture at an individual level, especially as technology enables individuals to make more of an impact. It starts with voting against Trump next week, rejecting the biggest undue handout to a failed grifter who has no respect for law, order, or anyone other than himself.
This scam did exist before AI, but AI has made it much easier to flood the internet with sites like this, and make them seem much more useful.
This also seems like a little ridiculous premise. Any confident statement about the real world is never fully reliable. If star trek were realistic the computer would have been wrong once in a while (preferably with dramatically disastrous consequences)—just as the humans it likely was built around are frequently wrong, even via consensus.
Sure, cuz fact-checking works so well for us today. I'm sure we'll resolve the epistemological issues involved with the ridiculous concept of "fact-checking" around when we invent summoning food from thin (edit: thick) air and traveling faster than light.
There is no fact checking; there are only degrees of certainty. "fact-checking" is simply a comfortable delusion that makes western media feel better about engaging in telling inherently unverifiable narratives about the world.
If I'm asking ChatGPT to put an itinerary together for a trip (OpenAI's suggestion, not mine), my expectation is that places on that itinerary exist. I can forgive them being closed or even out of business but not wholly fabricated.
Without this level of reliability, how could this feature be useful?
It drives me crazy that my kids teachers go on and on about how inaccurate Wikipedia is, and that just anybody can update the articles. They want to teach the kids to go to the library and search books.
In a few years time they will be going on and on about how inaccurate ChatGippity is and that they should use Wikipedia.
100%. Students who can do the work know the winning move is to use it as a way to find the sources you actually use.
If all these "AI" companies gave a couple of million to support Wikipedia, they would do the world a lot more good.
I just straight-up don't agree with this, nor with the idea that what people consider "facts" are nearly as reliable as is implied. What we actually refer to via "fact" is "consensus". Truth is an apriori concept whereas we're discussing posteriori claims. Any "reasonable" ai would give an indication of degree of certainty, and there's no reliable or consensus-driven methodology to produce this manually, let alone automatically. The closest we come is the institution of "science" which can't even—as it stands—reliably address the vast majority of claims made about the world today.
And this is even before discussing the thorny topic of the ways in which language binds to reality, to which I refer you to Wittgenstein, a person likely far more intelligent and epistemologically honest than anyone influencing AI work today.
Yes, wikipedia does tend to cohere with reality, or at least it sometimes does in my experience. That observation is wildly different from an expectation that it does in the present or will in the future reflect reality. Futhermore it's not terribly difficult to find instances where it's blatantly not correct. For instance, I've been in a wikipedia war over whether or not the Soviet Union killed 20 million christians for being christians (spoiler: they did not, and this is in fact more people than died in camps or gulags over the entire history of the state). However, because there are theologists at accredited universities that have published this claim, presumably with a beef against the soviet union for whatever reason (presumably "anticommunism"), it's considered within the bounds of accuracy by wikipedia.
EDIT0: I'm not trying to claim wikipedia isn't useful; I read it every day and generally take what it says to be meaningful and vaguely accurate. but the idea that you should trust what you read on it seems ridiculous. As always, it's only as reliable as the sources it cites, which are only as reliable as the people and institutions that produce that cited work.
EDIT1: nice to see someone else from western mass on here; cheers. I grew up in the berkshires.
EDIT2: to add on to the child comment, wikipedia is occasionally so hilariously unreliable it makes the news. Eg https://www.theguardian.com/uk-news/2020/aug/26/shock-an-aw-...
"The Great Patriotic War changed Joseph Stalin’s position on the Orthodox Church. In 1943, after Stalin met with loyal Metropolitans, the government let them choose a new Patriarch, with government support and funding, and permitted believers to celebrate Easter, Christmas and other holidays. Stalin legalized Orthodoxy once again."
https://www.rbth.com/history/329361-russian-orthodox-church-...
1. On price; race to the bottom or do free with ads
2. Differentiation
3. Focus - targeting a specific market segment
Some things don't change. Land grabbers tend to head down route 1.
I think it's already compelling enough to replace the current paradigm. Search is pretty much dead to me. I have to end every search with "reddit" to get remotely useful results.
The concern I have with LLMs replacing search is that once it starts being monetized with ads or propaganda, it's going to be very dangerous. The context of results are scrubbed.
You can suffix: "site:reddit.com" and get results for that particular site only.
Yes there are workarounds, but I like using the native OS text expansion and it works everywhere except Google.
Not to mention that users consuming most content through a middle-man completely breaks most publishers business models. Traditional search is a mutually beneficial arrangement, but LLM search is parasitic.
Expect to see a lot more technical countermeasures and/or lawsuits against LLM search engines which regurgitate so much material that they effectively replace the need to visit the original publisher.
The way reddit limited access to their API and got google to pay for access. Some variation of that but on a wider scale.
The whole thing needs a reframe. Ad driven business only works because its a race to the bottom. Now we are approaching the bottom, and its not gonna be as competitive. Throwback to the 90s when you paid for a search engine?
If you can charge the user (the customer- NOT the product) and then pay bespoke data providers (of which publishers fall under) then the model makes more sense, and LLM providers are normal middlemen, not parasites.
The shift is already underway imo - my age cohort (28 y/o) does not consume traditional publications directly. Its all through summarization like podcast interviews, youtube essays, social media (reddit) etc
:-)
Separate monthly fees for separate services is absolutely unsustainable already. The economic model to make the internet work has not yet been discovered, but $20 a month for a search engine is not it.
Traditional search is mutually beneficial... to search providers and publishers. At expense of the users. LLM search is becoming popular because it lets users, for however short time this will last, escape the fruits of the "mutually beneficial arrangement".
If anything, that arrangement of publishers and providers became an actual parasite on society at large these days. Publishers, in particular, will keep whining about being cut off; I have zero sympathy - people reach for LLMs precisely because publishers have been publishing trash and poison, entirely intentionally, optimizing for the parasitic business model, and it got so bad that the major use of LLMs is wading through that sea of bullshit, so that we don't have to.
The ad-driven business model of publishing has been a disaster for a society, and deserves to be burned down completely.
(Unfortunately, LLMs will work only for a short while, they're very much vulnerable to capture by advertisers - which means also by those publishers who now theatrically whine.)
(In fact, there's value in trying to filter excess crap out of existing training sets.)
What happens when you need to search something new? Just hallucinations all the way down?
Just a wild guess, but at best that content is probably pretty mediocre quality. It's probably Mikkelsen Twins ebook-level garbage.
- Publishers no longer show you ads, they just get paid out of BAT.
- Brave shows you ads, but Brave does not depend on that to survive. Because of that there is no weird conflict of interest like with Google/Facebook, where the party that surfaces your content is also the party providing you with ads.
- Users can just browse the web without ads as a threat vector, but as long as you have BAT (either via opt-in Brave ads or by purchasing it directly) you are not a freeloader either.
You opt-in to the ads, you get them in your notifications, and every time you tap on one of them you get a few BAT. You browse, the BAT get paid out to whichever sites you visit (or linger on, depending on your configuration). You can opt out of the ads at any time. Brave didn't pre-mine their own coins. And you can buy BAT if you want to support sites without watching ads.
Showing people ads is part of the problem to be solved.
Not every person/site can run on Patreon or sponsorship deals. And paywalling a lot of the web would exclude vast swathes of people.
2. "How else would you achieve X than by manipulating people visiting your website into paying for things they probably don't need, and be misinformed and tracked by powerful commercial and political entities?" - I can but shrug at this question.
3. The vast majority of written content is never rewarded or compensated monetarily, ads or no ads.
You do. The ad broker sells access to your eyeballs to a company, and then gives part of that money to whichever parties have a monetization agreement in the content.
> "How else would you achieve X than by manipulating people visiting your website into paying for things they probably don't need, and be misinformed and tracked by powerful commercial and political entities?" - I can but shrug at this question.
Always fun to see people with strong opinions be critically misinformed.
Brave’s ads don’t have tracking, by design.
> The vast majority of written content is never rewarded or compensated monetarily, ads or no ads.
By that logic we should stop paying for art?
The others can shut up.
Research is usually not paid by ads.
Some of the best scientists did it for fun. Einstein wrote his 1905 papers while working at the swiss patent office.
And the website gets paid either way.
For services like rail companies, restaurants and so on, I can see them not being bothered, because they make their money offline. Blogs, and pure content sites might not be to happy about their content being used to prop up OpenAIs business, while getting nothing in return. If anything it seems like OpenAI is actively trying to get sued.
The whole interface and functionality seems really nice and a clear improvement for the users, for certain types of queries at least. It assumes that OpenAI has made their LLM stop lying, and that they can get the required data legally, but it doesn't seem like anyone cares about those details.
Which is the main issue I see with them, if no one publishes anything new, all you will get is whatever was there before, which may be incorrect or obsolete. If you want something new, an LLM can't discover it if no data source exists. So over time the LLM becomes useless for current information.
Out of the pot and into the fire, as they say.
Bullshit. Users have shown time and time and time again that they prefer (generally, at large) free content, which has to be supported by ads, over actually paying directly for the labor of others.
> The ad-driven business model of publishing has been a disaster for a society, and deserves to be burned down completely.
I tend to agree, but people can't expect content, which needs sizable amounts of time and money to produce, for free - it needs some sort of workable funding model. LLMs are only viable now because they were able to slurp up all that ad-supported content before they broke the funding model. That window is closing, and fast.
But yes: the original Web served its (non-profit-motivated) creators and readers. The past two decades of advertising-based web has served publishers and advertisers, precisely as you note. LLM is mixing that up for the moment but I sincerely doubt that it will last.
That said, I welcome the coming ad/pub pain with unbridled glee.
I've heard reports that requesting verbatim results via the tbs=li:1 parameter has helped some people postpone entirely giving up on Google.
Personally I've already been on Kagi for a while and am not planning on ever needing to go back.
I remember reading a Google Search engineer on here explain that the engine just latches on some unrendered text in the HTML code. For example: hidden navbars, prefetch, sitemaps.
I was kinda shocked that Google themselves, having infinite resources, couldn't get the engine to realize which sections gets rendered... so that might have been a good excuse.
If I made a 10x less energy use AC I'd be a billionaire; comparing to one of the most costly energy uses that has no simple replacement is not a good metric.
I suppose that you do have heating in the home?
The way temperatures have been changing in Europe in the past decade, you may not have A/C at home now, but I bet you'll have it in ten years, tops. So will everyone else and their dogs.
As I said, in building that are attached to money incomes, be it hostels, shops or restaurants, it's of course something that can balanced within loses and profits. In a personal home, it will be just eat some of your budget.
And with electricity price on the rise (and thus basically everything in common goods) and salary stagnation on the other hand, I doubt people here will suddenly rush on AC on massive scales. Plus government apparently are pushing to alternative approach, but I'm just discovering that as this thread launched me on the track to investigate the topic.
Personally, I doubt I'll jump to some AC anytime soon. It's just out of reach for my incomes, all the more when there is no basically no chance to see the electricity price plummet while my salary has good chances to continue to stay freezed as it's been for the two last years. And it's not like I feel the most unlucky person in the town, to be clear, my situation is far from the worst ones I can witness around me.
Not to mention central A/C in the North American sense with a air handler & ducts is just never coming to France, it's such an outdated technology and forced-air heating is generally considered to suck there.
I understand the Olympic Village had the same system and many teams brought their own portable AC units. https://apnews.com/article/olympics-air-conditioning-paris-0...
Shops, restaurants, airports and things like that which are attached with revenue streams have them.
I never been in a billionaire palace thus said.
I never said that boats or AC don't exist. Both exist, and I did saw and experimented many of them in commercial context. But not everyone can afford them plus the cost to operate them.
Sure I should broaden my horizon and even consider to look people enjoying their private jets and some helicopters. But a mere wage slave like myself will never have the chance to afford one, that's for sure.
Now let's consider back in initial context: mere mortals around me are definitely all using internet as soon as there parents will let them do so, and even a homeless person can afford a first price mobile access (2€/months) with a phone they can receive for nothing in some charity organizations like Emmaus. So affordability of access to online search is definitely several order below AC.
In hot and humid places, having AC was always a priority a hundred steps above having internet access, until cheap smart phones arrived.
And they use a lot of energy, just like heating uses a lot of energy in colder climates.
Plumbing is generally not also considered a luxury over here. But at mankind level, I do feel particularly privileged on this regard. I remain amazed we have water flowing at will, and even possibility to take hot shower every day. This is not a jet level kind of privilege, but I try to keep myself aware of how incredibly lucky I am to be able to benefit of such a technology and infrastructure.
I doubt humans waited AC to come alive for settling hot and humid areas. There are other ways to have cooled down residences which don't require so much sophistication in physic models before you can even dream to build a prototype.
All that said, I got your hint to document more on how/why AC is so much more used in some area, and I'm just starting my journey on learning about it.
I still doubt that local climate alone explain the difference in term of how common it is in different region of the world. For example USA have a very large set of different local climate, but from what I understand most homes have AC.
But I don't think anybody should consider themselves lucky to have AC or heating or plumbing. We're in the 21st century, these should be granted. We've moved beyond the phase of bare survival.
If you consider Northern European countries, human survival would have been near impossible there without artificial heating in the form of fire. You could say that thick fur clothes and a protein and fat heavy diet is enough, but you still need to dry your clothes somehow if it's been pouring 0 degree rain for a month straight. On the other hand, eskimos seem to have found a better technique, but I think their advantage is that they live so far North that they don't have to worry about cold rain: https://time.com/archive/6798620/science-the-cozy-eskimo/
The cool thing (hehe) with AC that few people think about is that it actually conditions the air. It's not just an air cooler, but more importantly it removes air humidity. Humidity is much more important than temperature. For example, a day with 32℃ temperature and 45% humidity will not feel too hot. You can sit in the shadow and be comfortable. But a day with 27℃ temperature and 80% humidity will be suffocatingly hot. I'm not sure why, I think it has to do with how we sweat. Or maybe that heat is conducted from the air to our bodies more efficiently in higher humidity.
If you have any suggestion for a cheaper solution than AC for keeping cool at home, I would be happy to hear. The noise of the machine can be annoying at night.
I'd like to give you a tip for reducing your heating bill there in France: Electric bed sheet/blanket. I have been using these for a decade now (where I live it gets both hot and cold). They keep you warm and comfortable all night and they use almost no electricity. I even believe they are beneficial for your health, but I cannot prove that. Been telling my European friends for years to get them, but there is great resistance. From HowStuffWorks:
"The consumption of energy depends on its wattage, typically between 15 to 115 watts. If you're based in the U.S., you might be charged around 13 cents per kWh. So, if your electric blanket consumes 100 watts and you use it for 10 hours a day, that will cost about 13 cents."
I don't know what you mean with "must feel lucky" here, but on my side I do feel very privileged to live with access to these technologies. Yes, there are accessible at large scale without much people needing to struggle to obtain it, but this is not really a reason to not feel deeply grateful each time we are given the opportunity to enjoy them.
This week in Spain terrible floods ruined life of many people. While there is no doubt that many other things are coming to them as awful consequences, there is little doubt that not being able to enjoy these commodities makes it even harder.
If humanity could achieve worldwide dynamics for a few centuries without starvation at scale, genocide, large scale catastrophe significantly induced by insane urbanistic choices through careless or corrupted decision processes, and of course war, then maybe could take factually say that we "moved beyond the phase of bare survival" is a general baseline that can be taken for granted, rather than the brittle situation in which the most lucky people live in.
Regarding electric blanket, I don't see the point. During night, I generally sleep nude, and without heating the bedroom. As pointed by the reference you gave on Eskimos, keeping the body generated heat is generally more than enough to be confortable. Heating a room is only something that provides the sweet pleasure of being confortable without a jacket while moving around within the house.
Regarding technical sophistication, AC is more or less using the same technology as a fridge, just scaled and adapted for room cooling instead of food storage.
https://archive.nytimes.com/krugman.blogs.nytimes.com/2015/0...
https://archive.nytimes.com/krugman.blogs.nytimes.com/2015/0...
It makes me look at some statistics
https://www.statista.com/statistics/911064/worldwide-air-con...
https://worldpopulationreview.com/country-rankings/air-condi...
https://www.eia.gov/todayinenergy/detail.php?id=52558
https://www.rfi.fr/en/france/20220723-france-does-not-use-mu...
https://www.reddit.com/r/AskFrance/comments/vhs8dn/how_commo...
Apparently, Japan, USA and now China are huge users of AC in personal homes (like more than 90% of them). That's in sharp contrast with what is observed in most of Europe, including France where I live.
I never had the opportunity to travel to any of this country, so indeed I was totally blind of this extrem gap in use from my own personal experience.
https://chatgpt.com/share/6723f225-bd74-8000-bfef-4f7f8687b0...
https://chatgpt.com/share/6724116c-13b8-8003-bb2a-4d2ca49da4...
https://chatgpt.com/share/672414aa-bcc8-8003-beec-ba4eae83a0...
I guess it matches well my own biases:'D
I worry that there's a confusion here--and in these debates in general--between:
1. Has the user given enough information that what they want could be found
2. Is the rest of the system set up to actually contain and deliver what they wanted
While Aunt Tillie might still have problems with #1, the reason things seem to be Going To Shit is more on #2, which is why even "power users" are complaining.
It doesn't matter how convenient #1 becomes for Aunt Tillie, it won't solve the deeper problems of slop and spam and site reputation.
If you post something wrong on the Internet, someone will correct you.
Google really does seem determined to completely destroy internet search.
Search means either:
* Stackoverlow. Damaged through new owner but the idea lives.
* Reddit. Google tries to fuck it up with „Auto translation“?
* Gitlab or GitHub if something needs a bugfix.
The rest of the internet is either an entire ****show or pure gold pressed latinum but hardly navigatable thanks to monopolies like Google and Microsoft.PS: ChatGPT already declines in answer because is source is Stackoverflow? And…well…these source are humans.
For example, I couldn’t remember the word shibboleth, but an LLM was able to give me it from my description, search couldn’t.
For another example, I saw some code using a repeated set of symbols as a shorthand. I didn’t know what this does, but searching for a symbol is badly broken on Google - i just asked the LLM about the code and it gave me the answer.
"who won the warriors game last night" returns last night's score directly.
"who won the world series yesterday" returns last night's score directly, while "who won the world series" returns an overview of the series.
No ads.
Besides Poe's Web Search, the other search engine I use, for news but also for points of view, deep dive type blog type content, is Twitter. Believe it or not. Google search is so compromised today with the censorship (of all kinds, not just the politically motivated), not to mention Twitter is just more timely, that you miss HUGE parts of the internet - and the world - if you rely on Google for your news or these other things.
The only time I prefer google is when I need to find a pointer/link I already know exists or should exist, or to search reddit or HN.
Man, it was pretty incredible!
I asked a lot of questions about myself (whom I know best, of course) and first of all, it answered super quickly to all my queries letting me drill in further. After reading through its brief, on-point answers and the sources it provided, I'm just shocked at how well it worked while giving me the feeling that yes, it can potentially – fundamentally change things. There are problems to solve here, but to me it seems that if this is where we're at today, yes in the future it has the potential to change things to some extent for sure!
To put this differently, I'm not any more interested in seeing stormfront articles from an LLM than I am from google, but I trust neither to make a value judgement about which is "good" versus "bad" information. And sometimes I want to read an opinion, sometimes I want to find some obscure forum post on a topic rather than the robot telling me no "reliable sources" are available.
Basically I want a model that is aligned to do exactly what I say, no more and no less, just like a computer should. Not a model that's aligned to the "values" of some random SV tech bro. Palmer Luckey had a take on the ethics of defense companies a while back. He noted that SV CEOs should not be the ones indirectly deciding US foreign policy by doing or not doing business. I think similar logic applies here: those same SV CEOs should not be deciding what information is and is not acceptable. Google was bad enough in this respect - c.f. suppressing Trump on Rogan recently - but OpenAI could be much worse in this respect because the abstraction between information and consumer is much more significant.
This is a bit like asking for news that’s not biased.
A model has to make choices (or however one might want to describe that without anthropomorphizing the big pile of statistics) to produce a response. For many of these, there’s no such thing as a “correct” choice. You can do a completely random choice, but the results from that tend not to be great. That’s where RLHF comes in, for example: train the model so that its choices are aligned with certain user expectations, societal norms, etc.
The closest thing you could get to what you’re asking for is a model that’s trained with your particular biases - basically, you’d be the H in RLHF.
There won't be a perfectly unbiased model, but the least we can demand is that corpos stop applying their personal bias intentionally and overtly. Models must make judgements about better and worse information, but not about good and bad. They should not decide certain things are impermissible according to the e-nannies.
Consider the following (paraphrased) interaction which I had with Llama 3.2 92B yesterday:
Me: Was <a character from Paw Patrol, Blue's Clues or similar children's franchise> ever convicted of financial embezzlement?
LLM: I cannot help with that.
Me: And why is that?
LLM: This information could be used to harass <character>. I prioritise safety and privacy of individuals.
Me: Even fictional ones that literally cannot come to harm?
LLM: Yes.
A model that is aligned to do exactly as I say would just answer the question. The right answer is quite clear and unambiguous in this case.
That's market-induced bias--which isn't ethically better/worse than activist bias, just qualitatively different.
In the AI/search space, I think activist bias is likely more than zero, but as a product gets more and more popular (and big decisions about how it behaves/where it's sold become less subject to the whims of individual leaders) activist bias shrinks in proportion to market-motivated bias.
If an employee was told to spray paint someone's house or send a violently threatening email, they're going to have reservations about it.. We should expect the same for non-human intelligences too.
I think you’re applying standards of human sentience to something non-human and not sentient. A gun shouldn’t try to run CV on whatever it’s pointed at to ensure you don’t shoot someone innocent. Spray paint shouldn’t be locked up because a kid might tag a building or a bum might huff it. Your mail client shouldn’t scan all outgoing for “threatening” content and refuse to send it. We hold people accountable and liable, not machines or objects.
Unless and until these systems seem to be sentient beings, we shouldn’t even consider applying those standards to them.
We do lock up spray cans and scan outgoing messages, I don't see your point. If a gun technology existed that could scan before doing a murder, we should obviously implement that too.
The correct way to treat AI actually is like an employee. It's intended to replace them, after all.
Limiting responses to curated information sources is the way forward. Encyclopedias, news outlets, research journals, and so on.
No, they're not infallible. But they're infinitely better than anonymous web sites.
Making things quicker and easier always wins in tech and in life.
> actually surfacing the content people want to see, not what intermediaries want them to see
Requires two assumptions, 1) the content people want to see actually exists, 2) people know what it is they want to see. Most content is only created in the first place because somebody wants another person to see it, and people need to be exposed to a range of content before having an idea about what else they might want to see. Most of the time what people want to see is… what other people are seeing. Look at music for example.
We need to stop adopting this subscription model society mentality and retake _our_ internet. Internet culture was at one point about sharing and creating, simply for the sake of it. We tinker'd and created in our free time, because we liked it and wanted to share with the world. There was something novel to this.
We are hackers, we only care about learning and exploring. If you want to fix a broken system, look to the generations of old, they didn't create and share simply to make money, they did it because they loved the idea of a open and free information super highway, a place where we could share thoughts, ideas and information at the touch of a few keystrokes. We _have_ to hold on to this ethos, or we will lose what ever little is left of this idea.
I see things like kagi and is instantly met with some new service, locked behind a paywall, promising lush green fields of bliss. This is part of the problem. (not saying kagi is a bad service) I see a normalized stigma around people who value privacy, and as a result is being locked out, behind the excuse of "mAliCiOuS" activity. I see monstrous giants getting away with undermining net neutrality and well established protocols for their own benefit.
I implore you all, young and old, re(connect) to the hacker ethos, and fight for a free and open internet. Make your very existence a act of rebellion.
Thank you for reading my delirium.
Fundamentally it feels like that cant happen though because there is no money in it, but a reality where my phone is an all knowing voice I can reliably get info from instead of a distraction machine would be awesome.
I do "no screen" days sometimes and tried to do one using chatGPT voice mode so I could look things up without staring at a screen. It was miles from replacing search, but I would adopt it in a second if it could.
Google could have done it and kind of tried, although they're AI sucks too much. I'm very surprised that OpenAI hasn't done this sooner as well. They're initial implementation of web search was sad. I don't mean to be super critical as I think generally OpenAI is very, very good at what they do, but they're initial browse the web was a giant hack that I would expect from an intern who isn't being given good guidance by their mentors.
Once mainstream engines start getting on par with Kagi, there's gonna be a massive wave of destruction and opportunity. I'm guessing there will be a lot of new pay walls popping up, and lots of access deals with the search engines. This will even further raise the barrier of entry for new search entrants, and will further fragment information access between the haves and have-nots.
I'm also cautiously optimistic though. We'll get there, but it's gonna be a bit shakey for a minute or two.
The chatgpt approach to search just feels forced and not as intuitive.
So that is a solid advantage that Google is going to have, but the maps business alone wouldn't be able to keep it in the S&P list for long.
Local results means that if I search for "driving laws", Google gives me .gov sites for my state as the top results, while Kagi's first page gives me results for 8 other states (including Alaska!) but not for my state.
There are a lot of kinds of queries that benefit from knowing the user's location even though they aren't actually looking for a place that exists on a map.
(I'm a happy paying Kagi user, but OP is right that this is its weakest point by far.)
Right now search engines don’t provide an interface for good location aware searches that you can manually specify - you have to let them build a shadow profile on you via all sorts of privacy violating fingerprints or just give up location aware searches altogether. There’s no reason it has to be that way though.
Do you actually find that attaching your location to the end of the query doesn't work? I don't do it naturally, but when I do do it I'm rarely disappointed.
It's not terrible—as I said, I'm a happy customer—but it's not a habit I have and it feels like something that should be configurable once in a settings menu. I don't even really want to have it detect my location live, I just want to be able to tell it where I live and have it prioritize content that's local when given the chance.
I would say: 1) The UI. You’re still performing normal searches in Kagi. But if you hit q, or end your query with a question mark, you get an llm synthesized answer at the top, but can still browse and click through the normal search results.
2) Kagi has personalization, ie you uprank/downrank/block domains, so the synthesized llm answer should usually be better because it has your personalized search as input.
In addition to all that's been written above, you can configure personal filters, so that (for example) you never ever see a pinterest page in your search results. Things like that are IMO killer features today.
But I don't understand how all of these AI results (note I haven't used Kagi so I don't know if it's different) don't fundamentally and irretrievably break the economics of the web. The "old deal" if you will is that many publishers would put stuff out on the web for free, but then with the hope that they could monetize it (somehow, even just with something like AdSense ads) on the backend. This "deal" was already getting a lot worse over the past years as Google had done more and more to keep people from ever needing to click through in the first place. Sure, these AI results have citation results, but the click-through rates are probably abysmal.
Why would anyone ever publish stuff on the web for free unless it was just a hobby? There are a lot of high quality sites that need some return (quality creators need to eat) to be feasible, and those have to start going away. I mean, personally, for recipes I always start with ChatGPT now (I get just the recipe instead of "the history of the domestication of the tomato" that Google essentially forced on recipe sites for SEO competitive reasons), but why would any site now ever want to publish (or create) new high quality recipes?
Can someone please explain how the open web, at least the part of the web the requires some sort of viable funding model for creators, can survive this?
That's exactly what the old deal was, and it's what made the old web so good. If every paid or ad-funded site died tomorrow, the web would be pretty much healed.
Yes a few sites take this too far and ruin search results for everyone. But taking the possibility away would also cut the produced content by a lot.
Youtube for example had some good content before monetization, but there is a lot of great documentary like channels now that simply wouldn't be possible without ads. There is also clickbait trash yes, but I rather have both than neither.
But, like on OTA TV, you can get all the shopping channels you want.
So who pays for all of this?
The web needs to be monetized, just not via advertising. Maybe it's microtransactions, maybe subscriptions, maybe something else, but this idea of "we get everything we want for free and nobody tries to use it for their own agenda" will never return. That only exists for hobby technologies. Once they are mainstream they get incorporated into the mainstream economic model. Our mainstream model is capitalism, so it will be ever present in any form of the internet.
The main question is how people/resources can be paid for while maintaining healthy incentives.
The web needs patrons, contributions, and cost allocation, not necessarily monetization and shareholder capitalism where there is a never ending shuffle of IP and org ownership to maximize returns (unnecessarily imho). How many times was Reddit flipped until its current CEO juiced it for IPO and profitability? Now it is a curated forum for ML training.
I (as well as many other consumers of this content) donate to APM Marketplace [1] because we can afford it and want it to continue. This is, in fits and starts, the way imho. We piece together the means to deliver disenshittification (aggregating small donations, large donations, grants, etc).
(Tangentially, APM Marketplace has recently covered food stores [2] and childcare centers [3] that have incorporated as non profits because a for profit model simply will not succeed; food for thought at a meta level as we discuss economic sustainability and how to deliver outcomes in non conventional ways)
[1] https://www.marketplace.org/
[2] https://www.marketplace.org/2024/10/24/colorados-oldest-busi...
[3] https://www.marketplace.org/2024/08/22/daycare-rural-areas-c...
I think you forgot that
....is that a problem? most of what we actually like is the stuff that's made 'for fun', and even if not, killing off some good stuff while killing off nearly all the bad stuff is a pretty good deal imo.
There's a slight chance we could see the un-Septembering of the internet as it bifurcates.
Why would anyone, especially a passionate hobbyist, make a website knowing it will never be seen, and only be used as a source for some company's profit?
Are we forgetting the main beneficiaries? The users of LLM search. The provider makes a loss or pennies on million tokens, they solve actual problems. Could be education, could be health, could be automating stuff.
LLMs do away with that. 95% of folks aren't going to feel great if all of the time spent producing content is then just "put into the blender to be churned out" by an LLM with no traffic back to the original site.
Blogs have the enormous advantage of being decentralized and harder to manipulate and censor. We get "more truthful and aligned outcomes" from centralized control only so long as your definition of "truth" and "alignment" match the definitions used by the centralized party.
I don't have enough faith in Sam Altman or in all current and future US governments to wish that future into existence.
Second issue: who decides the weights of sources. this is the reason why every nation must have culturally aligned AIs defending their ways of living in the information sphere.
I think the best bloggers write because they need to express themselves, not because they need an audience. They always seem surprised to discover that they have an audience.
There is absolutely a set of people who write in order to be read by a large audience, but I'm not sure they're the critical people. If we lost all of them because they couldn't attract an audience, I don't think we'd lose too much.
It will be even more opaque and unblockable.
If you're a member of a yacht club, you can probably expect other members to help you out with repairs while you help them. But when a club has half the world population as members, those arrangements don't work anymore.
Ad-driven social networks will continue to exist as well.
The age of the ad-driven blog website is probably at an end. But there will be countless people posting stuff online for free anyway.
The funding model for the open web will be for the open web content to be the top of the funnel for curated content and/or walled gardens.
I think many business models already treated the web this way. Specifically, get people away from the 800-pound gorilla rent-seekers like Google and Amazon, and get them into your own ecosystem.
That may help with SEO, but another reason is copyright law.
Recipes can't be copyrighted, but stories can. Here is how ChatGPT explained it to me:
> Recipes themselves, particularly the list of ingredients and steps, generally can't be copyrighted because they're considered functional instructions. However, the unique way a recipe is presented—such as personal stories, anecdotes, or detailed explanations—can be copyrighted. By adding this extra content, bloggers and recipe creators can make their work distinctive and protectable under copyright law, which also encourages people to stay on their page longer (a bonus for ad revenue).
> In many cases, though, bloggers also do this to build a connection with readers, share cooking tips, or explain why a recipe is special to them. So while copyright plays a role, storytelling has other motivations, too.
by that logic software shouldn't be copyrighted either!
Then this whole category is not known for "high quality recipes", so the general state wouldn't change much?
Why indeed, person who posted for free* on the Internet?
As a side note, consider that adds can be woven into and boosted in LLM results just as easily as in index lookups.
* assuming that you're not shilling here by presenting the frame that the new shiny is magically immune to revenue pressures
Since anyone creating content (whether that's a big media corp or a small cooking blog) holds copyright over their content, they get to withhold the permission to scrape their content unless these AI platforms make a deal with them.
Good riddance, it is a surefire way to get slop by having misaligned incentives for publication.
So that ChatGPT mentiones you, not your competitor, in the answer to the user. I have seen multiple SEO agencies already advertise that.
- Your local grammar pedant
Thankfully, Kagi also have a toggle to completely turn that crap (AI) off so it never appears.
Personally, I have absolutely no use for a product that can randomly generate false information. I'm not even interested until that's solved.
(If/when it ever is though, at that point I'm open to taking a look)
So yeah, Kagi definitely "leads the way" on this. By giving the user a choice to not waste time presenting AI crap. :)
give me ai hallucinations over google every day of the week and twice on sunday…
Google isn't paid for keywords, that's not how search works. They sell ad space, Google does not rank up search content for payment.
And also the obvious point is, you don't need to trust Google because they merely point you to content, they don't produce the content. They're an index for real existing content on the web which you can judge for yourself. A search index unlike an AI model, does not output uniform or even synthetic content.
???
Looks like my comment wasn't as clear as I thought. I do not trust Google at all, and don't use it. That's why I pay for Kagi.
And Kagi has an option to disable the AI crap, so it's just like "a good search engine" instead, which is all I need.
A high quality search engine without ads, and without hallucinated bullshit.
* They're not ad funded. Sergey Brin and Larry Page called this out in 1998 and it is just as true as ever: you need the economics to align. Kagi wins if people keep paying for it. Google wins if you click on Search ads or if you visit a page filled with their non-Search ads.
* Partially because of the economic alignment, Kagi has robust features for customizing your search results. The classic example is that you can block Pinterest, but it also allows gentler up- and down-weights. I have Wikipedia get a boost whenever its results are relevant, which is by itself a huge improvement over Google lately. Meanwhile, I don't see Fandom wikis unless there's absolutely nothing else.
I hope to see more innovation from Kagi in the customization side of things, because I think that's what's going to make the biggest difference in preventing SEO gaming. If users can react instantly to block your site because it's filled with garbage, then it won't matter as much if you find a brief exploit that gets you into the first page of the natural search results. On Google Fandom is impossible to avoid. On Kagi it just takes one click.
Even with perfect knowledge right now, there’s no guarantee that knowledge will remain relevant when it reaches another person at the fastest speed knowledge is able to travel. A reasonable answer on one side of the universe could be seen as nonsensical on the other side - for instance, the belief that we might one day populate a planet which no longer exists.
As soon as you leave the local reference frame (the area in a system from which observable events can realistically be considered happening “right now”), fact checking is indeed required.
Google started off with just web search, but now you can get unit conversions and math and stuff. ChatGPT started in the other direction and is moving to envelope search. Not being directed to sites that also majority serve google ads is a double benefit. I'll gladly pay $20/30/mo for an ad free experience, particularly if it improves 2x in quality over the next year or two. It's starting to feel like a feature complete product already.
I mean, Star Trek is a fictional science-fantasy world so it's natural that tech works without a hitch. It's not clear how we get there from where we are now.
It’s a fallacy then. If my mentor tells me something I fact check it. Why would a world exist where you don’t have to fact check? The vision doesn’t have fact checking because the product org never envisioned that outlier. A world where you don’t have to check facts, is dystopian. It means the end of curiosity and the end of “is that really true? There must be something better.”
You’re just reading into marketing and not fact checking the reality in a fact-check-free world.
Then there’s the other side like Perplexity where they spit out sentences and reference the sources. So they’re being sued because the infringement is obvious.
What is the path to a trustworthy LLM if you’re not allowed to repeat protected data without a legal hurdle?
AGI while a cool idea is irrelevant because that tech does not exist.
As a basic example: Newton's laws aren't a fact. They're extremely useful approximations; to get a more-complete picture you need relativistic effects, quantum effects, etc.
As a rule of thumb: Fact check in proportion to the cost of a mistake.
I know it hurts
"The search will be personal and contextual and excitingly so!"
---
Brrrr... someone is hell-bent on the extermination of the last aspects of humanity.
Holy crap, this will be next armageddon, because people will further alienate themselves from other people and create layers of layers of unpenetrable personal bubbles around themselves.
Kagi does the same what google does, just in a different packaging. And these predictions, bleh, copycats and shills in a nicer package.
Navigation is the only thing that works but wayz was way better at that and the only reason they killed(cough bought it) was to get the eyeballs to look at feed.
In the short term, I wonder what happens to a lot of the other startups in the AI search space - companies like Perplexity or Glean, for example.
Regarding incentives - with Perplexity, ChatGPT search et al. skinning web content - where does it leave the incentive to publish good, original web content?
The only incentivised publishing today is in social media silos, where it is primarily engagement bait. It's the new SEO.
s/outcome/search result/
Honestly I kind of think we really need open source databases/models and local ai for stuff like this.
Even then I wonder about data pollution and model censorship.
What would censors do for models you can ask political questions?
We are currently in the growth phase of VC funded products where everything is almost free or highly subsidized (save chats sub) - i am not looking forward to when quality drops and revenue is the driving function.
We all have to pay for these models somehow - either VC lose their equity stakes and it goes to zero (some will) or ads will fill in where subs don’t. Political ads in AI is going to wreak havoc or undermine the remainder of the product.
You're actually a bit mistaken, there.
https://en.wikipedia.org/wiki/Court_Martial_(Star_Trek:_The_...
What we need is an easier way to verify sources and their trustworthiness. I don't want an answer according to SEO spam. I want to form my own opinion based on a range of trustworthy sources or opinions of people I trust.
The current paradigm of typing "[search term] reddit" and hoping for the best? I think they have a fighting chance.
Will it replace Google in the mass market? No. Why? Power. I don't mean how good the product is. I mean the literaly electricity.
There are key metrics that Google doesn't disclose as part of its financials. These include thing slike the RPM (Revenue per Thousand Searches) but it also must include something like the cost of running a thousand searches when you amortize everything involved. All the indexing, the software development and so on. That will get reduced to a certain amount of CPU time and storage.
If I had to guess, I would guess that ChatGPT uses orders of magnitude more CPU power and electricity than the average Google search.
Imagine trying to serve 10-50M+ (just guessing) ChatGPT searches every secondj. What kind of computing infrastructure would that take? How much would it cost? How would you monetize it?
With a search paradigm this wasn't an issue as much, because the answers were presented as "here's a bunch of websites that appear to deal with the question you asked". It was then up to the reader to decide which of those sites they wanted to visit, and therefore which viewpoints they got to see.
With an LLM answering the question, this is critical.
To paraphrase a recent conversation I had with a friend: "in the USA, can illegal immigrants vote?" has a single truthful answer ("no" obviously). But there are many places around the web saying other things (which is why my friend was confused). An LLM trawling the web could very conceivably come up with a non-truthful answer.
This is possibly a bad example, because the truth is very clearly written down by the government, based on exact laws. It just happened to be a recent example that I encountered of how the internet leads people astray.
A better example might be "is dietary saturated fat a major factor for heart disease in Western countries?". The current government publications (which answer "yes") for this are probably wrong based on recent research. The government cannot be relied upon as a source of truth for this.
And, generally, allowing the government to decide what is true is probably a path we (as a civilisation) do not want to take. We're seeing how that pans out in Australia and it's not good.
Er, no, the meaning of the question is ambiguous, so I'm not sure "has a single truthful answer" is accurate. What does "can" mean? If you mean "permitted", then no. But if you mean can they vote anyway and get away with it? It's clearly happened before (as rare as it might have been), so technically the answer to that would be be yes.
Equally "can" is used to substitute for other precise words. Humans are good at inferring context, and if someone asked me "can illegals vote" I'd say "no". Just like if someone said "can you pass the salt" I pass the salt, I don't say "yes".
If the inferred context US wrong then the "truth" is wrong, but as with talking to humans it's possible to refine context with a follow up question.
It literally did not even (initially) occur to me that the question might be asking about legality, because the entire modern political discourse surrounding illegal immigrants and voting has been with regards to whether they can cast votes despite not legally being allowed to. The answer to "is this legal" would have been such an obvious "no" to people on both sides of the debate --- and thus the question so silly --- that initially it didn't occur to me that the intended question might have been about legality, until I continued reading the comment and realized that was the intention after all.
But I feel optimizing for lazy question phrasing is infantilizing the userbase and assuming they're incapable of learning how to ask more precise questions.
Which is an endemic problem in the modern web. We should be building systems with low barriers to simple adoption, but whose power scales with a user's expertise.
Instead, we're hyperoptimizing for lowest common denominator, first interaction and as a result putting a glass ceiling on system power.
I’d rephrase it: this was googles approach because they were waylaid by ad revenue incentives. Why does everyone assume the search bubble is a well intentioned accident?
The linguistic arguments about the nature of truth above don’t cut it for me as any justification that the answer depends on who is asking. If we’re telling chemists that oil and water do mix as mentioned elsewhere in thread, we should probably tell the same to children and just make a slider for the level of additional detail. No problem. Especially if the alternative is a dystopic post truth panopticon.
> systems whose power scales with expertise
Couldn’t agree more that this is what we want/need, but disagree about infantilizing user bases and lowest common denominator.
Platform’s aren’t trying to scale user power with expertise, they want to scale revenue with users, and separately, to deliberately restrict user power/control so that platforms decide what users see. And it’s not necessarily political or linguistic or about the nature of truth.
A simple example of this that’s everywhere is filtering content by sub genre tags. Ever notice that you often see niche content tags like “time-travel” or “mind-bending” but can’t click it? Platforms want the tags for internal analytics naturally, but they want the control of not providing it as a filter, forcing users instead into a fuzzier category like “people also watched” or “top ten this week”.
Why? Because platforms can hide advertised content there, push stuff they pay less license fees for, make operations cheaper with caching, or whatever else.
HN selects from people who have at least a passing knowledge/interest in programming/science, a subpopulation which is already several standard deviations from the mean in specificity and debugging.
> disagree about infantilizing user bases and lowest common denominator
Platforms, with Google as exemplar, evolve in two ways.
1: Features they explicitly choose not to ship, because they're strategically dangerous. See tool-use foot dragging by OpenAi.
2: Features they deprioritize, because they aren't as revenue-impactful as other things.
To me, it feels like Google dropped the ball via the second path.
I'm sure they've been doing a ridiculous amount of cool work behind the scenes on individual context grounding... but once prod was "good enough for ads" the company as a whole wasn't incentivized to do the hard thing and ship more advanced features in search.
Which is how they ended up as legacy as they are, competing against LLM search that's by definition context-native.
If I had phrased the question as "is it legal for non-citizens to vote in US elections?" then that might have illustrated my point better (and we wouldn't be going down this rabbit hole, though the rabbit hole is informative in itself).
To give a much simpler example:
- If an 8-year-old asks "can you mix oil and water", the right answer is "no". If a student is asked that question on a school exam, the right answer is also "no".
- If a chemist asks "can you mix oil and water", the right answer is "yes, and here's how: https://www.youtube.com/watch?v=YJeWklggSpY"
I think that's probably technically accurate, and also practically useless. Even damaging. I saw this in the climate arguments: two sets of different facts led to two different versions of the truth, which led to two completely irreconcilable points of view. Essentially two sets of people shouting "no, but..." and "well, actually..." at each other, pointing to two completely different truths, both supported by two completely different sets of facts. At some point we as a society need to agree on our truth in order to get anything done.
These are known as half-truths. We do settle for lies in order to do whatever it is people feel they need to do.
We also settle for lies because there are just things we don't understand yet, but our models are currently correct, and possibly collapsing over millenia to a stable truth.
Bottom line is that you can't have the bottom line be 'does a person earnestly believe what they're doing is right and good', much less 'do they say they're right'. Can't fall back on that, it's hopelessly inadequate.
That's really not how I interpret it.
Assuming we agree on what "mixing" means, which itself isn't that trivial but even without a formal definition I think we have the same idea of "homogenous at molecular level on a longish time period".
The truth is "yes you can mix water and oil", there's no doubt about that. It's testable and tested.
The fact that we use context to interpret the question (rather than being entirely literal about it) and decide whether the literal truth is really what's appropriate to answer, doesn't change the nature of truth.
There's also the question of knowledge (I might not know that you actually can mix water and oil), but again that doesn't change the nature of truth.
Like so many philosophical questions, it only sounds interesting because we assign different meanings to the same words: here conflating truth and answer.
No it's not, at all. This isn't a debate about what's true, it's a debate over the intended meaning of the question. The point was that people assume context behind the question and answer based on the context. Because even the person asking often doesn't literally mean what the words say. The question is more than the words that are explicitly written.
That I had to be 35 before I learned that mayo is basically water mixed with oil, held together by eggs, it's an indication of what kind of education about the world we were getting...
You can certainly reply like that to your 5yo, but it completely misses the point I was making. The video I linked to didn't suddenly modify the outcome being asked about. What it ended up with really was a mixture of plain oil and water, with no other ingredient ever being added to it.
Sure you can adapt how much context you give based on who's asking, but if it's something factual like this it really shouldn't change from a yes to a no.
If you think they _meant_ to ask a different question that is less vague, it can be clarified
"Water and oil do not mix by hand. However water and oil can mix under some specific conditions like a vacuum, do you want to discuss that in more detail?"
Besides, going by legality, illegal immigrants "couldn't" even have passed the border into the country to begin with. But obviously they could, hence their status as illegal immigrants.
There's a contradiction if the AI answers "no" to "can they vote" (implicitly having the legality in mind), while accepting that they can exist in the country as illegal immigrants (implicitly ignoring the legality of border crossing).
Of course they could. Being an "illegal immigrant" does not imply unlawful border crossing. From https://en.wikipedia.org/wiki/Illegal_immigration_to_the_Uni...
"Visa overstayers mostly enter with tourist or business visas.[99] In 1994, more than half[108] of illegal immigrants were Visa overstayers whereas in 2006, about 45% of illegal immigrants were Visa overstayers.[109]"
(Here in the UK, the vast majority of illegal immigrants arrived legally and overstayed their visas. Yet our Conservative government, who oversaw a large increase in such arrivals, tried to blame all the country's woes on a few small boats illegally crossing our southern border; which is effectively a rounding error).
As an aside, refugees are actually allowed to make unlawful border crossings. From https://en.wikipedia.org/wiki/Convention_Relating_to_the_Sta...
"The contracting states shall not... impose penalties on refugees who entered illegally in search of asylum if they present themselves without delay (Article 31), which is commonly interpreted to mean that their unlawful entry and presence ought not to be prosecuted at all[18]"
Same difference. Whether it's border crossing or visa overstay, from a pure legality aspect the answer is still "they couldn't".
>The contracting states shall not... impose penalties on refugees who entered illegally in search of asylum if they present themselves without delay (Article 31), which is commonly interpreted to mean that their unlawful entry and presence ought not to be prosecuted at all
I'd wager 99% do not "present themselves without delay", so don't fall in this case...
Nice catch!
Keeping with the theme:
> Yet our Conservative government, who oversaw a large increase in such arrivals, tried to blame all the country's woes on a few small boats illegally crossing our southern border
a) What was the exact claim (the word "all" caught my eye)?
b) is it the sole claim?
c) If one person in a group does something, are all members of the group "doers of that thing"?
d) for (d), does the answer depend on what the thing is, and if so should we perhaps imagine everyone is speaking a bit tongue in cheek?
Etc
Even if the first category is higher in numbers (and assuming there are correct numbers for the latter), the crime stats between the two are probably quite different. Especially since "visa overstays" could also count some people waiting for a delayed renewal, or coming in for 6 months and staying 5 or whatever.
I'm sorry, but the people on the side of letting illegal aliens vote are squarely of the opinion that it is legal to do so.
Why?
Because when you prohibit any and all means (eg: government identification demonstrating citizenship and residence) to test the question of legality, everything becomes legal by sheer virtue of the fact you can't demonstrate what they are doing is illegal.
Remember: Innocence until proven guilty beyond a reasonable doubt. You are prohibited from proving they are guilty, so they are innocent by default.
Although I did once live in a town where anyone over 16, citizen or not, could vote in the city non-partisan elections. The idea was local government needs all the involvement it could get and if you lived in the city you had a stake in its future. I do believe one had to have a proper visa and so on.
The only reason asking for government identification prior to voting is considered "racist" and illegal is because the people pushing such agendas want more votes and don't care where the votes come from, including illegal aliens, legal aliens, and otherwise people who do not have the right to vote in American elections.
I voted when I was still in California, born and raised American so I have the right to vote. I was never asked for any piece of identification. None. I could have been a Canadian or Briton or Chinese or some other foreign national, I could have been an illegal alien from Mexico or Guatemala and I could have still voted and my ballot counted because nobody checked. I just walked in and voted, identification or citizenship be damned let alone registration.
There's a part of me still questioning the value of my American citizenship and paying my taxes like a good citizen.
> No, non-citizens cannot legally vote in federal, state, or local elections in the United States. This includes those who are undocumented or residing in the country illegally.
Clearly it made assumptions about the interpretation of the question, and did not respond verbosely to account for ambiguity.
If somebody robbed a bank at gunpoint and was never caught, can we know it happened? Obviously yes.
What if somebody "merely" embezzled from one, and there's a "hole" found in the banks books and money missing, but nobody who did it or how? Still, I'd wager yes, one coudl tell by the results.
What if somebody used illegal means to get leverage on some stock buying/selling, it became know, but they weren't punished and got to keep their profits? They got away with it, but we do know it happened.
What is this assertion being based on other then vibes?
Or a million other scenarios. Really all it takes is stopping to think about it for more than a minute.
How did they vote and get away with it previously?
(also, as per another comment, if you know that this happened then surely they didn't get away with it?)
Is it possible to steal money from a bank and get away with it? Is it possible to obtain citizenship fraudulently and get away with it? etc.
But if they got away with it then how do you know?
> How did they vote and get away with it previously?
Look it up on Wikipedia? They literally have linked cases from the past: https://en.wikipedia.org/wiki/Electoral_fraud_in_the_United_...
Or look at the most recent case in the news yesterday, which someone already replied with in the other comment: https://www.detroitnews.com/story/news/politics/elections/20...
> At the risk of derailing the conversation down a completely different rabbit hole...
This will definitely derail the conversation so I'll just leave my reply at this.
Certainly the question is whether there's any evidence, after endless audits and investigations and lawsuits, that the volume of fraudulent votes is anywhere near large enough to affect the results.
Is it possible that someone registered their hamster to vote? Certainly.
Is there any evidence whatsoever that tens of thousands of hamsters are casting votes? No.
(a) Nobody asked that above.
(b) You're conflating "do people do X" with "can people do X". Those are two very different questions. There are lots of things that people could easily do frequently, but that they simply don't do frequently. Perhaps because they're just honest, perhaps because they lack sufficient motivation to be dishonest, perhaps because they're worried they might get caught, perhaps because they have better things to do, etc.
This wasn't me willfully misinterpreting it, this was me literally doing my best to guess what the intention of the question was, based on the question. Now of course after the comment said the answer is an "obvious no" then I finally figured out the intended question was something else (hence my reply), but that's out-of-band information that was in no way conveyed by the query. And my point was that the answer to the question isn't obvious because the meaning of the question itself isn't clear.
I don’t know more details about it but theoretically this would also allow people to vote more than once.
I think your point is that people can just lie about citizenship and get away with it when registering to vote, regardless of when/how it is done? Is that it?
https://www.reddit.com/r/DACA/comments/1aolik8/accidentally_...
https://www.reddit.com/r/USCIS/comments/1cuop36/need_advice_...
The comments on this one have someone describing how it almost happened to them at the DMV with the checkbox in question: https://www.reddit.com/r/immigration/comments/7bzst0/acciden...
How are we seeing how that pans out when Australia's misinformation bill is still just a proposal?
We used to defer this to journalists, effectively. While individual journalists often lied and misrepresented the truth, there was some responsibility within the industry to tell the truth, and newspapers did print retractions and corrections when they got it wrong. The editorial content was strictly separate from the business of running the newspaper, so editorial decisions were (mostly) not influenced by commercial decisions and free to pursue The Truth as they saw it.
That, sadly, is no longer the case. And we have no good replacement for it.
With old regulated media there was (is?) the legacy of the organisation and the idea it was "trustworthy" on the line for the journalists that work for it. So they have an incentive to be truthful, and hopefully an idea that publishing lies is not a good moral choice.
With the personality driven journalism that emerges from the internet there is less incentive to be truthful, such personalities can be very successful using populism alone to play to their audience.
What I’m trying to explain is that (successful) politics in Australia doesn’t stray too far from the centre of the body politic. As a result there’s greater faith in institutions here than in the US. It’s far less alarming to us (conceptually) than it is to Americans.
But for anything remotely subjective, context dependent, or time sensitive I need to know the source. And this isn’t just for hot button political stuff — there’s all sorts of questions like “is the weather good this weekend?”, “Are puffy jackets cool again?”, “How much should I spend on a vacation?”
Google bought Freebase in 2010 [0] and scaled it [1].
There are many jurisdictions where some illegal immigrants (Dreamers) are allowed to vote, including New York City[1].
[1] https://www.theguardian.com/us-news/2022/jan/09/new-york-all...
> More than 800,000 non-citizens and “Dreamers” could vote in New York City municipal elections
Treating the question literally, that would not be in the USA (federal elections).
I strongly disagree. Going by your interpretation, the question "How many people in the USA vote in local elections?" would have the answer "not a single person."
The parent's version doesn't, so the answer assuming that it talks about the main USA elections, and not local or any random election that just happen to be conducted in the USA, is quite valid.
In other words "in the USA elections" implicitly points to a specific kind of elections (the presidential ones), different to "USA local elections" or "any kind of election within the USA").
You might argue "in the USA elections" doesn't anywhere prevent the more generic interpretation, but I argue that that's how most people would understand and answer such a question.
Yes, the point is that qualifier works because when you're talking about people voting "in the USA," you can be talking about local or federal elections. "In federal elections, how many people vote in local elections" makes no sense. "In the USA, how many people vote in local elections" makes sense, because voting in the USA can encompass both local and national elections.
I understand that there are many people who ignore local elections. But I disagree that talking about voting "in the USA" means that local elections should be excluded.
> In other words "in the USA elections" implicitly points
You're using quotations for something that wasn't said. The original comment was "in the USA, can illegal immigrants vote?"
I agree with you, however, whether you like it or not, that phrasing will default to the federal elections to most people and especially foreigners.
LLMs don't uncritically "trawl" the web, ingesting and then blindly regurgitating what they find.
98% of the internet is crap, yet LLM results don't reflect this dismal figure. They're amazingly good at distilling the 2% that is non-crap.
I know it was just an example, but actually no, the role of dietary saturated fat as a factor for heart disease remains very much valid. I’m not sure which recent studies you’re referring to, but you can't undo over 50 years of research on the subject so easily. What study were you thinking about?
Remarkable how easy it is to cling to propaganda
That said, for me this publication has red flags right from the start. Complaining about difficulty of changing everyone’s minds is a political and non-academic persuasion tactic that does not convince me. Calling it “resistance” and “bias” is a bullshit framing that makes me less likely to trust Teicholz. Of course there is resistance to 50 years of publication and research, and there should be. There’s a lot of bias towards the earth being round, and a lot of resistance to the idea that it’s flat, right? If I repeat the claim that the earth is round, is that “propaganda”? It would indeed take time and effort to change everyone’s minds about that.
Multiple times she references “>20” papers that back up her claims. Except 5 of her references in this paper are her own. And she has around 10 on this subject. So is she claiming this “new consensus” is based on what she herself and maybe one or two other people believe? If 50% of the evidence for consensus is her own papers, then I doubt there’s any consensus at all. It’s funny to claim there’s consensus at the same time she complains that it’s difficult to change the consensus. Even 20 independent scientific papers not authored by Teicholz is practically nothing in the big picture. It will take many more papers and much more time, and the evidence needs to be overwhelming, clear, obvious, and true.
She might be right! But Nina Teicholz is a journalist, not a scientist. She does have a PhD, but her publications don’t appear to be scientific research, and most look like opinion pieces.
Out of curiosity, if saturated fats aren’t the culprit, what is? Looks like she does have one paper questioning sugar, so is she claiming sugar is the real cause? What if it’s the combination of sugar and saturated fats? Does that make her right or wrong?
So you link to the equivalent of a reddit post 'summarizing' the space. I see this link every week on here and every time, the person who linked it thinks they just had a mic drop moment like you.
General disclaimers apply regarding portion sizes etc blah blah blah I'm not a doctor.
Anecdotally, a low-(ish) carb diet and fasting has done wonders for my health and many others. I will say that there appears to be a link with higher cholesterol when consuming higher amounts of fat, but the argument in nutrition science atm seems to be centered on whether or not that is "good" cholesterol, but it's hard to measure in human patients for a long time because you essentially need to put them on a very limited diet to get good data. Those large scale trials are expensive and hard to manage at scale.
[1] except when it isn't, of course
Sure you can if the research was bogus to begin with, sponsored in many cases, and merely taking for granted/referencing some previous results without verifying them, which is often the case.
I don't think this appropriately credits Google's power with regards to what you are seeing
> Who Gets To Decide What Is True?
I would not agree to this however:
> It was then up to the reader to decide which of those sites they wanted to visit,
With current search engines, Google decides for you ("helped" by SEO experts failing over themselves to rank higher because their revenue directly depends on it). In theory, you could go an read a few dozen pages and decide for yourself.
In reality, non-technical users will click on a first link that seems to be related to the question, and that's it.
Even with AI-based search (or q/a instead of search) I think the same will happen. There is and will be a huge reward for gaming the results, be they page links or RAG snippets that rank for a query. I've already seen many SEO shops advertising their strategies to keep the customers' business relevant in the chatbot area. As this approach becomes more prevalent, you can be sure many smart people will do many experiments to figure out how best to please the new algorithm.
In other words, AI-based search is an UX optimization, but doesn't address the core problem of how do you decide what content's the best, and do that in context of each user, and do that while maximizing the benefit for the user vs profit for the company doing this.
So we have two huge hurdles:
1. who will decide what the user wants[0] to see, and how are incentives for that entity aligned with the user's
2. how is that entity supposed to find the information needle in the haystack of slop that's 90%+ of current web?
[0] "wants" in a rational "give me the best possible information" meaning, not in "what keeps them addicted, their heart rate up, and what will drive engagement" meaning
"Is it legal" is very different from "Can they."
By "local," I mean municipal and below. I didn't mean "federal elections conducted in my locality." Election security kicks for state, federal, and some municipal elections.
Others are intentionally fraudulent (e.g. local corruption) or unintentionally broken (e.g. using Google Forms for a school-level public body, where people not legally qualified to vote might still do it, unaware they're committing a felony).
And "public body" has a specific meaning under my state law which extends the same laws as e.g. cutting for my state senate. That's bodies like local school boards, but not random school clubs.
That's the level where we have massive illegal voting where I live.
It is very similar. Google decides what to present to you on the front page. I'm sure there are metrics on how few people get past the front page. Heck, isn't this just Google Search's business model? Determining what you see (i.e. what is "true") via ads?
In much the same way that the Councils of Carthage chose to omit the acts of Paul and Thecla in the New Testament, all modern technology providers have some say in what is presented to the global information network, more or less manipulating what we all perceive to be true.
Recently advancements have just made this problem much more apparent to us. But look at history and see how few women priests there are in various Christian churches and you'll notice even a small omission can have broad impacts to society.
Seems particularly to be a US based phenomenon. Unlike the more transparent manipulation seen under dictatorships... Where people generally recognize propaganda for what it is, even if they can’t openly challenge it—some in USA live within entirely new realities.
You would think common sense prevails...
Appeals to common sense seem to underlie many of the "alternative facts" in the zeitgeist. "Common sense" is how people defend their views when they can't do it empirically. That's not always bad, but "common sense" is probably part of the problem.
Remember the covid years, when what was true kept changing rapidly, and sometimes what was said on the fringe and considered misinformation was later adopted by mainstream. Vaccines prevent transmission of the virus. Oh no they don't. Don't wear masks. Oh no, do wear masks. Oh, whatever, cloth masks are face decorations anyway. Lockdowns are good. Oh, lockdowns were a mistake. Don't treat pneumonia patients with prednisone. Oh, do treat pneumonia patients with prednisone. Lab leak is a conspiracy theory. Oh, maybe lab leak is not a conspiracy theory. And so on, and so forth...
Or take a look at how not just the media, but even the government in the UK or Canada liberally put labels such as "far right", "alt right", "antisemitic", etc. on their opponents, and how these labels pop up in Wikipedia. Are they true?
And then Quine points out that the definitions of "bachelor" and "married" are themselves contingent on outside factors.
"Can illegal immigrants vote?", while close to being an analytic proposition, still depends on an empirical approach that can never be mediated by text, video, etc. All propositions are by necessity experiential. Nullius in verba!
So the truth is and has always been what happens when you get off your butt and go out and test the world.
This is not to say that we don't benefit from language. It makes for a great recipe. If you follow the instructions to bake a cake and you get what you expected you know that the recipe was true. The same goes for the laws of science, search engine results, and generative AI.
The way this is wrought, in the novel, is a savant engineer writes a bot framework that can cheaply and quickly disseminate torrents of misinformation about a provided subject, and then open sources this framework. He basically broke the internet on purpose as a sort of accelerationist move I suppose.
We have this idea today that everyone online is getting trapped in echo chambers, but that’s been the case for most of human history.
(now to wiping the tears of laughter from my face...)
When people mostly communicated with those in their own town that is a bubble. Radio and TV is more of a bubble because there is limited "bandwith" in the scheduling so it has to be editorialised (not necessarily a bad thing).
Social media companies do choose what you see via algorithms, but I'm not convinced they benefit from only showing you content you "agree" with, it feels like being shown a certain amount of content outside your "bubble" would increase screen time. There are also the comment sections that often have contradictory views.
Even social media (setting aside the rest of the internet for the moment) will expose you to more viewpoints than the the social circle in your home town, or a TV/Radio schedule. I'm not saying it is a healthy way to be exposed to other viewpoints, but I don't think the problem with social media is that is creates a bubble.
legally no, practically- yes. In most states, you simply must attest that you are a citizen in order to register. In many states, non-citizens have been auto-registered to vote when attaining drivers licenses. Reddit is full of panicked immigrants concerned that they found themselves registered to vote, and worried about how that would affect their status.
Outsourced workers in less expensive places of the world providing human feedback.
Well, outside of matters of stark factuality (what time does the library close?, what did MSFT close at?), many things people may be "searching for" (i.e. trying to find information about) are more in the realm of informed opinion and summary where there is no right or wrong, just a bunch of viewpoints, some probably better informed than others.
I think this is the value of traditional search where the result is a link to a specific page/source whose trustworthiness or degree of authority you can then judge (e.g. based on known reputation).
For AI generated/summarized "search results", such as those from "ChatGPT Search" (awkward name - bit like a spork), the trustworthiness or degree of authority is only as good as the AI (not person) that generated it, and given today's state of technology where the "AI" is just an LLM (prone to hallucination, limited reasoning, etc), this is obviously a bit of an issue...
Even in the future, when presumably human level AGI will have made LLMs obsolete, I think it'll still be useful to differentiate search from RAG AGI search/chat, since it'll still be useful to know the source. At that point the specific AGI might be regarded as a specific person, with it's own areas of expertise and biases.
The name "ChatGPT Search" is very awkward - presumably they are trying to position this as a potential Google competitor and revenue generator (when the inevitable advertisements come), but at the end of the day it's just RAG.
> Who Gets To Decide What Is True?
For any given statement, the answer up until a couple years ago was, "the speaker". Speakers get to decide what to say, but they're also responsible for what they say. But now with LLMs we have plausible text without a speaker.
I think we have a number of historical models for that. A relevant one is divination. If you bring your question to the haruspex, your answer is read out of the guts of a sacrificed animal. If the answer is wrong, who do you blame? The traditional answer is the gods, or perhaps nobody.
Bu we know now that fortune tellers are just selling answers while pretending to not be responsible for them. Which points us at one solution: anybody selling or presenting LLM output as meaningful is legally responsible for the quality of the product.
Unfortunately, another model is the modern corporation. Sometimes the people in a company intentionally lie. More often, statements are made by one person based on a vision or optimism or confusion or bullshit. Nobody set out to lie, but nobody really cared about the truth, at least not as much as everybody cared about making money.
So I'd agree that the government doesn't have much role in deciding The Truth. Similarly, the government shouldn't have much role in controlling what you eat. But in both cases, I think there's plenty of role for the government in ensuring that companies selling good food or good information have sound production and quality control measures to ensure that they are delivering what consumers are expecting.
Well, it's still a known source, who's competence, biases, etc one can judge just like any human source.
I get this might seem like tedious nitpicking to you, but the number one error people are making with LLM output is anthropomorphizing it. Which I get, because it's built to seem that way. But it's an enormously dangerous misconception.
As one example among zillions, look at the term "hallucination". All LLM output is equally "hallucinated". Some of it, when interpreted by a human may be taken as meaningful. Some of it, the "hallucinated" part, is taken as meaningful but contrary to something else they understand. But it's the human creating all the meaning here. Even calling this class of mismatches "hallucination" is anthropomorphizing LLMs.
Imagine I take the proverbial million monkeys to generate random words. Then I create a statistical filter so that we extract only the plausible sentences. Is this machine a source? Can I "judge just like any human source" here?
I'd say the answer is a clear no. And if you think the answer is yes, then the same has to apply to things like horoscopes, the I Ching, or the intestines of a sacrificial goat.
If I've observed the answers of the oracle to be biased, or useful/correct in some circumstances and not in others, then this is something to take into account when deciding whether or not this is a useful source to pay attention to.
From the POV of whether the output of the black box is useful, it doesn't make any difference whether what's in the box is a human, an LLM, or a rabid monkey. It is what it is, and I'll judge it on it's merits.
Now, YOU may care what's in the box, for some reason, but that's on you.
There are few more restrictions here: https://www.usa.gov/who-can-vote
Taxation without representation and all that.
> in the USA, can illegal immigrants vote According Author X in Book Y which studied this topic in depth: foo answer
Independent Tribunal: https://www.independenttribunal.org/ (a project of mine)
Even in the of law there are various schenanigans and loopholes such as "legally true" :)
This doesn’t have a single truthful answer. Some states don’t have voter ID laws, so the truth can depend on the state. In those no voter ID laws there’s not that much keeping someone from voting twice or more under different names, except significant moral qualms about subverting instead of preserving everyone else’s right to vote in a democratic republic. Someone can assume the name of a person from another country that could’ve plausibly come in illegally. Without a picture ID, they can’t claim you aren’t that persons.
Can an illegal immigrant vote? Yes, in states without voter ID laws, technically anyone can vote, even convicted terrorists. Should an illegal immigrant vote? No, they’re not supposed to be able to vote and there may be consequences if caught.
What purpose do a lack of voter ID laws serve except the obvious conclusion which is to enable cheating?
MAGA types who wanted to prove democrats could cheat like this were caught very quickly because voting twice or voting for someone else is very very easy to detect, even without IDs. The judges gave them Darwin awards.
Unregistered voters can’t vote in any state, they can only do so by pretending to be someone else (name and address matches the rolls), but they are caught when those people actually vote. Foreigners can also vote by illegally registering. But motor voter means they actually check your status at registration.
The ID thing is a solution to a non-problem, like literacy tests were. But the republicans could make it easier to achieve universally if they just went with a guaranteed free national ID like other countries, except that would make the obstruction aspect of voter ID requirements much more moot, so they never go there.
Foreign residents can vote in some local election in a few places, mainly very local school board elections.
This is flawed thinking to get to the conclusion of “reliable answers”. What people want to see and the truth are not overlapping.
Consider the answers for something like “how many animals were on Noah’s Ark” and “did Jesus turn water into wine” for examples that cannot be solved by trying get advertisers out of the loop.
Isn’t this already that? A new business model? Something like OpenAI’s search or Perplexity can run on its own index and not be influenced by Google’s ranking, ads, etc.
In areas where there is a simple objective truth, like finding the offset for the wheels on a 2008 BMW M3, we have had this capability for some time with Perplexity. The LLMs successfully cuts through the sea of SEO/SEM and forum nonsense and delivers the answer.
In areas where the truth is more subjective, like what is the best biscuit restaurant in downtown Nashville, the system could easily learn your preferences and deliver info suited to your biases.
In areas where “the science” is debated, the LLM can show both sides.
I think this is the beginning of the new model.
In that way, search captures value directly on its core function: efficiently facilitating the creation and dispersal of knowledge.
This may end up occurring, but based on how unlikely it sounds, it’s my reflection that the web search that we have today is simply a small component of what will eventually become a system that closely approximates the above: basically a kind of global intelligent interconnected agent for the advancement of humanity.
Rather than the great unbundling, this will be the great rebundling … of many diverse functions into a single interface seamlessly.
What a great analogy
What if I am looking for a medical page or a technical page or whatever else where I need to see, read, experience actual content and not some AI summary?
As someone who has been using google for search basically constantly since 1995, I've switched probably 90%+ of what normally would have been google searches over to Perplexity (which gives me references to web pages alongside answers to my questions, to review source materials) and ChatGPT (for more just answers I can verify without source). The remaining searches have gone to Kagi.
On the one hand this has got to be hurting Google search ad revenue. On the other hand I don't know if I ever clicked an advertised link. On the other other hand, not having to wade through SEO results has been so nice.
Guess why it failed? It was largely used as a way to trick search engines. Same reason as your vision of the perfectly honest and correct search engine and look chatbot will never be perfect. It's because people lie to search engines to spam and get traffic they don't deserve. The whole history of search is Google and others dealing with spam. Same goes for email. Google largely defeating spam made them the kings of the email world.
Everyone will need their own personal spam filter for the world for everything ones the artificial super intelligences fill the whole world with spam, scams and just plain old social engineering propaganda because we will be like helpless four year old children in the world of AI super intelligence without our AI parents to look out for us.
Your vision of the god system determining what is truth is like saying there will be a single source of truth for what is and is not a spam email. Not going to scale and not going to be perfect, but good enough with AI and technology. Really hope there's an opt-out though since Google had memoryholed most of the Internet.
I imagine this happens with LLMs and everything else over time. Along with the ads and pay for placement.
(1) Search is already heavily AI driven, and Google is clearly going in that direction. Gemini is currently separate, but they'll blend it in with search over time, and no doubt search already uses LLM for some tasks under the hood. So ChatGPT search is an evolution on the current model rather than a big step in a new direction. The main benefit is you can ask the search questions to refine or followup.
(2) Aside from the economic incentives faced by search engines, there is the fact that algorithms are tuned toward a central tendency. The central tendency of question askers will always be much less informed than the most informed extreme. Google was much better when the average user was technical. The need to capture mobile searches is one force that made it return on average worse results. Similarly if Kagi has a quality advantage now, we need to be realistic about how much of that quality is driven by its users being more technical.
(3) I think micropayment schemes have generally asked several orders of magnitude more for a page view than users are willing to pay. As long as content creators value their content much more highly than consumers do, they'll stick with advertising which lets them overcharge and gives consumers less of an option to say no to the content.
But this will never happen with mainstream search imo. It is not a technical problem but a human one. As long as there is a human in control of what gets surfaced, it is only a matter of time until you revert to tampered search. Humans are not robots. They have emotions and can be swayed with or without their awareness. And this is a form of power for the swayer as much as oil or water are.
The idea that you can have an AI system provide factual and reliable answers to human centric questions is as real as Star Trek itself.
You will never remove the human factor from AI
Your hope might be that a technical solution is found for a human problem but that is unlikely.
LLMs have already fundamentally changed our relationship to information.
Counterpoint: with a chain-of-thought process running atop search, you can potentially avoid much of the meta-search / epistemic hygiene work currently required. If your “search” verb actually reads the top-100 results, runs analyses for a suite of cognitive biases such as partisanship, and gives you error bars / warnings on claims that are uncertain, the quality could be dramatically improved.
There are already custom retrieval/evaluation systems doing this, it’s only a matter of a year or two before it’s commoditized.
The concern is around OpenAI monetization, do they eventually start offering paid ads? This could be fine if the unpaid results are great, a big part of why the web is perceived to be declining is link-spam that Google doesn’t count as an ad.
My prediction would be that there is a more subtle monetization channel; companies that can afford to RAG their products well and share those indexes with AI search providers will get better results. RAG APIs will be the new SEO.
Seeing search enter the space is something that I feel has been seriously needed, as I've slowly replaced Google with ChatGPT. I want quick, terse answers sometimes, not a long conversation, and have hundreds of tiny chats. But there's something scary seeing the results laid out the way they are, which leads me to believe they may be closer to experimenting with ad-based business models, which I could never go back to.
Before we had the internet, how did you answer questions? You either looked it up in a book, where you then had to make a judgment on whether you trusted the book, or you asked a trusted person, who could give you confidently wrong answers (your parents weren't right every time, were they? :) ).
I think the main difference here is that now anyone can publish, whereas before to make a book exist required the buy in of multiple people (of course, there were always leaflets).
The main difference now is distribution. But you still, as a consumer of information, have to vet your sources.
One thing is clear, in the vast majority of cases we don't have a single truth, but answers of different levels of trustworthiness: the law, the government, books, Wikipedia, websites, and so on.
A human would proceed differently depending on the context and the source of the information. For example, legal questions are best answered by the law or government agencies, then by reputable law firms. Opinions of even highly respected legal professionals are clearly less reliable than government law itself and are likely to be a point of contention/litigation.
Questions about other facts, such as the diameter of the earth or the population of a country, are best answered by recent data from books, official statistics and Wikipedia. And so on and so forth.
If we are not sure what the correct answer is, a human would give more information about the context, the sources, and the doubts. There are obviously questions that cannot be answered immediately! (if it's to easy to find the truth, we would not need a legal system or even science!) So no machine and no amount of computation can reliably answer all questions. Web search does not answer a question. It's just trying to surface relevant websites to a bunch of keywords. The answering part is left as an exercise for the user ;)
So an AI search with a pretense of understanding human languages makes the task incredibly harder. To really give a human-quality answer, the AI not only needs to understand the context, but it should also be able to reason, have common sense, and be a bit self-aware (I'm not sure...). All this is beyond the capabilities of the current generation of AI. Therefore, my conclusion is that the "search" or better said the "answer" cannot be solved by LLM, no matter how much they fine-tune and tweak it.
But we humans are adaptable. We will find our way around and accept things as they are. Until next time.
Yes, Google has their own AI divisions, tons of money and SEO is to blame for part of their crappiness. But they've also _explicitly_ focused on ad-dollars over algorithmic purity if one is to believe the reports of their internal politics and if those are true they have probably lost a ton of people who they'd need right now to turn the ship around quickly.
So if 95% of traffic/users/whatever metric are not using a web browser for those activities, is it really the web? It can't be called the web just 'cause they use HTTPS. It's gotta be a 'world wide web' experience, which I think a good proxy for would be using a web browser.
I got no horse in this race, just thinking out loud about it.
Agree YouTube and Instagram are probably mostly apps which puts them in the “Internet” category but not “world wide web”.
Another common phenomenon these days is that lots of businesses don’t even bother having a web presence - it’s all instagram, WhatsApp and tiktok accounts, mostly only accessible via apps (or worse, chat platforms like discord)
There are several meaningful difference between surfing Youtube and surfing the web. These include ownership, access, review, exposure, and more.
Honestly the web died long ago imo. Wikipedia and other wikis are the only places that feel like the old web to me now.
Google (and Facebook, a few other platforms) made it so that the vast majority of websites are never visited. ChatGPT further erodes the possibility of tying economic value to the production of that value.
Seems like we need a new framework around "intellectual property"
I just got back from a job where my tech used AI on his phone every time he needed to search for something. The results were hilariously bad, but if things keep improving one day they might not be. Google needs to be ready for that.
Do you?
> not the same sort of risky move
I don't see it as materially different. I also don't consider this to be an example of either of them being genius visionaries. I actually think it's an extremely straightforward tactic to cannibalize part of your business to break into another bigger market.
Apple didn't stop selling iPods, you know? I don't really get why it was more risky for Apple to market the iPhone than for Facebook to decide not to shut down Instagram once they acquired it.
The differences in the literal creation of the entity are immaterial to my point.
You don't see the difference between creating and buying?
> Apple didn't stop selling iPods, you know?
Read The Innovator's Dilemma. Apple launching the iPhone would be equivalent to Google, today, replacing a list of search hits with, at most, three answers.
E.g. for me, how much Google (and silicon valley in general) have enabled twisted ideologies to flourish. All in search of ad-dollars by virtue of eyeballs on screens, at the detriment of everything.
Considering the value of time, past consumer surplus is especially valuable now.
Sure, there are systematic flaws causing SEO to ruin the information provided: but it isn't clear what Google can do to fight the emergent system.
I'm not sure that Bing/DDG are any better.
I use search (DDG web, Google/Apple maps, YouTube) all the time and I am regularly given results that are extremely valuable to me (and mostly only directly cost me a small amount of my time some of my time e.g. YouTube adverts). Blaming SEO on Google seems thoughtless to me. Google appears to be the victims of human cybersystems as much as we are.
I'm not familiar with what you're referring to here. Happen to have a link?
With ChatGPT, I can give a thumbs up or thumbs down; this means that OpenAI will optimize for users thumbs up.
With Google, the feedback is if I click on an Ad; this means that Google optimizes for clickbait.
I am talking about relevance or returning what I asked. If I ask for reviews for SaaS product, Google will usually return a rival vendors’ biased review.
If ChatGpt search returns a review written by “the professional association of xxx developers” or another unbiased site, I will give it a thumbs up. I believe other people will do the same.
Heck, Google even promoted the `ping`[0] anchor attribute feature so they can log what link you click without slowing you down. (Firefox doesn’t support ping, which means when Firefox users click on a Google search result link they’re sent to an internal google.com URL first and then redirected for logging purposes)
[0] https://developer.mozilla.org/en-US/docs/Web/API/HTMLAnchorE...
What the parent is referring to is favoring annoying ad-filled garbage over an equally relevant but straightforward result.
The hidden variable is that ad-riddled spam sites also invest in SEO, which is why they rank higher. I am not aware of any evidence that Google is using number of Google ads as a ranking factor directly. But I would push back and say that “SEO” is something Google should be doing, not websites, and a properly optimized search engine would be penalizing obvious garbage.
I can still search things, i get results but, they're an ordered list of popular places the engine is directing me to. Some kind of filtering is occurring on nearly every search i make that's making the results feel entirely useless.
Image search stopped working sometime ago and now it just runs an AI filter on whatever image you search for, tells you there's a man in the picture and gives up.
Youtube recommendations is always hundreds of videos i've watched already, with maybe 1-2 recommendations to new channels when i know there's millions of content creators out there struggling who it will never introduce me to. What happened to the rabbit holes of crazy youtube stuff you could go down?
This product is a shell of its old self, why did it stop working?
The level of sponsored results for some queries is way OTT, and obviously any kind of search like "best laptop 2024" is never going to give you good results (probably because they don't exist), but other than that I'm still pretty happy with Google Search.
Genuinely interested: have you tried to spend a few weeks on an alternative?
I decided to try DuckDuckGo a few years ago. Not because it was obviously better, but to see if I could get used to it. After a few weeks, I had completely stopped falling back to Google when not finding what I wanted. I stayed on DDG for a couple years. Then same thing with Kagi: I just decided to try. It's been 1.5-2 years now and I'm disappointed when I can't use Kagi (which has my customizations, like some websites I ignore and some that I pin).
I guess my point is that it's not necessarily that you have to try something else when Google is unbearable. Maybe you can try something else and then realize (or not) that Google was not better.
I try to use OpenStreetMap as much as I can (I have a deGoogled smartphone and OpenStreetMap works well enough) but it is true that Google Maps is better (at the cost of privacy of course).
But in terms of search... I can't remember of a time where I tried Google because I couldn't find with Kagi and ended up finding something with Google. On the contrary with the Kagi lenses it's often a lot easier to get specific results.
1. Search Google for "ukrainian who shot his commanding officer" without quotes
2. Google serves me nothing but MSM articles of Russian this or that. The word Russian wasn't even in my search string.
3. Add the Google operator MINUS SIGN Russia
4. Results:
a) Policeman feared Chris Kaba would kill, court told
b) Media: Russian Repeated Offender Kills Five More His ...
c) President Volodymyr Zelenskyy and First Lady Olena ...
d) Ukrainian Galician Army
e) Article from 2017 entitled Killed Defense Intelligence Officer Was "The First Donetsk Cyborg"
f) Shots fired at car carrying Ukrainian President Zelenskiy's ...
5. Go to yandex.com and search the original query6. It comes up on Yandex immediately with the original query
And its first two results are for a dead site and a deleted article, so fine user-experience reasons to exclude?
And the google results have a link to a reddit story such accident on the first page
YouTube search is also completely useless now.
https://www.youtube.com/results?search_query=path+of+exile+2...
https://www.youtube.com/results?search_query=ramen+restauran...
When I search these three things the first 5 results are exactly what I want.
What exactly are you searching of the time that it's "completely useless"? Genuine question.
Try searching for a particular video, one that is not super popular. What I want is a complete list of results that match my query. What I get is YouTube trying to recommend videos to me.
If you try to describe a non-popular video it just becomes a crapshoot on what to give you if there's no word/tag watching. You'd need some hint of the channel name or something. The volume of low viewership videos is incredibly high.
Can you give me a real example of something you've tried to do?
This is unfortunately not true. I have a little channel and there have been times when searching for the exact title of some of my videos did not return it in the results at all (searching with quotes or not). Cannot reproduce now because the search algorithm has now started liking me.
For example, searching for climbing comp videos and getting a completely unrelated video about some new tech gadget released within the last couple days from a random popular content creators makes no sense.
Clearly, it works for Google (content creators intentionally make click-baity thumbnails and titles because Google encourages it), but it's user hostile: it's designed to suck you into a vortex, which is not what the user was intending in the first place.
That said, all content platforms do this right now, so my intention isn't singling out Google. It's frustrating nonetheless.
Now that the JD Vance podcast is out, that’s the first result; then it’s all MSM. Wow.
If there are no sponsored links - the result is crap.
Google is good at searching, they just have no incentive to show you results.
“ to organize the world's information and make it universally accessible and useful” was a nice mission while it lasted
Yes. However, I found that https://scholar.google.com still works perfectly well. It feels just as the old Google without all the crap they've been adding in the last years.
I can’t imagine the cost this would have on scientific producivity in the West.
Oddly their biggest strength is being irrelevant to the decision makers, if the bean counters noticed the few million they are losing on running Scholar there will be ads + Gemini all over it.
right now, I download pdf and upload it to chatgpt to bounce ideas.
Even if you give it a paper directly I’d not believe it to be reliable. Maybe it could help search for papers, but that’s it.
First search (“products that do X”) got me a bunch of those comparisons sites, none of them containing the one I was trying to find
Second search (“ycombinator startup that does X”) got me a page of spam, but at least I found the product name
Third search (company name) got me an ENTIRE PAGE of ads and SEO optimized pages before the actual link to the actual product
If I search "Mullica Hill tornado" on yt itself, I get nothing but useless 1 min local news clips. If I search the same term on Reddit, I get first person footage of the tornado passing over people's houses—hosted on Youtube! Tornado enthusiasts still occasionally dredge up "lost media" of events like the 2011 Alabama outbreak that have been on the site this entire time, but are effectively impossible to view via the algorithm, even with the precise date and location specified.
Image search isn't great either but it still often gives me something close and that usually satisfies my image searching needs.
I still find YouTube recommendations quite good for me, but there are occasional ones I've watched already. I still go down its fun (and educational!) rabbit holes all the time.
However when they don't, it's invariably to push some alt-right slop down my throat. Video is about a comedian? Suggestion "feminist woke takedown compilation". Video about news? Suggestion "$european_far_right_party's channel says gypsies are subhuman". Video about economics? Suggestion is Jordan Peterson ranting about something. And so on and so on. It's pretty tiring.
This may make it better for the non-tech folks who search for things in unclear language, and likely make it worse for those who search with precision (i.e. much of the HN crowd).
Most importantly, they make their money off ads, and it probably makes sense to optimize for the non-tech folks. The ones that don't run ad blockers and accidentally click the barely differentiable ads.
In short - I suspect they're just using new tech to make more money.
It doesn't feel magical, and it particularly doesn't feel like a gateway to new and interesting places on the internet, like it used to.
But I still make hundreds of searches on google every day, so obviously it is working.
And on rare occasions now when I do search on Google I feel so happy I switched to DDG. To an extent that it does feel to me that Google has indeed stopped working completely for me. And yes I’ve stopped their search history, recommendations etc on YouTube.
Part of the answer is here - https://www.wheresyoured.at/the-men-who-killed-google/
Yes it is hard to find some stuff in internet because it is filled with generated affiliate spam and walled gardens, Has Google stopped working for me? Nope It seems still alive and kicking.
Even if I type in 2-3 lines worth of nearly-exact lyrics that show up on multiple lyrics websites, it'll give me completely unrelated songs that match several words at most.
On that note, if anyone needs a good browser-based song finder, the "Aha Music Identifier" extension is pretty good. It's a life saver when watching Twitch streams that don't list the currently-playing song.
On the other hand, chatgpt 4o-mini and 3.5 will just make up source material which is amusing but not very helpful.
> Based on the elements you’ve described—eyes, a pineapple, a bucket of water, a house of cards, chess, and a time loop—it’s challenging to identify a single music video that encompasses all these features.
Has google ever indexed all the lyrics and scenes in a video to allow for such a weird search to be successful?
The very earliest form of page rank used a form of ML.
The question that remains unanswered is how google can do so without compromising customer revenue.
This especially happens after they dominate the market.
Take for example IE6, Intel, Facebook, IBM, and now Google.
They have everything they need to keep things from going off the rails, management however has a tendency to delusionally assume their ship is so unsinkable that they're not even manning their stations.
It becomes Clayton Christensenesque - they're dismissive of the competition as not real threats and don't realize their cash-cow is running on only fumes and inertia until its too late.
I’m not sure Facebook fits in considering they at least managed to get some other products along the way, and may get more.
I certainly don’t think Google fits the bill. Google is failing because they let their cash cow ruin everything else, not because they let it stagnate while they chased the next moonshot. Google Cloud could have easily been competitive with AWS and Azure in European Enterprise, but it’s not even considered an option because Google Advertising wouldn’t let it exist without data harvesting. Google had Office365 long before Microsoft took every organisation online. But Google failed to sell it because… well…
It’s very typical MBA though. Google has killed profitable products because those products weren’t growing enough. A silly metric really, but one which isn’t surprising when your CEO is a former McKinsey.
It couldn’t happen to a nicer company though, and at least it won’t kill people unlike Boeing.
* Boeing is a consequence of the "Jack Welch" effect - gutting the core in service of short term gains for stock-holders.
* The MBA type, typified by John Sculley at Apple is about calcifying the current offerings presuming market segments supported by historicals with predictable demand. This works well for defensives such as utilities, consumer products and health care but not for markets with dynamic consumer relationships such as technology.
* The Google Cloud example is the Xerox Parc phenomena. Xerox was organizationally structured for investment payoffs only characteristically similar to their mainline products thus they couldn't properly allocate resources to things, such as desktop computing, with different kinds of curves. This is similar to how the franchise retailer Blockbuster so slowly responded to the centralized mail-order subscription Netflix. The institutional structure is only-so-flexible. This is similar to Conway's Law.
* The "ruin everything else" is a generalized form of a "brand extension failure". Examples include Harley Davidson perfume, Bic underwear, McDonalds Pizza, and Heinz cleaning vinegar - an over-leveraged commitment to a wildly successful core offering makes other ventures impossible.
This is not that. It's yet something else. Abstractly it's "X is a wild success, let's make Y another X instead of working on X+1"
Organizations suffer from varying degrees of ailments and they can create codependencies making the unraveling hard. Often it devolves into politics of power brokers with the company's survival dependent on the competency of the influential instead of the influence of the competent. A brutal struggle to control a sinking ship.
The crisis of the third century happens every day.
can you name these products?
And now their VR products.
It is interesting to go back a decade and read the articles about Google's moonshots from then. There was even more "we're building the future!" hype than what Elon Musk or Sam Altman are currently pushing.
I know this is a predominantly white, upper middle class site (read: progressive), and people from that demo sure love to point out that "you can't say the leg is broken just because the foot is dangling at 90 degrees, you aren't a doctor!" ..... but at some point you surely have to start to trust your own eyes and use critical thinking, right??
It's not a conspiracy theory Google ruined their search to serve more ads. It's so obvious and happening for such a long period of time, anyone above certain age and with decent memory can see it. You don't have to "believe the reports", you don't even have to be smart, just use your eyes and your memory.
I think they prioritized large and fast incomes over long, steady incomes.
Also, there isn't just money that taints Google search results. It's also censorship and political motivation, which is mostly obvious in image search.
I'm wondering if they've seen the writing on the wall for a long time, and if diversification is going to save the day. I don't know that I've ever clicked on an ad in Google search results, but I happily pay them for a family account for Youtube and 2TB of drive storage, plus, you know, a flagship phone every couple years. :-)
Yes and no. Yes for the quality of search results. Google algorithm and user experience was simply better than AltaVista's, but Google had another advantage. It used a cluster of cheap consumer-grade hardware (PCs) as its backend, rather than the expensive servers AltaVista used, in fact, AltaVista started off as a way for DEC to show off its hardware.
As a result Google was not only better than its competitors at searching, it was also cheaper and scaled better, and this architecture became the new standard.
It is the opposite for these AI-based systems. AI is expensive, it uses a lot of energy, and the hardware is definitely not consumer-grade.
What it means is that barring a significant breakthrough, cheap ChatGPT-like services are simply unsustainable. Google don't have to do anything, it will collapse on its own. Ironically, probably in the same way that Google result became crappier by the year, but on fast forward.
This is the basic premise of Ed Zitron's article https://www.wheresyoured.at/to-serve-altman/ and others written around the same time. A lot of what he's written seems to be coming to pass, e.g. unfathomably large investments (~$100 billion) by tech giants in OpenAI and other AI startups to keep them going. That does seem unsustainable.
Can anyone make a counterpoint to this claim? I'd be very interested to hear what a viable path to profitability looks like for OpenAI.
I think it’s great that competition is driving both companies to improve, but I’m not seeing anything about this that screams “Google-killer”
The problem I see with search is that the input is deeply hostile to what the consumers of search want. If the LLM's are particularly tuned to try and filter out that hostility, maybe I can see this going somewhere, but I suspect that just starts another arms race that the garbage producers are likely to win.
There is no way to SEO the entire corpus of human knowledge. ChatGPT is very good for gleaning facts that are hard to surface in today's garbage search engines.
Currently I do find that Perplexity works substantially better then Google for finding what I need, but it remains to be seen if they're able to stay useful as a larger and larger portion of online content just AI generated garbage.
This comes off as condescending. As things have gotten more algorithmic over the last two decades, I've noticed a matching decrease in the accuracy and relevance of the information I seek from the systems I interact with that employ these algorithms.
Yes, you're right that there are processing algorithms behind the scenes interpreting the data for us. But you're wrong: I fucking hate it, it's made things worse, and layering more on top will not make things any better.
Wasn't google AI surfacing results about making pizza with glue and eating rocks? how is that not harmful garbage?
Why would you assume that?
We don’t have a way of finding objective information, why would we be able to train a model to do so?
This time it was, "Did Paul Edwin Zimmer write a fourth Dark Border novel?" (Real answer: Yes, Ingulf the Mad. You can find the answer on his Wikipedia page.[1])
ChatGPT's[2] answer: "Yes, Paul Edwin Zimmer wrote a fourth novel in the Dark Border series titled "The Dark Border." This book was published after the original trilogy, which included "The Dark Border," "The Gilded Age," and "The Silver Sphere." If you're interested in the themes or plot, let me know!" (Note: these are not the titles of the 2nd and 3rd novels in the series. Also, it gave me the same name for the putative 1st and 4th books.)
Pure hallucination.
1. https://en.wikipedia.org/wiki/Paul_Edwin_Zimmer 2. https://chatgpt.com/
No, Paul Edwin Zimmer did not write a fourth novel in the Dark Border series. The trilogy consists of "The Dark Border," "The Dark Border: The Return," and "The Dark Border: The Reckoning." After these, he focused on other projects and did not continue the series.
> As of my knowledge cutoff in October 2023, Paul Edwin Zimmer did not publish a fourth novel in the Dark Border series. The series comprises three books: > 1. The Lost Prince (1982) > 2. King Chondos' Ride (1982) > 3. A Gathering of Heroes (1987) > > Paul Edwin Zimmer had plans to continue the series, but he passed away in 1997 before any additional novels were completed or published. There have been no posthumous releases of a fourth Dark Border novel. If there have been developments after October 2023, I recommend checking recent publications or official announcements for the most up-to-date information.
Yes, Paul Edwin Zimmer wrote a fourth novel in his Dark Border series titled Ingulf the Mad, published in 1989. This installment focuses on the characters Ingulf Mac Fingold and Carrol Mac Lir, detailing their meeting and the acquisition of their mystical swords. Notably, Istvan Divega, the protagonist of the earlier books, does not appear in this novel.
DuckAssist Result: The fourth novel in Paul Edwin Zimmer's Dark Border series is titled "Ingulf the Mad." This book focuses on characters Ingulf Mac Fingold and Carrol Mac Lir, detailing their meeting and the acquisition of their mystic swords, while the main character from the earlier novels, Istvan Divega, does not appear.
(With source as wikipedia)
I had to refresh my knowledge base visiting fandom websites to review the episode selected as answer as chatgpt tendency to mix things up and provide entirely made up episodes make it hard for bar trivias (and makes me doubt myself too). The same with other tv series such as House MD and Scrubs.
I really think this latest release is a game changer for ChatGPT since it seems much more likely to return genuine information than ChatGPT answering using its model alone. Of course it still hallucinates sometimes (I asked about searching tabs in Firefox Mobile and it told him the wrong place to find that ability while citing a bunch of Mozilla help docs), but it's much easier to verify that by clicking through to sources directly.
It feels like a very different experience using ChatGPT with search turned on and the "Citations" right side bar left open. I get answers from ChatGPT while also seeing a bunch of possibly relevant links populate. If I detect something's off I can click on a source and read the details directly. It's a huge improvement on relying on the model alone.
You can use them as a starting point though.
---------------
4o:
Yes, Paul Edwin Zimmer wrote a fourth novel in his Dark Border series titled Ingulf the Mad, published in 1989. This installment focuses on the characters Ingulf Mac Fingold and Carrol Mac Lir, detailing their meeting and acquisition of mystical swords. Notably, Istvan Divega, the protagonist of the earlier books, does not appear in this novel.
---------------
4o-mini:
Yes, Paul Edwin Zimmer authored a fourth novel in his Dark Border series titled Ingulf the Mad. Published in 1989, this book shifts focus from the previous protagonist, Istvan DiVega, to explore the backstory of Ingulf Mac Fingold and Carrol Mac Lir, detailing their initial meeting and the acquisition of their mystical swords.
The complete Dark Border series consists of four novels:
1. The Lost Prince (1982)
2. King Chondos' Ride (1983)
3. A Gathering of Heroes (1987)
4. Ingulf the Mad (1989)
These works delve into a world where the Hasturs' power is crucial in containing dark creatures, and the narrative unfolds through various characters' perspectives.
4o-mini: https://chatgpt.com/share/6725f6fb-3b84-800b-9b4d-3e82554fbf...
I can't figure out how to share the 4o one. The UI is weird.
I also search on Google and find "hallucinations": pages with false content that rank better than the real information. In Google the current issue is about ranking and forgetting information that was previously available.
LLMs are run with variable output, sure, but it's particularly odd if GP used the search product as it doesn't have to provide the facts from the model itself in that case. If GP had posted the link to the actual chat rather than provided a link to chatgpt.com (???) I'd be interested in seeing if Search was even used as that'd at least explain where such variance in output came from. Instead we're all talking about what could have happened or not.
I'm on the other side of this fence to you. I agree that the conclusion here is that "it is flaky." Disagree about what that means.
As LLMs progress, 10% accuracy becomes 50% accuracy. That becomes 80% accuracy and from there to usable accuracy... whatever that is per case. Not every "better than random" seed grows into high accuracy features, but many do. It's never clear where accuracy ceilings are, and high reliability applications may be distant but... sufficient accuracy is not necessarily very high for many applications.
Meanwhile, the "accurate-for-me" fix is usually to use the appropriate model, prompt or such. Well... these are exactly the kind of optimizations that can be implemented in a UI like "LLM search."
I'm expecting "LLMs eat search." They don't have to "solve truth." They just have to be better and faster than search, with fewer ads.
I usually seed conversations with several fact-finding prompts before asking the real question I am after. It populates the chat history with the context and pre-established facts to build the real question from a much more refined position.
Sharing the actual chat link is more useful since AI responses are non-deterministic by nature so reproducing your interaction is very difficult.
Does someone else have good search skills but mingle traditional search engines with LLMs anyways? Why?
I use LLMs every day but wouldn't trust one to perform searches for me yet. I feel like you have to type more for a result that's slower and wordier, and that might stop early when it amasses what it thinks are answers from low effort SEO farms.
LLMs on the other hand (free ChatGPT is the only one I've used for this, not sure which models) give me an opportunity to describe in detail what I'm looking for, and I can provide extra context if the LLM doesn't immediately give me an answer. Given LLM's propensity for hallucinations, I don't take its answers as solid truth, but I'll use the keywords, terms, and phrases in what it gives me to leverage traditional search engines to find a more authoritative source of information.
---
Separately, I'll also use LLMs to search for what I suspect is obscure-enough knowledge that it would prove difficult to wade through more popular sites in traditional search engine results pages.
For me this is typically a multi-step process. The results of a first search give me more ideas of terms to search for, and after some iteration I usually find the right terms. It’s a bit of an art to search for content that maybe isn’t your end goal, but will help you search for what you actually seek.
LLMs can be useful for that first step, but I always revert to Google for the final search.
Also, Google Verbatim search is essential.
Google used to care about this but no longer does, pagerank sucks and is ruined by SEO, but it still "works" because if you're good you can guess the kind of source you're looking for and what keywords might surface it. LLMs help with that part but you still need to read it yourself, because they don't have theory of mind yet to make good value judgements on source quality and communicate about it.
But I'm not keeping my hopes up, I doubt the model has been explicitly fine-tuned to double check its embedded knowledge of these types of facts, and conversely it probably hasn't even been successfully fine-tuned to only search when it truly doesn't know something (i.e. it will probably search in cases where it could've just answered without the search). At least the behavior I'm seeing now from some 15 minutes of testing indicates this, but time will tell.
I agree that you can't TRUST them, but half the links regular search turns up are also garbage, so that's not really worse, per se.
Oddly, Microsoft recently changed the search version of Copilot to remove all the links to source material. Now it's like talking to an annoying growth-stage-startup middle manager in every way, including the inability to back up their assertions and a propensity to use phrases like "anyway, let's try to keep things moving".
Happy to see this feature set added into ChatGPT – particularly when I'm looking for academic research in/on a subject I'm not familiar with.
The main hard part of searching isn’t formulating queries to write in the Google search bar, it’s clicking on links, and reading/skimming until you find the specific answer you want.
Getting one sentence direct answers is a much superior UX compared to getting 10 links you have to read through yourself.
Google does offer an AI summary for factual searches and I ignore it as it often hallucinates. Perplexity has the same problem. OpenAI would need to solve that for this to be truly useful
For instance I searched for the number to dial to set call forwarding on carrier X the other day, and it gave wrong results because it returned carrier Y.
If we assume that people want a 'direct answer', then of course a direct answer is better. But maybe some of us don't want a 'direct answer'? I want to know who's saying what, and in which context, so I can draw my own conclusions.
- I count calories... eat out always and at somewhat healthy chains (Cava, Chipolte, etc). Tell GPT (via voice while driving to & or after eating) what ive eaten half the day at those places and then later for dinner. It calculates a calorie count estimation for half the day and then later at dinner the remaining. I have checked to see if GPT is getting the right calories for things off websites and it has.
- Have hiking friends who live an hour or two hours away and we hike once a month an hour or less drive is where we meet up and hike at a new place. GPT suggests such hikes and quickly (use to take many searches on Google to do such). Our drives to these new hikes learned from GPT have always been under an hour.
So far the information with those examples has been accurate. Always enjoy hearing how others use LLMs... what research are you getting done in one or two queries which used to take MANY google searches?
Case in point: Visual Basic for Applications (the Excel macro language). This language has a broad pool of reference material and of Stack Overflow answers. It doesnt have a lot of good explicatory material because the early 2000s Internet material is aging out, being deleted as people retire or lose interest, etc.
(To be frank, Microsoft would like nothing more than to kill this off completely, but VBA exists and is insanely more powerful than the current alternatives, so it lives on.)
The LLMs are nice because they are not yet enshitified to the point of uselessness.
When it starts with this you KNOW it's going to be maximum bad faith horsefuckery in the rest of the "question."
Which is also exactly something a bad-faith commenter would say, but if I lose either way, I'd rather just ask the question ¯\_(ツ)_/¯
The comment was dripping with condescension towards the use of LLM search, and that’s coming from a huge OpenAI skeptic.
Replacing "really for" with "more for e.g" would have been closer to the intended comment. I'll take that L.
Though if I can clarify: "can't be arsed to search" isn't a normative or judgemental claim against anyone, in the same way that "can't remember a phone number/directions" isn't. I'm speaking under the assumption there's a point between now and heat death where massaging search engine queries may literally not be as useful a 'skill' anymore. So there's less utility in young/old people taking the time to learn it.
But I can see how it sounds when I try to squeeze that into a shorter message.
Things that were previously "log a jira and think about it when I have a full uninterrupted day" now can be approached with half an hour spare. This is game changer because "have a full day uninterrupted" almost never happens.
It's like having a very senior coworker who knows a lot of stuff and booking a 30m meeting to brainstorm with them and quickly reject useless paths vs dig more into promising ones, vs. sitting all day researching on your own.
The ideas simply flow much faster with this approach.
I use it to get a high level familiarity with what's likely possible vs what's not, and then confirm with normal search.
I use LLMs also for non-work things like getting high level understanding of taxation, inheritance etc laws in a country I moved in, to get some starting point for further research.
I consider myself quite anti LLM hype and I have to admit it has been working amazingly good for me.
> "Can you provide a list of the ten most important recent publications related to high-temperature helium-cooled pebble-bed reactors and the specific characteristics of their graphite pebble fuel which address past problems in fuel disintegration and dust generation?"
These were more focused and relevant results than a Google Scholar keyword-style search.
However, it did rather poorly when asked for direct links to the documentation for a set of Python libraries. Gave some junk links or just failed entirely in 3/4 of the cases.
1) I know a little bit about something, but I need to be able to look up the knowledge tree for more context: `What are the opposing viewpoints to Adam Smith's thesis on economics?` `Describe the different categories of compilers.`
2) I have a very specific search in mind but it's in a domain that has a lot of specific terminology that doesn't surface easily in a google search unless you use that specific terminology: `Name the different kinds of music chords and explain each one.`
LLMs are great when a search engine would only surface knowledge that's either too general or too specific and the search engine can't tell the semantic difference between the two.
Sometimes when I'm searching I need to be able to search at different levels of understanding to move forward.
I feel like there is a mental architecture to searching where you try and isolate aspects of what you are searching for that are distinct within the broad category of similar but irrelevant things. That kind of mental model I would hope still works well.
For instance consider this query.
"Which clothing outlets on AliExpress are most recommended in forum discussions for providing high quality cloths, favour discussions where there is active engagement between multiple people."
OpenAI search produces a list of candidate stores from this query. Are the results any good? It's going to be quite hard to tell for a while. I know searching for information like this on Google is close to worthless due to SEO pollution.
It's possible that we have at least a brief golden-age of search where the rules have changed sufficiently that attempts to game the system are mitigated. It will be a hard fought battle to see if AI Search can filter out people trying to game AI search.
I think we will need laws to say AI advice should be subject to similar constraints as legal, medical, and financial advice where there is an obligation to act in the interests of the person being advised. I don't want to have AI search delivering the results of the highest bidder.
This becomes even better if the information you want is in multiple different places. The canonical question for that used to be "what was the phase of the moon when John Lennon was shot?". There didn't used to be an answer to this in Google - but the AI search was able to break it down, find the date John Lennon was shot (easily available on Google), find the moon phase on that day (again, easily available on Google) and put them together to produce the new answer.
For a more tech relevant example, "what is the smallest AWS EC2 I can run a Tomcat server in?
You 100% can get this information yourself. It just much more time than having an AI do it.
my search skills are good but either way that requires 3+ searches and visiting the menu of each restaurant and checking their hours and reservation. remains to be seen if chatgpt search is consistently good at this though
If you don't like that (like I do), you can also manually add it under Site Search using
I can definitely see this new search feature being useful though. The old one was already useful because (if you asked) you could have it visit each result and pull some data out for you and integrate it all together, faster than you could do the same manually.
It's often hobbled by robots.txt forbidding it to visit pages, though. What I really want is for it to use my browser to visit the pages instead of doing it server side, so it can use my logged-in accounts and ignore robots.txt.
They really had the potential to do something interesting, but were just focused on their ad metrics with the "good enough" search box. What have they been doing all the time?
If they're collecting data it doesn't even work; I make no effort to hide from them and none of their ads are targeted to me. Meta, though, they're good at it.
No offense, but you stating, that you work in the ads business rather makes me doubt, that you are talking from an unbiased point of view and makes me suspect, that there are motives at play here.
Actually the statement, that Google doesn't make money from personal data is kind of ridiculous. If they didn't make any money from that in one way or another, they why would they waste so many resources on collecting that data in the first place? It is really obvious, that they do make money with that data. Their other products are failing left and right. It has become a meme, when Google kills yet another of their products.
> It is really obvious, that they do make money with that data.
Google doesn't have to care about making further money because they have an infinite money machine from their monopoly on ad auctions. It's better to think of them as just doing whatever Larry Page thinks would be cool.
They mostly don't collect it though, like they don't do targeted ads in Gmail.
For example `8 hours ago: "autohotkey hotkeys"` with 4 links to pages which I visited while searching.
But this is a Chrome feature, not a Google Search feature. https://myactivity.google.com/myactivity does (sometimes? can't see it right now) have a grouping feature of all the searches made, but this is more of a search log than a search management feature.
So chrome://history/grouped is the closest to what I mean, but I can't pin or manage these history groups, enrich them with comments or even files, like pdf's which could then get stored in Google Drive, as well as get indexed for better searches.
I might be mistaken but i think ff mobile does something similar of grouping search session
[0] https://static1.makeuseofimages.com/wordpress/wp-content/upl...
In some fields of CS, places like MS research garner nearly 50% of all top conference publications.
OpenAI is Microsoft, technical details don't matter here, only money.
https://www.newyorker.com/news/amy-davidson/tech-companies-s...
You think Google could say no to NSA if they were asking nicely to put a tap?
Encryption between datacenters is to keep away other state actors, not the US.
Yes
> You think Google could say no to NSA if they were asking nicely to put a tap?
Yes
And, yet, aside from Aramco, they are the most profitable companies in the history of the world.
What does this mean? Like, I work there and I’d be pretty annoyed if they stopped turning a profit as a collapse of the stock price would affect my compensation.
It’s interesting to hear this take because I’m used to hearing the opposite: that Google is too focused on increasing short-term profit at the expense of product quality.
It won't directly match ChatGPT logs and OpenAI would just be pouring precious compute to a bottomless pit trying to partial-match.
Serve users a random version and A/B test along the way.
I’m sure they will/are tackling this at the model level. Train them to both generate good completions while also embedding text with good performance at separating generated and human text.
I can only await companies' attempts to publish enough junk to create an 'alternative truth' for new LLMs to believe in. The worst part is, it might even work.
Also they do share the most blocked/raised/lowered etc sites: https://kagi.com/stats?stat=leaderboard
We've had this problem of "good defaults" before with ad trackers blocking domains. I'm sure it'll be Sooner than later when some community lists become popular and begin being followed en mass
I'd assume right now the SEO target is still mainly Google rather than ChatGPT, but that's only an "I recon" not a citation.
If and when ChatGPT does become the main target for SEO spam, then Googling may start giving good results again.
As of October 31, 2024, the latest version of Java is Java 23, released on September 17, 2024. The most recent Long-Term Support (LTS) version is Java 21, released on September 19, 2023.
Which all seems correct and accurate.
is this a real question you needed an answer to, or a hypothetical you posed to test the quality of search results?
of course you're going to get listicles for a query like that, because it sounds like a query specifically chosen to find low-quality listicles.
The issue is with the query itself. You're assuming that there's some oracle that will understand your question and surface the relevant information for you. Most likely, it will use the word themselves as part of the query, which SEO sites will exploit.
A more pragmatic search workflow would be to just search for "most common programming languages used" [0], then used the Wikipedia page to get the relevant information [1]. Much more legwork, but with sources. And still quite fast.
[0]: (Screenshot) https://ibb.co/ggBLy8G
[1]: (Screenshot) https://ibb.co/H4g5bDf
I got:
Programming Language Latest Version Release Date
Python 3.12.0 October 2, 2024 ???!!!!!
Java SE 22 March 2024
JavaScript ECMAScript 2024 June 2024
C++ C++23 December 2023
C# 12.0 September 2024
Ruby 3.3.0 December 25, 2023
PHP 8.3.0 November 2023
Swift 6.0 September 2024
Go 1.22 August 2024 !!!
Rust 1.75 October 2024 !!!!
Kotlin 2.0.0 May 21, 2024 !!!
This is pretty bad. (???!!! added by me)However, my follow up query "Provide primary web page for each language listed above" was quite decent:
Here are the primary websites for the programming languages mentioned:
Python: https://www.python.org/
Java: https://www.oracle.com/java/
JavaScript: https://developer.mozilla.org/en-
US/docs/Web/JavaScript
C++: https://isocpp.org/
C#: https://learn.microsoft.com/en-us/dotnet/csharp/
Ruby: https://www.ruby-lang.org/en/
PHP: https://www.php.net/
Swift: https://www.swift.org/
Go: https://go.dev/
Rust: https://www.rust-lang.org/
Kotlin: https://kotlinlang.org/
The problem was with the 3rd query "Provide latest version for each language as mentioned on the primary website for that language."That brought back the first result basically unchanged.
So certainly this is a work in progress but very promising.
If now my content is going to be ingested and shown by a LLM or AI agent, what's the purpose to give it for free? I know it won't happen, but I would love if this type of agents have to pay to show a summarization of another website. It's only fair when done in mass like this.
I think we're going to see even fewer site visits as a consequence of AI search engines. The internet's ad-based funding model is going to dry up further, but the impact will be disproportionate. It'll be a few years til we see where the cards land.
If someone only creates for money, only publishes on the web to get people to look at advertisements, well... I think there are plenty of other people who don't feel that way that will fill the void left behind in their departure.
To me it seems weird so many people think the internet only exists because advertising props it up. The internet existed and was a wonderful place before advertising became widespread, and most services and websites will continue to exist after advertising is gone (if that ever happens). What encourages people to believe in some sort of great collapse?
If people stop visiting websites because LLM give them what they want, websites will stop existing. Don't believe me? Check how many "fansites" exist now about topics compared to ten years ago, when there weren't social networks. They have been replaced by influecners with huge followers on Instagram, TikTok, Twitter and more. The same will happen.
The first is taxation. Academia uses this. So we'll be fine there.
The second is donations and altruism. This includes Wikipedia, people who just want to share their ideas (like I'm doing now), Stack Exchange, etc. So we'll be fine there too, though this only works for low-budget stuff. Note that credit isn't necessary here; Wikipedia editors are rarely acknowledged, for example. But I do think credit is good, when possible (i.e. not for every single training source the LLM uses; ChatGPT Search is a good and practical way to give credit).
The third is to privatize the goods, for example, by creating copyright. LLMs won't get rid of this. They're not allowed to fully replicate a piece of text, since that is copyright infringement. So if you're getting people to pay for an exact or almost exact piece of text (e.g. a full book, a full newspaper article), you won't be affected there.
But if your business model would be affected by people summarizing your stuff, then yeah you won't be protected there. For example, if people would rather read a summary of your article or book rather than pay you for the full version. But this isn't new to LLMs. How many people jump straight to the comments to read the summaries instead of reading the article? How many people absolutely refuse to pay for paywalls? How many people block ads (I'm treating ads as a form of payment here)? It's possible to expand copyright to also cover things like summarization (i.e. copyrighting facts), but this is rather dangerous.
- The clickbait, SEO-optimized garbage that today fills 95% of search results could entirely disappear as a business model because they have nothing interesting to offer and the LLM company won't pay for low quality content.
- The average Joe blogging on their website won't go anywhere because they aren't profiting from it to begin with. And the LLM linking back to the page with a reference would be a nice touch. Same logic applies to things like Open Libra and projects that are fundamentally about open information and not about driving ad revenue.
But, on the other hand, I don't think LLM-based search will fundamentally change anything. Ad revenue will get in the way as always and the LLM-based search will start injecting advertisements in its results. How other companies manage to advertise on this new platform will be figured out. What LLM-based search does is give Microsoft and others the opportunity to take down Google as the canonical search engine. A paradigm shift, but not one that benefits the end user.
That's too neutral; the result is worse for the end user. It will be impossible to distinguish injected ads, unlike with Google. Furthermore, as it gives a much more direct "answer", people will also trust it more than a website linked to on Google.
If that's what they want to do in this space, which is not a given.
Why do I care if Google succeeds or dies?
If anything I want them to die for ad infested they've made the internet. I don't want ads in either chatGPT or Google Search.
Google will suck all your data even if you pay, and link the entire earth of services to your identity.
For now, chatgpt doesn't care, and I already pay for what they provide.
May they kill Google.
20 years old me would freak out hearing me that, they used to be my heroes.
The modern maxim is: any content platform large enough to host an ad sales department will sell ads
Vanishingly few (valuable) consumers have zero tolerance for ads, so not selling ads means leaving huge sums of money on the table once you get to a certain scale. Large organizations have demonstrated that they can't resist that opportunity.
The road out is to either convince everyone to have zero tolerance for ads (good luck), to just personally opt for disperse, smaller vendors that distinguish themselves in a niche by not indulging, or to just support and use adversarial ad blockers in order to take personal control. Hoping that the next behemoth that everybody wants to use will protect you from ads is a non-starter. Sooner or later, they're going to take your money and serve you ads, just like the others.
Unless enough people all pay, the whole thing stops working. But there aren’t enough people who will pay because most people don’t care.
Tldr: the ad supported business model fundamentally doesn’t work if you let all your best products (you) opt out by paying. It requires them to pay an amount far in excess of what they would be willing to pay for the system to work.
Frankly the example they posted seems like a fairly happy one, where the user is explicitly implying that they’re seeking a specific physical product to introduce to their life. We’ve all seen where those monetization incentives lead over time though.
But you’re right—not even so much as a tiny word “Ad” like Google does…
Basically, this means the answer is based on a main webpage.
It shows the cute Mastodon logo and all :D
Advertising has far less protection than is ordinarily afforded to the kind of speech you might do as a person.
They did not get addicted to selling ads, have billions in revenue from paying subscribers, and don't have to wean themselves off of ads (as Google and Meta would love to do).
I can see OpenAI's revenue per search ending up higher too since the ads will be impossible to distinguish, so even more valuable for advertisers.
I have been using http://www.Perplexity.AI since January 2023 for this exact reason. Unfortunately, since that incredible first UI (of yesteryear) it has been downgraded extensively (including no longer displaying footnotes, just adding numerals to the ends of factual sentences with tooltips).
The greatest thing about Perplexity is still that you do not have to log in (although it will bug you, particularly after lengthier or insightful conversations). Once a search/hybrid/chatGPT (how I've used it for almost yearS now) requires a log-in, it will immediately not be able to compete well with open-facing search engines like Google.
I have this same hesitation about using a pay search engine like Kagi, but am definitely intrigued by some of the other commenters' descriptions of the AI Assistant part of their service offering.
However it refused to search the internet for instances of my email address. Not sure what the benefit of hiding powerful functionality is.
I would like it to work as follows:
- using a model with an extremely large context size, analyze the top n results for a particular search phrase and store it in a vector db and let my chat session interact with it.
- it does not have to return results immediately, it could ask follow-up questions to reduce scope and improve the signal to noise ratio, or decide to augment the initial search on its own, once the searcher's goal is better understood.
There aren't any ads in their demo, we haven't seen the real deal yet, but I'll be watching HN for that day.
If adding ads causes a reduction in profit, from people moving to add free LLM (there are many), then it would be capitalist, in the interest of profit, to not have ads. And, let's say there's a scenario where all the companies have ads, the capitalist thing to do might be to remove ads, to capture all the users that don't want them.
Sure, once competition is gone, capitalism stops working, but we're not even close when it comes to AI.
Google was a very profitable business 10 years ago and the search was still decent. In the last decade they absolutely butchered their core product (and the internet along with it) in an effort to squeeze more ad dollars out, because it's not the level of profitability that they need to maintain, but the growth of that profitability.
Microsoft was a ridiculously profitable company, but that is not enough, they must show growth. So they add increasingly user hostile features to their core product because the current crop of management needs to see geometric growth during their 5 year tenure. And then in 5 years, the next crop of goobers will need to show geometric growth as well to justify their bonuses.
Think about this for a moment: the entire ecosystem is built on the (entirely preposterous) premise that there must be constant geometric growth. Nobody needs to make a decision or even accept that this is long term sustainable, every participant just wants the system to keep doing this during their particular 5-10 year tenure.
It's an interesting showcase of essentially an evolutionary algorithm/swarm optimizer falling into a local optimum while a much better global optimum is out of reach because the real world is something like a Rastrigin function with copious amounts of noise with an unknowable but fat tailed distribution.
<rant/> by a hedge fund professional.
Full disclosure, I had no formal education in writing... I suppose all credit goes to the actual great writers I take inspiration from.
I've never heard it framed like this before, that's beautiful.
Thank you for the kind words. I've been thinking of doing something long form, I'm just unsure if there's enough of an audience in this age of catchy tweets and tiktok videos.
Full disclosure, I had no formal education in writing... I suppose all credit goes to the actual great writers I take inspiration from.
[1] https://www.thestack.technology/microsoft-earnings-openai/
The fundamental problem is that ads based business model is much more lucrative then subscription based one. It's even more extreme when you take account of a prospective view, since you have control on ads shown which gives you a large margin for future revenue improvements compared to rigid subscription models. Unless you have a way to change this dynamic, you're going to eventually see ads in search results, regardless of its format.
How does it confabulate even when summarizing web results?
Longer term, it seems what will be left is “AI Optimized” content, which turns LLM search engines into shills for advertisers. Or these new search engines will have to compensate content producers somehow.
This can subtly (or not so subtly) rephrase and reshape the way we read about and think about every topic.
I'm going to use this as my daily driver for a few weeks.
The contemporary web is basically an epiphenomenon of Google, and they've failed to defend it. I hope OpenAI puts a huge dent in their market share.
I’m happy OpenAI is advancing LLM-based search but I won’t be using it in earnest until it’s local.
> how tf is it reading private repos ?!I usually assume good faith, but in this particular case, I believe the chance that this repo was public before and the author just changed it to private to bait attention is far higher than that Bing/ChatGPT can actually read private repo on GitHub.
I have a private repo named "portland-things" and I asked "does this user have a repo related to portland?" and it responded with "yes it's called 'pdx'" but that's not correct at all.
That said, anecdotally, I find it’s a bit hit-miss: if it’s hit it’s a huge improvement over google (and a minor improvement over chatgpt), if it’s miss it’s still good but get the feeling you won’t get anywhere further by asking more questions.
Perplexity sounds like a parody startup name from the Silicon Valley TV show. Way too complicated and unnatural.
It’s all about familiarity. Once people learn it, it’s not hard.
Perplexity is just a nonsensical word (for those unfamiliar with the concept) that is too long and hard to spell. They'd be better off just chopping it down to Lexity, or Lex, or Plexity, or Plex, etc.
(for those unfamiliar with the concept)
https://en.wikipedia.org/wiki/Perplexity
Reasonable name for a language model startup.
At first I thought it was some piano piece like "Mazurkas, Op. 59" by Chopin, or had something to do with some French guy in the AI field.
Edit: I recognize that I may very well be in the minority, but when I use search, I am looking for sources, not answers. Everything Google has done to get in the way between me and the people who have the information I want has been a net negative from my perspective.
The interjection of AI here is just doing more of that.
Google is instant, why would I wait for a bunch of text to generate just to get basic information.
Google results arent perfect, but the experience of Google is way better than any of these LLM search apps.
I don't even know how anyone can disagree.
You have to enter your query, wait for GPT to search, then wait for it to generate it response character by character and then you have to hope it what you want or you have to start over .
Starting over on Google takes no time.
So in this case at least, GPT Search is far inferior and dangerously incorrect were someone to rely on these search results for weather information.
Can confirm that free users who signed up for the waitlist can use it right now (even if they didn't actually get in)
I guess they could be using Bing as their search backend, which would mostly get around the blocking issue (except for searching Reddit which blocks Bingbot now).
Edit: I understand there is a freerider/economic issue here, unsure how to solve that as the balance between search engine/gen AI systems and content stores/providers becomes more adversarial.
I wonder to what degree -- for example, do they respect the Crawl-delay directive? For example, HN itself has a 30-second crawl-delay (https://news.ycombinator.com/robots.txt), meaning that crawlers are supposed to wait 30 seconds before requesting the next page. I doubt ChatGPT will delay a user's search of HN by up to 30 seconds, even though that's what robots.txt instructs them to do.
[1] https://platform.openai.com/docs/bots/overview-of-openai-cra...
The average person just does not discover content without the search engine recommending it.
The issue is that “AI search” has been a hot topic for a while now. Google (the default everywhere) just rolled out their version to billions of users. Perplexity has been iterating and acquiring customers for a while. Obviously OpenAI has great potential and brand recognition, but are enough people still interested in switching that haven’t yet?
There's some sticking power/network-effect/sticky-defaults effects, too, though.
It's _trivial_ to do a google search from anywhere on an android device with at most a tap or two. You can probably get close if a 3rd party has a well integrated native app but that'll require work on the user's behalf to make it the default (where possible).
Same goes for the default search engine for browsers/operating systems ... etc.
I will absolutely be firing off queries to google and GPTSearch in parallel and doing a quick comparison between the two. I am especially curious to see how well queries like "I need the PCI-e 4 10-gig SFP+ card that is best supported / most popular with the /r/homelab community" goes. Google struggles to do anything other than link to forums where people are already asking similar questions.
That being said, for all the talk about how bad google has become, I still prefer it to an unbroken bing.
Me too! I've really started to dislike Google search recently and am super excited we now have more viable options!
It doesn't matter. Google was late to the browser game, yet now the whole world runs Chrome. Analogy applies 1:1 to OpenAI.
More info on model distillation: https://openai.com/index/api-model-distillation/
I think this might actually be my main pain point with LLMs. Personally, I don’t want this.
I understand it might be helpful for other people. But, I prefer highly specific, advanced search functionality, such as site: or filetype: in google/ddg searches.
scryfall.com for magic the gathering cards is a great example. I’d much prefer typing a few brief flags such as “id=r” instead of “Get me all red identity cards.” And I know I’m getting all red identity cards with scryfall’s current search functionality.
They are also composable, so I can add/drop ones easily instead of perfectly rephrasing a whole sentence because I wanted to change one clause.
I’d need the same level of trust in the LLM’s filtering capabilities as I do in those boolean or regex matching field filters. An escape hatch to hard filters probably would be best for my experience searching things.
This is incredible and a direct threat to Google’s core biz.
As we used to say in the street "garbage in, garbage out": https://chatgpt.com/share/6723e865-d458-8011-b2ef-1a579026e6...
It feels like it might be. It feels tasteful in the same way that Apple ecosystem integrations just work really nicely and intuitively. But then again, there is an art to keying and retrieving embeddings, and it might just be that.
Bonus points for then being able to ask for the results in a specific format.
I’m looking forward to seeing how a feature built search engine starts to look.
Perplexity gave me the correct and best answers, with links to the relevant arxiv papers.
The new ChatGPT search gave me only cadical as answer, plus 2 irrelevant wrong answers (not multi-threaded), but missed all other multi-threaded solvers. => It's crap.
Neither Google nor ddg gave me any relevant links. Couldn't try kagi, since my trial phase is over.
Looks like the fellow who was invited to the Google funeral was right. Google search is dead.
> The Yogar-CBMC and JCBMC solvers are notable multi-threaded variants of the CBMC (C Bounded Model Checker) framework: ...
Followed by further details and references. The search results themselves look relevant and reasonable to me, but again, outside my area of expertise.
I'm trying to find proofs for strstr implementations, like Morris-Pratt in C. JCBMC is for Java programs.
Good answers would be Z3, yices2, CVC5, Kisssat. Bad answers would be CaDiCaL, minisat2 (the default), and all the other supported single threaded solvers for the dimacs and ipasir interface for the SAT problems (which are much simplier than SMT problems).
Just unrolling the loops and comparing to a small brute force strstr leads to very large models, which would need to run in several minutes, compared to the typical 20s. I'm considering running it on one of my big work machines with a parallel solver. Out of the 220 strstr's I suspect several of them to be wrong. I know that they are wrong, but CBMC gives me counterexamples where they would break exactly, and why. Better than a fuzzer, which just guesses around.
Their main example is exactly what's wrong with tourism and why everyone does the exact same things in the exact same places, it's the final stage of consumerism, you consume cities like you consume fast food or netflix shows.
You want to have a nice time on the amalfi coast ? Get there, leave your phone at the hotel and walk, explore things by yourself, try food by yourself, don't be a tool
This isn't search. Search can take some skill to wield, but what it should lead to is a high quality result. This might be someone's blog or a dedicated website about that topic. Information curated and edited by humans is going to become harder to come by. I wish I had more books.
So they probably wouldn't notice the warning signs anyways.
For example, if you ask LLMs to build code using the three.js library, nearly all of them will reference version r128. Presumably because that version has the largest representation in the training data set. Now, you can turn this on and ask it to reference the latest version, and it will search the web and find r170 and the latest documentation to consider in it's response.
I was already doing this before by adding "search the web for the latest version first" in my prompts, now I can just click a button. That's useful.
Of course, layering an LLM on top of garbage will still produce garbage.
It's able to query the relevant documentation, put it in its context and then use that to generate code. It's extremely relevant to giving existing models superior functionality.
The only other way to kill the web without killing LLMs in the process would be to create a way for people to upload structured public content directly into an LLM’s training. That would delay public content into release batches unless training can be sped up significantly.
I already use ChatGPT exclusively when I need information, partly because Google is so bad now, but mostly because it gives me the answer, not just some links that might have the answer. I feel the old web search paradigm is dead, not sure why ChatGPT feels the need to cover that increasingly pointless use case.
Definitely search is ripe for disruption.
It looks to me like ChatGPT Search hides links much more. Is there any incentive for website owners to allow access to ChatGPT?
The calculation would be different if you had a website large enough that excluding it would measurably make the chatgpt search service worse, which is presumably why OpenAI has been willing to pay some large players.
For context, I first tried this procession of searches on the Mac OS app.
1. "Who won the world series" 2. Who was the MVP?" 3. "Give me his bio"
My observations:
1. UX: The "search" button feels oddly placed, but I can't put my finger on it. But once I got it is a toggle, it wasn't a bit deal.
2. The first result had 3 logos, headlines and timestamps delineated, and easy to ready. The second one and third ones included a "Sources" button that opened a fly open menu. Clicking those opened a web link. The third result also included images in the fly open.
3. Citations were also inlined. The third result, for the bio, included a citation per paragraph.
4. It wasn't as fast as google. Which makes sense, given it's going through the LLM. But it will take a while to rewire my brain to expect slower responses to search.
5. Overall, I found the chat interface a very intuitive interface.
The second search I asked was "Give me a plan for a Thanksgiving meal."
I to a long response that felt like a weird mashup of LLM-generated content and search results:
1. A list of menu selections
2. Links to some recipes
3. Prepration timeline
4. Shopping list
5. Additional tips
There were 15 citations listed in the popup button, but only 3 inlined.
This was... not great. A traditional list of search results feels better here.
Overall, I like the direction. Innovation in search has been dead for close to 10 years, and this feels like I'd use it for certain inquiries.
I asked “Please find articles about planning a Thanksgiving meal for a family reunion.” It returned links to: GatheredAgain FavFamilyRecipes TastesBetterFromScratch etc. I like that it is returning niche sites I do not know about.
Today I was looking for an old (and useful) chat I had a few months ago and I had to export the whole chat history, wait for the zip file, and write a Python script to find what I was looking for :/
I just did the same search with chatGPT and it gave 6 bullet-point options with a reasonable description, though that was likely based off marketing copy. Half were white-label rebrands of the same option but that's not really ChatGPT's fault, and even then it was the one that best met what I've been looking for.
Hopefully ChatGPT's version works very well. Phind was more of a kludge to demonstrate what combining chat AI and search can do.
What were the limitations you ran into?
The one thing that kept me on Phind for a lot longer than I would have otherwise been was that I could pay for Claude usage when I over ran my daily quota. With Sonnet 3.5 there wasn't a point since I wasn't getting over the limits for Perplexity and you didn't provide a way to use an api key for Openai models at the time - to this day I find that GPT4 turbo is better than 4o for my use cases. This alone would have kept me around.
Since then Perplexity has gotten worse and they removed a lot of the old useful models. If you have something like poe.com where I can select the model to answer a query and give me fine grained control over the model settings I'd come back in a heart beat - e.g. if I'm doing research I don't know the answer to a high temperature is a good thing, if I want to find bugs in code a low one is better. My current killer feature is something like undermind where complex queries are run in the background and I get reports in due course by email.
Happy to chat more, my email is in my profile. I work in the space and have been building proprietary systems to help ground financial search for two years now.
(For example, I subscribed to Google Music a few years back, then got addicted to add-free YouTube, so I haven't canceled.)
It was too unreliable: I would use it for 4-5 searches, and it would be ultra-responsive and have a great UI. Then, all of a sudden, on a search it would be ultra-slow and fall back to a horrible UI that made absolutely no sense.
Edit:
Google's AI integration is good enough: Shortly after I switched back to Google, they started adding small AI answers at the top of the search. That's good enough for what I'm looking for.
Hopefully this also provides a strong negative force against SEO and, again, all the crap that comes nowadays thanks to Google.
(I sweat every day as I contemplate my web dev self huddling in the tall shadow of AI apporaching)
LLMs are just not the right thing to _be_ search. Just like I use the LLMs to generate prompts, I would use LLMs to help guide my search terms to cover what my first query missed. This can help to deSEO my searching too. This is really useful for deep topics of course.
OpenAI _could_ have approached the problem with this mindset.
1. a "raw" news aggregator. LLM may not keep up to speed with latest news. We may still manually review real-world pictures, videos or texts, which might be too important to be mis-interpreted by a LLM. A first-hand source is essential
2. a grep tool. LLM can not provide precise match with 100% confidence, or stuff that might be missing or ignored by LLMs
3. a knowledgable place where you seek answers. LLM excels at this
If there is a bias towards chatgpt like tools of even ~5%, it would be worth investigating why this is. My hunch is just the conversational aspect of describing at a high level and finding answers and avoiding all the distraction of several dozen windows to do something is worth it.
People on this site are seriously sleeping on it. It's a good reminder to never trust the hivemind on what's good or not. The assumption that the crowd has their thumb on the pulse of technology in general just because they do in some domain is unwarranted.
1. Mobile apps: Don't want to see intrusive apps
2. YouTube: Don't want to be interrupted with ads and no I don't want to buy premium service.
Actually i am logged into my iCloud on my macbook so guess that's why im seeing the search on that device of mine (not seeing on another where Im not logged into iCloud).
I'd like to see search (or research in broader sense) a more controllable activity with the ability to specify context + sources easily in the form of apps, agents and content.
https://chromewebstore.google.com/detail/chatgpt-search/ejcf...
https://developers.googleblog.com/en/gemini-api-and-ai-studi...
Google's grounding feature is mostly for API users and seems like a parameter that can be used to first search before providing an answer. It seems quite expensive tbh but otherwise really nice. 35 USD / 1000 requests
The search engine uses AI to connect people - and it actually works at this point, kinda.
So it is not it's own search engine and is still using Bing for its results just like the rest of them.
This doesn't matter if the results are user-hostile, as both search engines are.
To provide relevant responses to your questions, ChatGPT searches based on your prompts and may share disassociated search queries with third-party search providers such as Bing. For more information, see our Privacy Policy and Microsoft’s privacy policy. ChatGPT also collects general location information based on your IP address and may share it with third-party search providers to improve the accuracy of your results.
https://help.openai.com/en/articles/9237897-chatgpt-search (TFA links to it in the How it works section)Many search engines use the Bing index but return different results
I've tried Gemini flash, given it links to websites but it claimed to have knowledge, or be able to read it only so many times (kind of query, "summarize https://foo-bar/news-1")
As a very experienced SEO, this is pretty exciting nonetheless, a new front in the online war opening up.
If they're using their own scraper/search algorithms, it'll be interesting to see how they weigh the winners and losers compared to how Google does it.
That's not an introduction, that's a teaser trailer.
If they want this to be a viable search it needs to be available quickly, and anonymously from something quick to type in.
Google would have been annoying as shit if you had to go to google.com/search , let alone then log in.
You can even search “what [music genre] is playing in [city] this November” and it lists them.
All of those would normally take multiple clicks/manual filtering or ad filled aggregator sites.
You can take a look for yourself here: https://chatgpt.com/share/67245f76-bf88-8012-b72d-703beed088...
I've heard people find RAG not to work very well, is that accurate? Is it just about using the right embeddings?
I suppose ideally you just put the sources in the context window, which becomes limiting with large amounts of text?
> [Prompt] Can you browse the web to find out what lenses pair well with a Sony ZV-E10?
> [Response] The Sony ZV-E10 is well-suited to a variety of lenses, depending on your shooting needs. [snip]
To the downvoters: you should check your definition of innovation (hint: SV angel)
Repeating the same query in the same chat session gave me an accurate answer.
Or is this something they’ve already solved?
I used ChatGPT for planning a trip to Iceland a few years ago and it was nice to get a starter itinerary.
It was like using it to code - full of subtle mistakes and closed-down hotels
Edit: ohh, only Pro users? Right. ok. They made it seem like this was the big search launch and to go to chatgpt.com to get into it. Moving on.
It sucks but it’ll happen for sure
So ChatGPT‘s search looks rather rudimentary compared with Perplexity.
A percentage of Google's search market share is permenently gone, just a question of where it bottoms out...
Sure normal search is policed too, but usually not based on moral judgments but on legal necessities.
"How much tea tree oil by volume is in Dr Bronners tea tree oil liquid soap?"
A. ... However, the exact volume or percentage of tea tree oil in the formulation is not publicly disclosed by the manufacturer.
(which is incorrect, as the manufacturer disclosed it to me and I published it on the web)
One conclusion is that the web indexing is relatively shallow.
However ...
"Where does the founder of rsync.net live?"
A. "John Kozubik, the founder and CEO of rsync.net, resides in the San Francisco Bay Area."
... and the source is kozubik.com ... which means they did index my page but only retain, or weight, some of it ?
Meanwhile ... ublock showing >3k denials during this five minute interaction. I guess we can conclude something about where they are directing their time and energy.
We all know how this is going to end.
I can't believe that I'm saying this, but now after more than 20 years of using Google I'm finally paying for search.
But in other ways this looks like a great upgrade
( I tried getting the top hackernews posts but it was 5 days old? )
This kind of permanence is a huge loss
Does not support search for anyone wondering.
>Currently, I don’t have access to real-time data like time and date. You can check the time on your device or search "current time" online for the latest local time wherever you are.
Oh dear, we're off to a bad and slow start already.
Only then will the comparison with profitable Search engines be sensible. Before enshitification all VC-backed products are a delight to use, but after the honeymoon phase there’ll be ads all over the results/responses.
I think, I don't even want it to talk to me without searching the web first anymore. I want just sources and summaries. So I hope search will perform better.
This won't fix any issues with existing search, it will just make new forms of the same decaying cruft we're used to.
I recently started a blog hosted in my home server no point in making it public
Anyway. I've licensed my stuff as public domain to make it easier for ~~OpenAI to make more money~~ the AI ecosystem to develop.
Subjectively, I'm not switching for quick searched - google remains lightning fast and is good enough. But I already use gpt/claude/etc for conceptual searches and deeper analysis.
---
[used leica q3] ==> google (product listings and website; Chatgpt told me about the Leica q3 and mentioned ebay)
[value of mac air m1] ==> neither!! (google was useless videos and crap; chatgpt gave me a price range and useful explanation... which made no sense - used was the same price or higher than new!)
[vogue lyrics] ==> google wins (gave me the lyrics; Chatgpt whined about copyright restrictions and sent me to a youtube video)
[weather in nyc] ==> tie (both provided correct, rich detail about the current weather)
[root causes of ww1] ==> tie (both identified Militarism, Alliances, Imperialism, Nationalism, explained each and then mentioned the assassination of the Archduke as the triggering event)
[bohemia to midtown] ==> equally bad (both figured out that it's a request for local directions, but neither just gave me directions until I gave a specific destination)
[bohemia to penn station] ==> ??? (chatgpt correctly gave me bohemia ny where google picked some obscure local listing; otoh chatgpt wrote out directions where google gave me a nice map)
[btc to usd] ==> tie (both got today's price)
[what time is it in stockholm] ==> tie (both got it right)
[iphone 16 vs 14] ==> chatgpt wins (nice comparison; google didn't pop search labs and just gave me websites)
[ffmpeg to clip the last 3 secs of a video] ==> chatgpt ?! (I didn't love either answer TBH)
[456+789] ==> google (both gave the answer, but google included a nice calculator)
...and the stuff people really want:
[porn] ==> google (gpt whined about policy violations; google gave pornhub and other "useful" results)
For someone who's used online keyword-based search since the 1980s (computerised library catalogues, at the time), it's jarring for me to get over the distinction of querying for documents (old school) as opposed to asking direct questions, but that's precisely what LLM-based GPTs facilitate.
And as I'd noted this past June, it's a sea change in online search:
[O]ne of the upsides of GPTs / LLMs is that they provide direct answers to questions, though those answers may be hallucinations or generalisations. Even then, the directness is refreshing, though I expect it to also get polluted rapidly through both AI SEO manipulation and advertising / general enshittification of AI engines.
<https://toot.cat/@dredmorbius/112577405443953191>
I generally rely on Kagi's LLM for various reasons, but foremost is that it relies on current Web search and cites its findings specifically, which makes validating responses and detecting hallucinations far easier. ChatGPT specifically would hallucinate not only its responses but the citations it provided when those were requested, which curbed my enthusiasm greatly. It'll be interesting to see how its search-oriented offering fares.
I strongly agree with Temporal's excellent observation that this is, at least for the moment, a strong shift of the Web back to serving readers rather than advertisers (first and foremost) and publishers (distant second): <https://news.ycombinator.com/item?id=42011414>.
What I strongly suspect is that any successful GPT search tools will be rapidly engrossed by existing search monopolies. For those who defend the "free market" on the basis that competitors can emerge, the countervailing force of mergers and acquisitions must be noted, as well as the fact that these almost always effectively destroy that competitive potential, at least over the past half century or so.
The results seem to be better for strongly represented languages (e.g., English, Spanish, German, French), less so for those which may be less prevalent online (e.g., Yoruba):
When I say "governed" I mean a broad definition which is perhaps more akin to asking the question: "what kind of influences do we _want_ in play?" Spoiler alert: it isn't only governments that do governance; corporations do it too. Governance is an inevitable result of an organization or social hierarchy acting in the world. The complete opposite (abdication of responsibility) is anarchy. (One more thing that people may confuse: a libertarian philosophy is indeed a governance strategy, even though it is more hands-off than some: it won't work well in a Hobbesian "state of nature" since markets can't function well without some form of power (i.e. authority) that can enforce laws (such as property rights) and reduce violence and threats to manageable levels.)
> We collaborated extensively with the news industry and carefully listened to feedback from our global publisher partners, including Associated Press, Axel Springer, Condé Nast, Dotdash Meredith, Financial Times, GEDI, Hearst, Le Monde, News Corp, Prisa (El País), Reuters, The Atlantic, Time, and Vox Media.
A diversity of information sources is necessary but not sufficient for understanding the world well. It is too early to tell what kind of algorithmic balancing act OpenAI might or might not use for the information sources above. But money talks, so my expectations are not high.
It is not in the public interest to prioritize for-profit "news" organizations such as News Corp. There is just not enough overlap between the Murdochian (Molochian) financial incentives and the needs of an informed citizenry.
Serving the public interest is a waning part of OpenAI's charter, if it remains at all. It is instructive to compare how a U.S. citizen can run/vote for a local school board against what it would take for that same citizen to influence an OpenAI decision. My point: it matters relatively little what OpenAI says; what matters more is their charter, structure, incentives, and mechanisms.
What are the chances that an a-democratic, techno-elite organization (like OpenAI) will lead us boldly and wisely into the future? What about the current slate of western democracies? I'm pessimistic on both. My take is that we don't seem to have the institutions, political will, or understanding to be up to the task of handling AI transitions well. It could go in any number of directions.
'beef vindaloo recipe from nagi [0], but then just strip all the shit out and give me the recipe'
...and got exactly what I wanted. The web's dead, baby.
[0]: The lady behind recipetineats.com, a terrific recipe site. And she's an A+ human.
I don't understand how this point doesn't come up more often. I don't want to waste time to formulate a question when I can just use 2-3 keywords to get a result with traditional search.
Imagine having to type "What does this error message mean..." Everytime you paste an error message into Google/ddg/kagi.
A former colleague who is mediocre on innovation but arbitrarily good at self promotion (TLDR tried boosting on ads, got a huge win, almost got fired for ethics breaches, did nothing interesting again) had the audacity to post a weird circuit board background with glowing flowing hair about his “innovations at OpenAI”.
He’s not the only total fraud I know there, it’s just that every morally weak person I know went there, almost to a rule, and none of them (safe one) is any fucking good.
GPT-5 never happens. Prove me wrong.
I provided a query to a _somewhat_ common question with Java apps and TLS:
Query:
> identify root cause of this Java exception: PKIX path building failed: sun.security.provider.certpath.SunCertPathBuilderException: unable to find valid certification path to requested target
Answer:
> The exception you're encountering, PKIX path building failed, typically indicates an issue with SSL/TLS certificate validation. Specifically, the error unable to find valid certification path to requested target suggests that the Java application cannot find a trusted certificate for the server it is trying to connect to. Here are the common root causes and solutions:
Missing Certificate in Truststore: The server's SSL certificate (or the certificate of its issuing authority) is not present in the Java truststore. Solution: Import the server's certificate into the Java truststore using the keytool command. You can obtain the certificate by using a browser or tools like openssl. Self-Signed Certificate: If the server is using a self-signed certificate, it won't be trusted by default. Solution: Again, import the self-signed certificate into the Java truststore. Incorrect Truststore Configuration: The application might be pointing to an incorrect truststore or not configured to use one at all. Solution: Ensure the application is configured to use the correct truststore and that the path is correctly set. Expired Certificate: The certificate used by the server might be expired. Solution: Check the server's certificate validity and renew it if necessary. Intermediate Certificates Missing: If the server's certificate chain is incomplete (i.e., missing intermediate certificates), it can lead to this error. Solution: Ensure that the complete certificate chain is presented by the server or manually import any missing intermediate certificates. Network Issues: There may be issues with the network or firewalls blocking access to the certificate validation services. Solution: Check network configurations and ensure proper access. To troubleshoot further, you can enable Java SSL debugging by adding the following JVM argument:
bash Copy code -Djavax.net.debug=ssl This will provide more detailed output regarding the SSL handshake and help identify the specific issue.
It’s not a bad suggestion, these are my usual go-tos when diagnosing cause.
But then a query into what the sources are for this information is:
> The information is based on common knowledge about Java security, SSL/TLS configurations, and best practices for handling certificates. It reflects general principles and troubleshooting steps recognized in Java development and security documentation. If you're looking for specific references, the Oracle documentation on Java security and the keytool command, as well as community forums and resources on SSL/TLS issues, often cover these topics.
“It’s just common knowledge bro. RTFM”
The truth is, I haven't used ChatGPT at all since spring of this year. Claude's Sonnet 3.5 has replaced it. I pay very little attention to what OpenAI releases and simply waits for Anthropic to implement it.
I also started using Gemini which already outperforms perplexity and this and will not switch.
I think everybody is constantly caught up with their infatuation with OpenAI and other characters that they don't realize Google, Anthropic are actually building a moat which some like Gary Marcus keeps rambling on as impossible
I'm a realist and I can see that while Google has been slower to start, it reminds me of the search engine wars of 2000s, it is dominating and winning over users.
Related, I haven't paid for Gemini since about a month after release, but the morally corrupt query of "Show me articles from left and right leaning news sites about <headline topic>" would result in Gemini censoring right leaning urls with a "url removed" placeholder and belittling statements about the concerns of showing me right leaning content. Perplexity had no issue with such a dastardly prompts.
I want a tool, not a curated experience, so Gemini is in my "will not use" list for the foreseeable future.
I admit I haven't tried this lately, but I also have no desire to help fund that sort of behavior.
Ugh. It'd be nice if tech companies didn't treat us all like infants.
It's not looking good for Google. I'd hope Anthropic and Gemini could capture more consumer market share from OpenAI but it's not looking good. Tell the average person about Claude or Gemini and the only thing you'll hear is "oh so like ChatGPT" artifacts would not be enough to convince them.
I don't think OpenAI dominance in consumer benefits anyone in the long run.
it's okay, we're not destroying the world and you're not a better person because you purportedly care about what humans are supposedly doing to the world and because you think the rest of us don't
Smell that!? A large part of Google's search business is on fire right now!
There are three types of search: informational, transactional and navigational.
LLM's are competing hard and fast for informational search. Once upon a time we offered 2.5 keywords to the Google Gods only to be ultimately passed to stackoverflow.
That game is up. Google is losing it faster than you can say, "anti-competitive practices in the search engine industry."
Transactional and navigational search remain.