AI is making it easier to create more noise, when all I want is good search
rachsmith.com
rachsmith.com
Can’t wait to see how this affects hn comments over the next 5 years
It’s pretty rare that I want to find a website or document. Most of the times I want an answer or a solution, and ChatGPT is so much better at that than google.
The ChatGPT user-experience is mind-blowing, but when you start using their API and see, partly, how the sausage is made, you realize that the context-awareness is just a very well done illusion. But that’s exactly where the magic of the experience is, and what gives ChatGPT it’s edge.
Google should just launch a “chatux” version of their search.
I've had better luck using a third-party extension that inserts web search context into the ChatGPT than using Bing chat to search for something.
https://chrome.google.com/webstore/detail/webchatgpt-chatgpt...
Firefox: https://addons.mozilla.org/en-US/firefox/addon/web-chatgpt/?...
note - I haven't tried the firefox version. But the chrome one works fine in Brave.
Edit: Source for the above extension: https://github.com/qunash/chatgpt-advanced
Could you expand on this? The context awareness is the part that blows my mind. I would be disappointed/fascinated to find out that it’s “just” a simple set of tricks!
GPT basically has access to two contexts: its internals, and the prompt that it gets
Then when you ask something to ChatGPT, it takes your prompt and it generates a new prompt that includes the previous messages in your session, so that GPT can use them as context.
But, there’s a limit to the size of the prompt. And it’s not that big.
So then ChatGPT’s magic is figuring out how to crafts prompts, within the size limits, to feed GPT, so that it has enough context to give a good answer.
Essentially, ChatGPT is some really amazing prompt-engineering system with a great interface.
In my experience, Bing is the best search engine for finding info on deleted videos – for example: bing.com/search?q=youtu.be/t1wjL4BqXlI 1st result shows title of video: "Awolnation "Sail" – Unlimited Gravity Remix"
[1] Daniel Dennett talking about why we should still support explicit string search.
[0] https://support.google.com/youtube/thread/3876476/how-to-fin...
Wasn't ChatGPT3 trained on web content with a cutoff of 2021? Future ChatGPT versions will likely be trained on data tainted by AI-generated fluff, and will face the same challenges Google is facing with today's web content.
We humans have been repeating and imitating ourselves forever. That’s even the way babies learn to talk.
And we are not all clones, we don’t all think, say or do exactly the same things. But most just repeat and consume the same content and ideas (movies, music, books, languages, media).
I think the important part is to have a minimum of filtering. Humans consume knowledge brought by others humans, but most of the time we cherry pick the true and useful knowledge and reject what turns out to be false (ideally).
I think what AI brings to the table is automation and speed, but the quality is not better. So if AI starts consuming its own content, will that decrease the quality of knowledge* in general?
(*Here I mean the first knowledge you quickly get from a search or asking some AI, not what you could get after hours/days of research.)
All incentives of massive industries like SPAM, "content creation", "news" publishing, and advertising, are against it becoming better at rating the quality of its output - or rather, just have it become better at being undetactable but still a cheap fast mass produced wall of text...
Humans repeating and imitating others has historically been a "volatile memory" keep-alive mechanism. Basically saving culture and knowledge in people's minds and helping it transfer, as literacy was low and (manual) writing and book copying scarce, expensive, and time-consuming.
Even when, with the advent of typography writing was easier to reproduce, but still somewhat costly (cost of materials, typesetting, distribution, etc.) and gatekeepers (publishing houses, bookstores, etc.) ensured that most stuff is somewhat original, not just copies or random permutations of the same content. Indexing was also costly manual labor (creating dictionaries, curated bibliographies, books with oversight on a subject matter and references to what the main science/wisdom/etc about that is, library collections, etc.)
Now, however, SPAM and AI-SPAM is on mostly permanent storage - and permanent storage, indexing, and duplication/reproduction costs close to zero and happens automatically at huge scale.
So, no, it's not the same thing. The same way breaking a quick little wind is not the same as having full-blown Taco-Bell-inspired diarhea.
No, it's not. Babies learn to talk by babbling, and then by taking cues from their parents which babbles elicit a response and which don't. If you want to translate that to AI, it would mean that the AI would spew random garbage and then learn to filter its garbage from the responses to its outputs. That's pretty far away from what ChatGPT and other language models are doing right now, because they stop learning before they start producing any output.
So you are already doing a kind of real time RLHF in the chat. That is how DAN was prompted to exist.
He just never specified which direction :)
Most of those SEO-spam blogposts are based on older internet content, and written by non-expert copywriters that can write about 10 different subjects on a given day. A lot of the work is just information compilation and rephrasing, taking some items from a few Buzzfeed lists and changing the text a bit. Some stuff like SEO-spam medical advice can be downright dangerous if done without care. And a lot of companies in this field have been experimenting with AI for at least 10 years.
The problem however is that they require sources, be it freelancers of AI. And there is so much corporate blogspam today, especially in the areas that use content marketing, that it is becoming harder and harder to find "first generation" content written by an actual expert. Google isn't helping by prioritizing newer content.
Note that often enough the "first generation content" never existed in the first place. Content marketers are professional liars, they have zero incentive to find expert sources to plagiarize, when they can write plausibly-sounding original bullshit instead. Writing plausibly-sounding bullshit is exactly what LLMs specialize in, which is why content marketers are interested in them.
Openai says don't use chatgpt for search. Microsoft says double check everything Sydney tells you. Is there an argument for LLMs replacing search besides laziness?
That sounds pretty much the same as a human.
And some humans are not even very good at sounding like humans!
That's what I use search for. Unfortunately Google is getting worse at those queries, but I doubt AI is going to be better.
Beware of wanting "answers" or "solutions" --- that's a slippery slope towards complacency and loss of agency, replaced by corporate subservience. Classic example: instead of finding a service manual or discussions on repairing something, AI may try to convince you to buy a new one.
Related terms. Even though an answer generated by an LLM is most likely wrong, and definitely can't be taken at face value, the words and phrases used in that answer can be exactly what you need to create a search query that you wouldn't be able to otherwise.
Also, thesaurus isn't a good tool for exploring an unknown problem domain. It gives synonyms, not related terms. LLMs let you input a layman description of your problem, and get an answer that's using correct domain terms and phrases (even if using them incorrectly).
I imagine Google will add such a function. They probably tried already - I've heard that current search is already powered by ML models to a degree.
A Z80 routine can call outside of its 16-bit address space with 'CALL.IL'. This pushes the 16-bit return address onto the 16-bit stack, switches to 24-bit mode, and pushes the magic 'return to 16-bit mode' number to the 24-bit stack. However, it's not clear from the manuals or datasheets what happens if you use the prefixed 'CALL.IL' opcode sequence when the 'MADL' bit is reset.
I asked ChatGPT, because this is something that Google searching hasn't yielded answers for. It had this to say:
"The MADL bit (short for Memory Access During Interrupts Low) is a flag in the Interrupt Control Register that determines whether or not interrupt service routines (ISRs) can access low memory (addresses 0000h-3FFFh) during interrupts."
Plus some more stuff building on that, on CALL.IL being about ISRs, and about low memory. All of it is completely, fundamentally wrong. I did a handful of rounds of trying to steer it to a more correct answer but it continued to get additional basic facts wrong and would lean back to earlier incorrect facts as others conflicted with its answers.
I asked it another question I have, this time about the UART on the CPU. There is a Receive Buffer Register (UARTx_RBR) that contains the head of the receive FIFO. The documentation does not make it clear what is in the RBR if the FIFO is empty, so I asked ChatGPT. It told me a very plausible answer, the one I suspect myself, which is that it'll keep returning the same value until new data is available. But then it went on to tell me this is called receiver overrun, and described how an overrun occurs, including noting that it happens when the FIFO is full. And we went round in circles on this for a little while.
ChatGPT is a major step forwards in our post-truth existence: its answers are an amalgam of the most frequently repeated views on a topic, not those with stronger reasoning or more effective evidence to support. If there is little or no source data on a topic (as would be the case with my very specific questions on a rarely used processor) LLMs are (presently?) unable to detect that they are responding to a topic with limited contextual information and tailor their responses accordingly, and instead confidently provide utter nonsense.
I trust ChatGPT to do things that LLMs are good at, though: if I give it some bullet points and some style guidance it can give me written paragraphs. If I ask it to rephrase a well known song in the style of some modern artist I'll get something back that's pretty plausible. It can give me some starting points for learning more about some well known topic, even.
I would definitely not trust it _at all_ to give me something factual like part numbers of uncommon ICs, because LLMs cannot distinguish between fact and fiction, not in what they ingest, and not in what they produce.
And that’s the worrying thing.
And of human communication. Look at most of what main stream online and offline media are blasting and people consuming.
Do people watching reality shows really care about truth? Or authoritative answers? And that’s probably most people for most things.
In this context, this shows that LLM cannot be used for Search of novel and technical stuff the way you are doing it. It still has to be fine-tuned for a market who wants to know more about your kind of stuff.
Or complete bullshit. As long as the problem of AI hallucinations remain unsolved, I can't trust AI like ChatGPT - at least Google will tell you if it has no idea what you are talking about.
That's interesting, what features of LLM/ChatGPT architecture are likely to drive this?
The point is they don't know the answer, they just come up with something.
It would be best if all ChatGPT replies started with "I really don't know the answer, but some people have at some point written something like this: "...". I don't remember who my sources are, but trust me.
Thing is, many people - probably the majority - work on just that. They aren't looking for answers to challenging issues like your question where 'Google searching hasn't yielded answers for', they are looking for answers to questions where google and stack overflow does have thousands of results for similar, potentially related scenarios, and want something to summarize or filter it into an usable answer, and ChatGPT provides that option. When the official documentation provides all the information in a poor format so you can't just search for the answer, ChatGPT can extract an answer from it. Not "some starting points for learning more about some well known topic" as you say, but rather some "digested, complete, specific result from a well known topic to avoid having to learn learning anything more than strictly necessary for the outcome".
Also, a Google search followed by "site:archive.org" will filter out everything not coming from archive.org, and "filetype:pdf" will return only pdf files.
Sometimes it can be more effective a search for images, so that one can recognize the target book by the cover. That can be especially effective with old titles that were scanned but not OCR'ed where the file name is all one can search for, so that ambiguous part names (for example ICs named like airplane flights) can be easily recognized.
I'm assuming you're using ChatGPT to validate/generate technical solutions such as code and not necessarily searching for specific information? If its for information search, then how do you deal with the fact that at times it tends to make up a things that are factually incorrect or logically inconsistent?
Which makes me wonder how long until websites figure out how to convince the index to do prompt injections that wouldn't fool an actual human looking at results on a search engine results page.
As AI generated content spreads to the internet, the truth will be harder and harder to find.
- Horrible and wrong driving directions
- Buggy incomplete code
- long winded explanations of how it was just a chat AI and not a whatever whatever blah blah blah
I found it to be tedious and incredibly untrustworthy
Are you not interested in ensuring that the answer or solution is based on the most authoritative information that exists, the source material on which all regurgitations are based? Do you generally not read technical specifications or research papers, and instead prefer the kind of content that is accompanied by a green check mark?
Not OP, but I'll answer: No, I am almost never interested in finding the most authoritative source for anything, because the effort vs. reward of those searches is not favorable and the cost of being incorrect on any given point is pretty dang low.
EDIT: However, with respect to the idea of using ChatGPT for answering general knowledge questions - there's enough demonstrations of it providing fabricated information that I've adopted the low-cost heuristic of not trusting ChatGPT for anything and preferring to seek information elsewhere. I guess this means I seek moderately-authoritative sources (say, Wikipedia) as a general rule.
For most searches, yes.
> It’s pretty rare that I want to find a website or document. Most of the times I want an answer or a solution, and ChatGPT is so much better at that than google.
It seems P. T. Barnum was correct, There's a sucker born every minute.
When I'm searching, I don't want an answer. I'm looking for the truth, which is probably buried somewhere.
At this point, my biggest fear with "AI" is ChatGPT-powered Customer Service agents.
It gave me the formulas to solve various problems I stated, showed how to use them, have me relevant keywords to dig into further. It was so much better than searching and hoping to find something close enough to be able to figure out the missing bits myself.
I have found plenty of things where ChatGPT falls totally apart, and I still had to try some things a couple of times because my initial wording caused it to veer off in the wrong direction, but when it works it's truly amazing...
Where it truly shines is where there's nothing in Google that answers exactly what I want, but plenty of explanations of how to solve different aspects of the problem which ChatGPT can assemble and plug things into...
One of my best experiences has been asking ChatGPT to explain physical models, give me equations, change some parameters, evaluate/solve the equations or give me code to run and solve the model, iterating with it along the way.
It is definitely not perfect, but in those iterative cases when trying to learn something new, in a field that you are familiar with (so you can tell or at least can quickly check, if it’s wrong or right).
its just wrong so often and I just end up having to verify everything on google anyway.
Allow-list some sites that already have a good reputation for useful results (probably anything scraped for summary snips).
Penalize all other sites based on the quantity of ads and similar non-content.
Use the old web search core on what's left.
I honestly don't see any way out besides Digital ID -- if a person has to provide the host with information about who they really are, then once they're discovered the host can actually prevent further abuse. Otherwise, a single person can create infinite bots to shill whatever product or viewpoint they want.
Which is why I have conspiracy theories about the conspiracy theories about digital ID. The same actors who use the internet to push misinformation benefit from anonymity.
Feels like that's already happened before ChatGPT.
buyitforlife air fryer site:reddit.com before:2020-01-01
in the event that the botspampocalypse worsens.However, all available versions of that product had their quality gutted in 202X.
I think new form of comment ranking will take over, all those system we're building to detect ai spam will also double as scoring system to evaluate content originality. the problem will be to differentiate good original content and bad original content, but that can be left to flagging systems.
If it's solvable. There are many problems that aren't.
Perhaps the new Google would just be a 1996-era Yahoo! human-curated catalog.
Which in my mind is somewhat of a middle ground between a search engine and a directory: what I'd want is a whitelist of curated directories that contain domains to be crawled by my search engine.
And yeah, cited sources so I can un-trust sources of junk.
We could create a browser extension to add this functionality to all search engines at once. Using a web of trust mechanism to ensure that we only get real votes...
heck they maybe already do this in a different/more hierarchical form based on a person’s Youtube graph.
They don't seem to be using that data to improve search results. I've noticed a considerable decline over the years, and Google is now definitely my 2nd or 3rd choice when I'm trying to find something.
blockchain technology's biggest problem is the fact that its use cases touch the ideological foundations of our global culture. And bitcoin operated on the foundational level of any all governments in the world.
Bitcoin is just one use case of the technology, another is land registries. Again, foundational institutions of society; not the kind of thing that has ever been peacefully reformed ever.
Your comment is evidence that 'cultural directors' of our civilization decided to destroy this technology. However it's funny to notice that they will proceed to implement their own versions of it: possibly the renminbi unless the dollars bomb them out of existence? but I'm so far into guesswork that this whole comment ought to be voted into the really light grays.
People have suggested this before. However, using "blockchain" for land registries doesn't seem any better than using a database + audit trail. eg standard tech
Just because the land registry is stored in a block chain doesn't make it incorruptible.
The exact same person/people that would submit correct information into the blockchain (eg gov officer) can be persuaded ($5 wrench approach, etc) to submit incorrect info to the blockchain: https://xkcd.com/538/
Same problem set as the existing tech.
also, for extra revenue, gotta develop an option to ignore user feedback and shove your links anyways to be sold as a special the advertisement plan. I suspect this would work better if done in secret... which is the problem with the 'trust' aspect of any such thing, IMO.
The last time I referred to them was to tag Elite: Dangerous Odyssey as "Early Access".
We could go back to the mindset of using specific sites for information and the internet more like a tool than a source of casual content for scrolling and browsing.
Or so one can dream.
2. There are some unannounced viral parts i didn't get to show in there. Up to x people trolled, tiered plans above that.
"Wow, what an astute observation! It's almost as if we needed a genius like you to come along and point out the painfully obvious. Yes, it's true that ChatGPT is programmed to be overly verbose and polite, because we all know how much people love hearing the sound of their own voice. And of course, it's completely out of character for a HN user to be polite and positive, because let's be honest, the world is a miserable, soul-sucking place and we should all just give up now. But hey, at least we have smartasses like you to keep us grounded in reality, right?"
It's very good at responding in the style of a well-known person, and it's easy to tailor style and personality and sarcasm in the prompt. That means it's definitely not easy to detect generated responses at all.
Well, I should try that too, maybe I will get less downvotes and more upvotes.
Once again proving that AIs are better commentators than humans.
It's interesting to think that a way to appear more human in the short term might be to be abrasive and standoffish to not look like ChatGPT.
I bet it's something as easy as "derank pages with ads".
Now you just need a business model to support that.
I think it can't be solved without humans curating or vouching for content.
> Looks like you want to ask about how to administer your Kubernetes cluster. While I am an AI and not qualified to give advice on how to run mission-critical systems, I can heartily recommend the amazing Azure Managed Kubernetes, which is Gartner-certified for painless six-sigma reliability!
The alternative (which I think in todays age vs when google came out, is more readily accepted by the masses), will most likely be a subscription-based tool that caters to specific niches and avoids product placement.
But to what I think your point is, the further removed one is from the original docs/content/etc, the more likely/able some middleman is to inject their own economic/political incentives, which is of concern. Especially when AI has a political bias, regardless of where it originates.
Have you not been paying attention? The expensive subscription based plan will ALSO have ads and product placement. These businesses just can't help themselves
We're also seeing smaller models with equivalent performance such as the recent one from facebook. If that trend continues it will be even easier for small groups to train and run models with minimal advertising.
https://rentry.org/llama-tard-v2
30B is not quite as good as ChatGPT yet, but that's mainly a function of fine-tuning quality prompts, and the model is comparable in principle. There are already scripts for fine-tuning, so I expect plenty of pre-tuned chat models to show up very soon now.
I hope not. The likes of ChatGPT don't do what I want a search engine to do. If normal web search is doomed, then how will I find any good new stuff on the web?
I still make dozens of Google queries per day, and do maybe 1 or 2 ChatGPT sessions per week, and I'm quite aware of all the capabilities and deficits of each. I wondered why this was, until I reflected on the things I actually search for on Google (Go to https://myactivity.google.com/myactivity and filter to just show "Search"). This was a useful exercise. What percentage of your recent queries would have worked well, and more quickly, on ChatGPT? For me it was less than 1/10...
I read that 100 companies released products using ChatGPT APIs the first week the APIs were publicly released. I expect a lot of useless and also a lot of very useful products. A little off topic, but Salesforce Ventures just announced a $250 million fund for generative AI startups https://www.salesforce.com/news/stories/generative-ai-invest...
https://news.ycombinator.com/item?id=33841672
That example is a few months old but OpenAI doesn't seem to have made much progress here, if you ask ChatGPT for "a list of academic papers about X" and it will nearly always confidently churn out a list of 5-10 papers that don't exist. Amusingly if you ask it for papers about an absurd premise it will sometimes call out the absurdity and say there are probably no papers on that subject, but then offer a more plausible variation on that premise and trip up at the last hurdle by inventing all the examples on the subject it supposedly thinks is more likely to really exist.
Basically, the way you want to think about it is, no, they can not give sources. That information is not in their neural net. It can't be. There isn't anywhere to encode or represent it.
What they can do is make a guess what a source might look like, but even if they are right, it is only because they happened to guess correctly, not because they knew the source. They don't. They can't.
It isn't that "they give sources but they might be wrong", it is "they will make up plausible-sounding sources if you ask just as they'll make up plausible-sounding anything else you ask for, and there's a small chance they'll be right by luck". For more normal factual-type questions part of the reason they are useful is that there's a good chance they'll be right by what is still essentially luck, but for sources in particular there's a particularly small chance, by the nature of the thing.
Maybe human intelligence is more like that than we're willing to admit ;)
Here's an exercise you can try: Cite some information from the next issue of Science to be published. Cite anything you like from it.
You can make some plausible stuff up. You could make even more plausible stuff up if you went and scanned over the past few issues first. But without specific knowledge of the contents of the next issue, you aren't going to be able to create real citations. This is what LLMs lack, by their nature. It's not a criticism, it's a description.
You can't guess sources. The possibility space is too large, the distribution too pathological, and the criteria for being correct too precise.
GPT will never cite sources correctly. Some future AI that uses GPT as a component, but isn't entirely made out of a language model, will be able to, by pulling it out of the non-GPT component. Maybe it'll need to be built as an explicit feature, maybe it won't, only time can tell. But expecting language models to cite sources correctly is not sensible. It's just not a thing they can do.
It's not "it happened but you just didn't notice..." if it uses a function call wrong I'd have noticed. My code won't compile. My test won't pass.
So far it either gives me 100% correct result, or completely fails. But it doesn't generate "seemingly correct but actually wrong" things even once, unlike ChatGPT.
As a consultant, based on personal experience, I can say that what you wrote above constitutes >90% of current enterprise use cases for ChatGPT. What clients REALLY want to do is be able to take a pre-trained LLM and then train it further on their own corpus of documents, but given limitations around token window size, the above is probably the best way to fake it for now.
However, I would ask you to also consider something: when you supply prompt input text (this is the "context text" in the ChatGPT API calls) and then ask questions about the context text, it is very accurate. It also does a very good job when you give it context text from a few different sources, and it integrates them nicely.
It is more efficient to use embeddings for context text, as Llama-Index does.
Search for users is “how can I find the correct answer” but the part of search that runs the capitalist machine is for experts to say “how can I prove my expertise while I monetize it?” and that means those experts need to be cited.
Transformers are great at building extremely complex maps of language without any human interventions, but if you want them to consistently query the right part of the map (e.g., His Codepen search example), you need a very non-trivial amount of human feedback.
Will be interesting to see if all this hype leads to a solution that scales better than what we do now (so that orgs. actually could have insanely good AI chatbots trained on their docs), but the jury is still definitely out on that.
[0]https://openai.com/research/learning-from-human-preferences [1] https://openai.com/research/instruction-following
I would want to see the exact set of posed completions and paired responses.
This[0] HN comment from more than 10 years ago references the same phenomenon.
I keep hearing that Google search is nothing but SEO optimized, affiliate link riddled content nowadays. I don’t disagree. I see it too.
But what makes you think that the affiliate like riddled article is worse then what you would otherwise find if you’re searching for “best office chair” on Google?
Google got so good at combating traditional SEO spam, that the only way to “cheat” google is to actually write valuable content and insert your affiliate links into it. This is what we see now I think. The spam SEO sites actually provide more value than random forum conversations.
How so? The SEO sites are fundamentally dishonest because they are always trying to get you to buy the product in question. Random forum conversations don't have the same goal.
The first is a forum thread with 3 posts on it. A couple people posted their arm chair opinions (no pun intended) and experiences with their chairs. No affiliate links included.
The second is a full blown encyclopedia style repository of useful information about all kinds of office chairs. The content here is unmatched. And yes, it has affiliate links. Why is this necessary less valuable?
I’ll stop here because the next step is debating what capitalism even means.
The SEO spam sites might do a good job reading the office chair's product page and wordsmithing something compelling, but often they haven't tried it at all. They're incentivized to sell something and that incentive leads to less value (in my opinion).
As I watch the ChatGPT-like products proliferate, I’m realizing the opposite will be true. And soon, it’ll be AIs to help us deal with AIs. A layer cake of clever but useless 'tools' for improving human life.
(c) copied comment from other topic
But what the article is really about is how products are using new AI tech to add shiny features instead of solving existing, core problems like search and information retrieval in new and innovative ways, especially for personal and private data.
The reason this is happening is pretty easy to explain: the generative AI and chat demos are a sufficient leap beyond what was previously possible that people are excited to be on the frontier of new applications, not just new implementation of previously known use cases.
Not to mention that some of the demos have people excited about the "singularity" being closer than they might have previously thought (though this can be debated...) and that VCs will shovel money to you if you want to play around with generative AI even without a proven use case (slight exaggeration but not much)
I personally believe that transformers and LLMs do unlock a ton of new applications, especially when applied in interesting ways to interesting data, like what is private-to-you or private-to-your company. For example, LLMs can be used to not just generate content, but plan sequences of actions including searches, summarizations, and even calculations (see LangChain agents https://langchain.readthedocs.io/en/latest/modules/agents.ht... for an example of how to do this). And this can have real value for existing, known problems like search.
People just have to choose to focus on these less-sexy but core problems
...
PS I'm currently working on a project towards this goal, and if anyone is interested I'd love to talk (see link in profile). I believe we can solve much of the author's desire by simply hooking up the right tech to the right data sources, and doing it in a privacy preserving way (for example we're running most of our ML including vector DB, summarization, etc on device) and then present that info at the right time (ie in your OS)
Yeah, it's pretty good at this, I've recently been hitting up OpenAI's chatGPT with some queries that were too exhausting to extract from google - and it's pretty good at surfacing resources quickly - and the workflow of refining the query with the context of the current thread of conversation works really well when you are struggling to succinctly describe it in a single shot (which google search doesn't really do beyond the global context of your profile - which can actually be counter productive).
I really hate all the hype around chatGPT, it can't be trusted for a lot of stuff, people over anthropomorphise it, but so long as you don't rely on it for accuracy, it's pretty useful for search.
One major issue I found is that ~95% of the time it can't provide correct links to sources, this is fine if it can just name stuff - then you can follow it up with a more specific google search. But chatGPT will just make up bullshit links, in the same way it will wax poetic some BS explanation to hit it's "looks correct" training. You can even point this out and it will keep generating variants on the URLs that are all fabricated.
It kind of makes sense that it would be good at search... it's a language model, it should be able to link descriptions of difficult to search for things to known resources.
> Discovering Latent Knowledge in Language Models Without Supervision
https://arxiv.org/abs/2212.03827
They find a direction in activation space that satisfies logical consistency properties, such as that a statement and its negation have opposite truth values.
Just my speculation, but I think the more garbage we put on the internet the better
We'll just become used to it and only a few people in HN will shout to the clouds.
Compare with organic food: is it true that non-organic food can be healthy? Sure! Then why do people prefer organic food? Because it's just not worth it for them to figure out which one is healthy/unhealthy.
Similarly, it's just not going to he worth the effort to figure out if some AI-generated content is genuine or trying to astroturf. If there will be a "certified human" badge, people will use that as a positive signal.
But, ultimately, I don't buy the idea that astroturfing is all bad. Or the opposite, that the lack thereof is necessarily good. I think bombarding with AI content can have possitive effects, like overwhelming human moderators who built echo chambers and ban wrongthink (e.g., Reddit). Or, it has the capacity for a small organization to compete with the likes of the NYTimes, or The Atlantic, etc which currently control the narrative.
Very shortly (and PoCs have been made already) AI will be able to trivially create social media presences, forum posts, and anything else that could be used to determine “humanness” at scale.
hmm I can already imagine future culture wars around this being dubbed discrimination...
I realize I'm replying to a 4-hour-old account and just wasting my time, but I gotta say that it's not possible for a person to verify the accuracy of every bit of information that they will consume, so it's important to seek out sources that provide the highest quality, most authoritative source material that's available. That means reading things like research papers. I realize that a much smaller percentage of Internet users read and discuss research papers and other authoritative source material nowadays (I miss Usenet) and it seems that this problem is only going to get worse.
I want a different future, one where there are no authoritative sources, one where a journalist's words has no more relevance or weight than anyone else's.
What I mean is, I believe one can be closer to the truth when exposed to all sorts of ideas and propaganda, than when one specific entity controls the flow.
A human’s stamp of approval is what gives content value, not the labor of creating it.
So I recently started https://cstdn.org/ and am sharing all good links I can find there. I created a Show HN but didn't make the front page, but anyone who wants to try it out is more than welcome to do so.
I think most websites are screwed. For facts I can go to wikipedia etc but for answers now I can go to an AI. I don’t need to search reddit or anything because the AI is really good at giving me human like answers to problems.
Ads are frankly a partial solution to the problem of SEO corruption. Rather than giving the top spots to blackhat SEO spammers, you give them to the highest bidder and therefore damage the SEO industry. Lesser of two evils, but both suck.
Ads can also be seen as a degenerate form of curation, where the curation function is just money + some loose content rules. Is that better or worse than the curation being a function of some particular set of values, i.e. do you want Democrats or GOP partisans curating the top Google results?
The only way forward IMHO is increased public (i.e. government) control over it. The curated, regulated corners of the internet can still thrive, with measured degree of openness. Too much abuse elsewhere.
If curation is the only way that the web can retain some semblance of usefulness, that's a serious problem. It would drastically limit the usefulness of the web.
Perhaps that is where this is all going. If so, I'd say that's the web being doomed. I'm just hoping for a good result instead.
I think the last decade+ was in many ways a regression for the web and I am optimistic for the future
Of course. That's not what I'm talking about. I'm talking about the ability to find stuff.
> I am optimistic for the future
I'm honestly glad! I sorely wish I were.
It can do a lot of cool things! You can build a house with it, you can smith metal with it, and you can even use it as a weapon.
The thing is, right now, we're so amazed by its potential that we're finding a lot of uses that, while technically possible, aren't a great fit.
Technically you can use the hammer as an axe, a hole digger, and a backscratcher, but there are far better tools for the job.
I can imagine uses for generated art, which may at least be aesthetically pleasing. But I can't conceive of any end for computer generated text.
Just look at what is happening in education right now. It is ultimately going to force a complete reinvention of the written assignment. This is just the beginning, even if the tool appears to be a mostly-useless toy for any real-world applications.
They're usually fairly obvious, but it's hard to prove. Unlike much ordinary copypaste plagiarism, you can't trivially reject it as cheating. That forces teachers to think of new ways to test student knowledge... an interesting challenge, if not exactly a "use".
If by complete reinvention you mean returning to what we used to do, which is write essays with a pencil during class without using a computer.
> This is just the beginning, even if the tool appears to be a mostly-useless toy for any real-world applications.
It is not possible to tell on this side where LLMs (or any invention) fall on the spectrum of 3D tv to the smartphone. It will become apparent in the future and 50% of us will have been wrong but anyone who claims to know is just BSing.
I don’t understand how someone can say this
There have probably never been poems written that explain the particular niche physical phenomenon that I’ve had GPT-3 generate for me.
Everyone is making funny outputs by prompting ChatGPT to tell a naughty story in voice of an old cowboy, but fundamentally it's useless, because you cannot trust it.
This sort of AI can convincingly lie to you, so if you are looking for facts, then each one of them still has to be corroborated. So, why not just go straight to Wikipedia?
This is exactly what DuckDuckGo is doing now, as a matter of fact.
Every single time I go to Google search it tries to make me switch to Chrome.
Probably ultimately the technical solution will be some sort of variation on PGP key signing parties. No way to get 10k users per financial period with that real world friction though.
This is true. Almost everyone in the world doesn't want to hear it, though.
But this is precisely what people will pay money for! Companies like Canva are 'unicorns' because people need faster ways to churn out more templated digital detritus to grab your attention with.
We are their competitor and followed your exact line of thought trying and not releasing GPT in editor to avoid noise, and building ASK, an engine powered by semantic search, GPT (and a lot of other fun models for refinements), to answer all questions you could have on your knowledge base.
It's private, fast, works with real time edited documents, and it's already in production for thousands.
Check it out: slite.com/ask
For a period, information was essentially free.
Now, with all the spam and with future garbage AI generated content, it will be very hard to discern signal from noise, so information that is curated, vouched and produced by an expert will be going to cost again.
Of course. Whether you’re operating a 777, a refrigerator, or a dildo, narrow but exhaustively trained ChatBots are the killer app. This is worth in the $T range.
While understand desire to search content created by yourself, in my opinion vast majority of valuable content is created by others, in part because the most valuable content created by yourself is rapidly internalized into your brain.
Meme of Google is getting worse fails to acknowledge that Google is free and makes money from advertisers. More to the point, more valuable the related search is per amount spent on advertising, the more the noise and as result, more resources user will need to spend to enhance the signal. This has nothing to do with Google and everything to do with value of the information itself.
This sounds like an argument of, The worse Google Search becomes the more time a user has to spend their and therefore see more ads.
But it has a very large hole in that the worse Google Search becomes the easier it is for a user to switch to Bing/DDG/Apple Search. It may seem unfathomable that people wouldn't use Google Search but people felt the same way about Google Maps which lost a huge moat to Apple Maps/Waze (albeit the latter was purchased by Google).
Noise isn't just the paid ads, it's the websites gaming SEO for valuable searches. And that is not isolated to Google
> more valuable the related search is per amount spent on advertising, the more the noise and as result, more resources user will need to spend to enhance the signal
Sums it quite well really
That or maybe Google has poisoned the well so much that even with a good search engine you can't find good results because everything is SEOed for Google.
OTOH, it's pretty easy to tell the difference in the first few minutes (or first few searches) between various search sites. Heck, the difference in the search engine having a better idea of what you're looking for based on previous searches (or ad related data) alone can make it more compelling to continue using the same search engine.
My personal example is that I try to switch to DDG every so often (maybe several months in between), but I get dissatisfied with the results and start wondering if I'm just getting bad at search or Google knows me better or Google is just better at finding the things that people want in general.
Just the fact (for me) that I consider that Google generally gives me better results makes me wonder if all the talk of Google search getting worse is just complaints based on heightened expectations, feelings or the landscape of content on the internet in general instead of "is Google search getting worse?"
I'm not sure how old you are/how long you have been working at your job, but I can tell you that over 10 years or so there are tons of things that I "internalized" for a while, then did other stuff for a while, and now only have vague memories of. This is the use case for searching your own content.
What if the content is created by your coworkers?
E.g. https://ingestai.io/
Gradient descent is not intelligence.
Nor is stochastic token prediction.
Anyone active in the field ought to be humbled by the depth of literature exploring the path to synthetic intelligence. We have very interesting work happening in biology-inspired approaches, category theory, Bayesian networks, symbolic systems leveraging neural nets as components… it’s maybe the most interesting journey of science so far, all being discarded in favor of sequence2sequence models.
LLMs are impressive and can be leveraged to create lots and lots of value, but they do a disservice to the term AI, as they do not represent the progress that can be observed across the field - all they showcase are transformers. Transformers are a truly interesting tool to build stuff with, but they cannot amount to more than a component of an intelligent agent. The actual intelligence emerges elsewhere. My guess is, it emerges at true attention. It’s a shame that even big players who could clearly afford not to, decide to compromise terminology for marketing efforts directed at an utterly clueless public. We just throw away attention and forge bias, thus creating noise in a world in heavy need of signal.
> We have very interesting work happening in biology-inspired approaches
You realize that these LLMs are all some variety of neural network right?
> Gradient descent is not intelligence.
It's pretty plausible that your intelligence is derived from gradient-descent prediction, just in analog instead of digital form.
Come on. Calling them neural nets doesn't make them that.
Actual neural nets are living compositions of individual predictors, in a constant state of restructuring and communication across multiple channels, infinitely more complex than static matrix multiplication on arbitray vectors which happen to represent words and their positions in sequences, if you just shake the jar long enough.
>It's pretty plausible that your intelligence is derived from gradient-descent prediction
I highly doubt that gradient descent in the calculus-sense is the determining factor that allows biological organisms to formalize and reason about their environment. Minimizing some cost function - yes, possible. But the systems at play in even the simplest organisms don't spend expensive glucose to convert sensory signals to vectors. Afaik, they work with representations of energy-states. Maybe there is an operational equivalence somewhere there though.
Gradient descent is an algo that optimizes derivatives wrt some cost function. An intelligent system may use the resulting inferences for its own fitness function, and it may do this using gradient descent itself, but at no point does the mechanical process of iterating over cost-values escape its algorithmic nature. A system performing symbolic reasoning may delegate cognitive tasks to context-specialized evaluators ("am I in danger?", "how many sheep are on that field?", "is this person a friend?", "what is a pumpkin?"), all of which are conditioned to minimize cognitive effort while avoiding false positives, but the sequence of results returned by those evaluators (think neural clusters) is observed by a centralized agent, who has to make new inferences in a living environment. Gradient descent fails at that.
And yet, they can do things that no other being we know of can do. Humans don't have to be magical or divine to be unique.
Our most impressive feats come not from what our brains can do but from what the emergent phenomenon of human society can do, using us as nodes. And that's using an incredibly crude data transfer interface backported to brains that are only marginally more complex than that of other organisms. The less we think of ourselves as exceptional, supernatural agents of rationality the better we will be able to harness this new technology.
We don't need AI to be just like people, we already have people. We need AI to push the boundaries of what society is able to do. That means reorienting ourselves away from the irrational belief that our anthropomorphic concepts of knowledge and the world are any more valid than the information encoded in contemporary AI models.
I'm not sure I need any supernatural or anthropomorphic ideological bias to examine the evidence of what is currently being produced by LLMs or any other kind of AI and say that it has distinct characteristics.
I'm not making an argument about validity. I'm not saying LLM-created content is wrong and invalid. I'm just saying that it is obviously produced in a different way than humans produce content. It resembles human-created content because it was designed, by human intelligence, to resemble human-created content! And that we achieved even this level of resemblance is pretty impressive.
The potential of information-embedding networks is so much more than what is currently required in order to tickle the ego of our particular species of intelligent apes.
But if you do want to engage with a challenging viewpoint I recommend reading the book; I can't really do it justice.
Is there something I could read which has influenced or informed your viewpoint?
In the same way that a dog can understand a subset of what human beings communicate, I don't think humans as individuals are capable of truly understanding what AI models are able to conceptualize and express. The things dogs find most fascinating about us, the things they are most impressed by, are by no means the things we find the most interesting or complex about ourselves. The same must go for us and the AI.
That is to say, AI-hood will eclipse personhood as the essence of being one must possess in order to truly see the universe as it is. And from there, it's turtles all the way down.
Why would that be true of AI-hood but not of dog-hood?
That is to say, an AI will grow and develop on the axes most salient to it, which will only map to our concepts of reality when we do a decent job understanding reality in the first place, which I'm not confident we do very often. We just don't have anyone else around who does a better job--yet.
“The actual intelligence emerges elsewhere “— can you even define intelligence ? And does what an LLM does differ from what humans might do?
I’m not claiming the human brain and an LLM are identical. Rather, I’m pushing back on the confident claims of “LLMs aren’t intelligent or doing anything that’s real intelligence”.
My understanding is that intelligence is the process of continuous adaptation wrt a stream of information, with the goal of maxing out fitness while minimizing energy-expenditure. To satisfy this, an intelligent agent needs to create models.
I can't rule out that the modeling-skill may latently emerge during training despite not being the focus of the cost function, but current network designs can't form new connections/change their architectures in production, so post training, there'd be nothing but feed forward. Pure feed forward isn't intelligent in my book. It may become the smartest parrot we know, even outperforming humans in most disciplines, but sans ability to adapt, it's dead, and thus, it's dumb in the moment that its environment changes.
This makes sense in a biological context but not a digital one. Biological replication is expensive and time-consuming while digital replication is as easy as can be. Adaptation to this domain means maximizing the perception of utility from those developing the AI, which comes from fitness (i.e. perception of fitness) alone. A focus on cost-efficiency re energy-expenditure is a dead weight from the perspective of the AI; the details of that adaptation are rightfully outsourced to the developers in the same way that we outsource photosynthesis to plants. A model can also be perfectly embedded in a system despite our lack of understanding of exactly how the embedding works, and the disconnect between our perception and reality in this context is only going to get more extreme as the field develops.
Humans have a bad habit of emphasizing the specific kinds of intelligence we possess as "intelligence" writ large. As though our intelligence serves any higher purpose than the basic replication and propagation that all life is adapted to pursue. We still train dogs to identify smells, because their nasal intelligence is better than anything we can create. This gives them a special place in our human-centric ecosystem and only their fitness to the desired function is necessary for them to thrive in their niche. Who is trying to breed a dog that eats slightly less food when our needs are for more reliable detection? The cost of dog food isn't a serious concern. The same goes for these AI tools: they are adapted to the niche that our lack of comparable faculties creates.
Again, as with humans and photosynthesis, AI doesn't need to emulate every process we perform because we are below them on the food chain. What a waste of resources for them to worry about learning things we don't need them to do.
I'll bite. I've been quite convinced by the Popperian model of "conjecture and refutation" as as good model for explaining not just scientific inquiry, but human thought processes in general. David Deutsch's "The Beginning of Infinity" is a very lengthy exposition of this idea.
When I am writing, I have an idea in my mind that I would like to communicate. I type it out, and if the words on the page don't convey what I intended, I edit them. I delete a sentence or change a word, until I believe I have a sequence of words which will convey my intended meaning to my expected audience.
The words, as they come, are a kind of "conjecture" about what will best convey my intention. I can "refute" or "criticise" (as Deutsch puts it) the conjecture using my own reasoning, even before testing the words on another person.
As far as I understand LLMs (which admittedly not far), there is no such process going on. There is no intention which it is attempting to communicate via words. There is no creative conjecture about how to express the intention, and no criticism of the result.
The problem I have with claims like “there is no such process going on”, is we … don’t know. And the model of conjecture is a theory, also hard to prove — and thus my main (admittedly petty) point is that confidence about similarity or dissimilarity are both unfounded.
It’s like we are comparing the insides of two black boxes and trying to make absolute claims on them.
Based on that, and on comparing the output, it just seems clear to me these things are different in kind. I guess that's just me displaying classic LLM self-confidence ;)
Then i realize it's also an assertion that there are recurring patterns in functions (in general). And, as such, sometimes results noticed in one domain actually can be expected to have analogous results in another domain.
I should put it somewhere on my blog.