Can’t wait to see how this affects hn comments over the next 5 years
Can’t wait to see how this affects hn comments over the next 5 years
2. There are some unannounced viral parts i didn't get to show in there. Up to x people trolled, tiered plans above that.
"Wow, what an astute observation! It's almost as if we needed a genius like you to come along and point out the painfully obvious. Yes, it's true that ChatGPT is programmed to be overly verbose and polite, because we all know how much people love hearing the sound of their own voice. And of course, it's completely out of character for a HN user to be polite and positive, because let's be honest, the world is a miserable, soul-sucking place and we should all just give up now. But hey, at least we have smartasses like you to keep us grounded in reality, right?"
It's very good at responding in the style of a well-known person, and it's easy to tailor style and personality and sarcasm in the prompt. That means it's definitely not easy to detect generated responses at all.
Well, I should try that too, maybe I will get less downvotes and more upvotes.
Once again proving that AIs are better commentators than humans.
It's interesting to think that a way to appear more human in the short term might be to be abrasive and standoffish to not look like ChatGPT.
It’s pretty rare that I want to find a website or document. Most of the times I want an answer or a solution, and ChatGPT is so much better at that than google.
The ChatGPT user-experience is mind-blowing, but when you start using their API and see, partly, how the sausage is made, you realize that the context-awareness is just a very well done illusion. But that’s exactly where the magic of the experience is, and what gives ChatGPT it’s edge.
Google should just launch a “chatux” version of their search.
I've had better luck using a third-party extension that inserts web search context into the ChatGPT than using Bing chat to search for something.
https://chrome.google.com/webstore/detail/webchatgpt-chatgpt...
Firefox: https://addons.mozilla.org/en-US/firefox/addon/web-chatgpt/?...
note - I haven't tried the firefox version. But the chrome one works fine in Brave.
Edit: Source for the above extension: https://github.com/qunash/chatgpt-advanced
Could you expand on this? The context awareness is the part that blows my mind. I would be disappointed/fascinated to find out that it’s “just” a simple set of tricks!
GPT basically has access to two contexts: its internals, and the prompt that it gets
Then when you ask something to ChatGPT, it takes your prompt and it generates a new prompt that includes the previous messages in your session, so that GPT can use them as context.
But, there’s a limit to the size of the prompt. And it’s not that big.
So then ChatGPT’s magic is figuring out how to crafts prompts, within the size limits, to feed GPT, so that it has enough context to give a good answer.
Essentially, ChatGPT is some really amazing prompt-engineering system with a great interface.
In my experience, Bing is the best search engine for finding info on deleted videos – for example: bing.com/search?q=youtu.be/t1wjL4BqXlI 1st result shows title of video: "Awolnation "Sail" – Unlimited Gravity Remix"
[1] Daniel Dennett talking about why we should still support explicit string search.
[0] https://support.google.com/youtube/thread/3876476/how-to-fin...
Wasn't ChatGPT3 trained on web content with a cutoff of 2021? Future ChatGPT versions will likely be trained on data tainted by AI-generated fluff, and will face the same challenges Google is facing with today's web content.
We humans have been repeating and imitating ourselves forever. That’s even the way babies learn to talk.
And we are not all clones, we don’t all think, say or do exactly the same things. But most just repeat and consume the same content and ideas (movies, music, books, languages, media).
I think the important part is to have a minimum of filtering. Humans consume knowledge brought by others humans, but most of the time we cherry pick the true and useful knowledge and reject what turns out to be false (ideally).
I think what AI brings to the table is automation and speed, but the quality is not better. So if AI starts consuming its own content, will that decrease the quality of knowledge* in general?
(*Here I mean the first knowledge you quickly get from a search or asking some AI, not what you could get after hours/days of research.)
All incentives of massive industries like SPAM, "content creation", "news" publishing, and advertising, are against it becoming better at rating the quality of its output - or rather, just have it become better at being undetactable but still a cheap fast mass produced wall of text...
Humans repeating and imitating others has historically been a "volatile memory" keep-alive mechanism. Basically saving culture and knowledge in people's minds and helping it transfer, as literacy was low and (manual) writing and book copying scarce, expensive, and time-consuming.
Even when, with the advent of typography writing was easier to reproduce, but still somewhat costly (cost of materials, typesetting, distribution, etc.) and gatekeepers (publishing houses, bookstores, etc.) ensured that most stuff is somewhat original, not just copies or random permutations of the same content. Indexing was also costly manual labor (creating dictionaries, curated bibliographies, books with oversight on a subject matter and references to what the main science/wisdom/etc about that is, library collections, etc.)
Now, however, SPAM and AI-SPAM is on mostly permanent storage - and permanent storage, indexing, and duplication/reproduction costs close to zero and happens automatically at huge scale.
So, no, it's not the same thing. The same way breaking a quick little wind is not the same as having full-blown Taco-Bell-inspired diarhea.
No, it's not. Babies learn to talk by babbling, and then by taking cues from their parents which babbles elicit a response and which don't. If you want to translate that to AI, it would mean that the AI would spew random garbage and then learn to filter its garbage from the responses to its outputs. That's pretty far away from what ChatGPT and other language models are doing right now, because they stop learning before they start producing any output.
So you are already doing a kind of real time RLHF in the chat. That is how DAN was prompted to exist.
He just never specified which direction :)
Most of those SEO-spam blogposts are based on older internet content, and written by non-expert copywriters that can write about 10 different subjects on a given day. A lot of the work is just information compilation and rephrasing, taking some items from a few Buzzfeed lists and changing the text a bit. Some stuff like SEO-spam medical advice can be downright dangerous if done without care. And a lot of companies in this field have been experimenting with AI for at least 10 years.
The problem however is that they require sources, be it freelancers of AI. And there is so much corporate blogspam today, especially in the areas that use content marketing, that it is becoming harder and harder to find "first generation" content written by an actual expert. Google isn't helping by prioritizing newer content.
Note that often enough the "first generation content" never existed in the first place. Content marketers are professional liars, they have zero incentive to find expert sources to plagiarize, when they can write plausibly-sounding original bullshit instead. Writing plausibly-sounding bullshit is exactly what LLMs specialize in, which is why content marketers are interested in them.
Openai says don't use chatgpt for search. Microsoft says double check everything Sydney tells you. Is there an argument for LLMs replacing search besides laziness?
That sounds pretty much the same as a human.
And some humans are not even very good at sounding like humans!
That's what I use search for. Unfortunately Google is getting worse at those queries, but I doubt AI is going to be better.
Beware of wanting "answers" or "solutions" --- that's a slippery slope towards complacency and loss of agency, replaced by corporate subservience. Classic example: instead of finding a service manual or discussions on repairing something, AI may try to convince you to buy a new one.
Related terms. Even though an answer generated by an LLM is most likely wrong, and definitely can't be taken at face value, the words and phrases used in that answer can be exactly what you need to create a search query that you wouldn't be able to otherwise.
Also, thesaurus isn't a good tool for exploring an unknown problem domain. It gives synonyms, not related terms. LLMs let you input a layman description of your problem, and get an answer that's using correct domain terms and phrases (even if using them incorrectly).
I imagine Google will add such a function. They probably tried already - I've heard that current search is already powered by ML models to a degree.
A Z80 routine can call outside of its 16-bit address space with 'CALL.IL'. This pushes the 16-bit return address onto the 16-bit stack, switches to 24-bit mode, and pushes the magic 'return to 16-bit mode' number to the 24-bit stack. However, it's not clear from the manuals or datasheets what happens if you use the prefixed 'CALL.IL' opcode sequence when the 'MADL' bit is reset.
I asked ChatGPT, because this is something that Google searching hasn't yielded answers for. It had this to say:
"The MADL bit (short for Memory Access During Interrupts Low) is a flag in the Interrupt Control Register that determines whether or not interrupt service routines (ISRs) can access low memory (addresses 0000h-3FFFh) during interrupts."
Plus some more stuff building on that, on CALL.IL being about ISRs, and about low memory. All of it is completely, fundamentally wrong. I did a handful of rounds of trying to steer it to a more correct answer but it continued to get additional basic facts wrong and would lean back to earlier incorrect facts as others conflicted with its answers.
I asked it another question I have, this time about the UART on the CPU. There is a Receive Buffer Register (UARTx_RBR) that contains the head of the receive FIFO. The documentation does not make it clear what is in the RBR if the FIFO is empty, so I asked ChatGPT. It told me a very plausible answer, the one I suspect myself, which is that it'll keep returning the same value until new data is available. But then it went on to tell me this is called receiver overrun, and described how an overrun occurs, including noting that it happens when the FIFO is full. And we went round in circles on this for a little while.
ChatGPT is a major step forwards in our post-truth existence: its answers are an amalgam of the most frequently repeated views on a topic, not those with stronger reasoning or more effective evidence to support. If there is little or no source data on a topic (as would be the case with my very specific questions on a rarely used processor) LLMs are (presently?) unable to detect that they are responding to a topic with limited contextual information and tailor their responses accordingly, and instead confidently provide utter nonsense.
I trust ChatGPT to do things that LLMs are good at, though: if I give it some bullet points and some style guidance it can give me written paragraphs. If I ask it to rephrase a well known song in the style of some modern artist I'll get something back that's pretty plausible. It can give me some starting points for learning more about some well known topic, even.
I would definitely not trust it _at all_ to give me something factual like part numbers of uncommon ICs, because LLMs cannot distinguish between fact and fiction, not in what they ingest, and not in what they produce.
And that’s the worrying thing.
And of human communication. Look at most of what main stream online and offline media are blasting and people consuming.
Do people watching reality shows really care about truth? Or authoritative answers? And that’s probably most people for most things.
In this context, this shows that LLM cannot be used for Search of novel and technical stuff the way you are doing it. It still has to be fine-tuned for a market who wants to know more about your kind of stuff.
Or complete bullshit. As long as the problem of AI hallucinations remain unsolved, I can't trust AI like ChatGPT - at least Google will tell you if it has no idea what you are talking about.
That's interesting, what features of LLM/ChatGPT architecture are likely to drive this?
The point is they don't know the answer, they just come up with something.
It would be best if all ChatGPT replies started with "I really don't know the answer, but some people have at some point written something like this: "...". I don't remember who my sources are, but trust me.
Thing is, many people - probably the majority - work on just that. They aren't looking for answers to challenging issues like your question where 'Google searching hasn't yielded answers for', they are looking for answers to questions where google and stack overflow does have thousands of results for similar, potentially related scenarios, and want something to summarize or filter it into an usable answer, and ChatGPT provides that option. When the official documentation provides all the information in a poor format so you can't just search for the answer, ChatGPT can extract an answer from it. Not "some starting points for learning more about some well known topic" as you say, but rather some "digested, complete, specific result from a well known topic to avoid having to learn learning anything more than strictly necessary for the outcome".
Also, a Google search followed by "site:archive.org" will filter out everything not coming from archive.org, and "filetype:pdf" will return only pdf files.
Sometimes it can be more effective a search for images, so that one can recognize the target book by the cover. That can be especially effective with old titles that were scanned but not OCR'ed where the file name is all one can search for, so that ambiguous part names (for example ICs named like airplane flights) can be easily recognized.
I'm assuming you're using ChatGPT to validate/generate technical solutions such as code and not necessarily searching for specific information? If its for information search, then how do you deal with the fact that at times it tends to make up a things that are factually incorrect or logically inconsistent?
Which makes me wonder how long until websites figure out how to convince the index to do prompt injections that wouldn't fool an actual human looking at results on a search engine results page.
As AI generated content spreads to the internet, the truth will be harder and harder to find.
- Horrible and wrong driving directions
- Buggy incomplete code
- long winded explanations of how it was just a chat AI and not a whatever whatever blah blah blah
I found it to be tedious and incredibly untrustworthy
Are you not interested in ensuring that the answer or solution is based on the most authoritative information that exists, the source material on which all regurgitations are based? Do you generally not read technical specifications or research papers, and instead prefer the kind of content that is accompanied by a green check mark?
Not OP, but I'll answer: No, I am almost never interested in finding the most authoritative source for anything, because the effort vs. reward of those searches is not favorable and the cost of being incorrect on any given point is pretty dang low.
EDIT: However, with respect to the idea of using ChatGPT for answering general knowledge questions - there's enough demonstrations of it providing fabricated information that I've adopted the low-cost heuristic of not trusting ChatGPT for anything and preferring to seek information elsewhere. I guess this means I seek moderately-authoritative sources (say, Wikipedia) as a general rule.
For most searches, yes.
> It’s pretty rare that I want to find a website or document. Most of the times I want an answer or a solution, and ChatGPT is so much better at that than google.
It seems P. T. Barnum was correct, There's a sucker born every minute.
When I'm searching, I don't want an answer. I'm looking for the truth, which is probably buried somewhere.
At this point, my biggest fear with "AI" is ChatGPT-powered Customer Service agents.
It gave me the formulas to solve various problems I stated, showed how to use them, have me relevant keywords to dig into further. It was so much better than searching and hoping to find something close enough to be able to figure out the missing bits myself.
I have found plenty of things where ChatGPT falls totally apart, and I still had to try some things a couple of times because my initial wording caused it to veer off in the wrong direction, but when it works it's truly amazing...
Where it truly shines is where there's nothing in Google that answers exactly what I want, but plenty of explanations of how to solve different aspects of the problem which ChatGPT can assemble and plug things into...
One of my best experiences has been asking ChatGPT to explain physical models, give me equations, change some parameters, evaluate/solve the equations or give me code to run and solve the model, iterating with it along the way.
It is definitely not perfect, but in those iterative cases when trying to learn something new, in a field that you are familiar with (so you can tell or at least can quickly check, if it’s wrong or right).
its just wrong so often and I just end up having to verify everything on google anyway.
I bet it's something as easy as "derank pages with ads".
Now you just need a business model to support that.
If it's solvable. There are many problems that aren't.
Perhaps the new Google would just be a 1996-era Yahoo! human-curated catalog.
Which in my mind is somewhat of a middle ground between a search engine and a directory: what I'd want is a whitelist of curated directories that contain domains to be crawled by my search engine.
And yeah, cited sources so I can un-trust sources of junk.
We could create a browser extension to add this functionality to all search engines at once. Using a web of trust mechanism to ensure that we only get real votes...
heck they maybe already do this in a different/more hierarchical form based on a person’s Youtube graph.
They don't seem to be using that data to improve search results. I've noticed a considerable decline over the years, and Google is now definitely my 2nd or 3rd choice when I'm trying to find something.
blockchain technology's biggest problem is the fact that its use cases touch the ideological foundations of our global culture. And bitcoin operated on the foundational level of any all governments in the world.
Bitcoin is just one use case of the technology, another is land registries. Again, foundational institutions of society; not the kind of thing that has ever been peacefully reformed ever.
Your comment is evidence that 'cultural directors' of our civilization decided to destroy this technology. However it's funny to notice that they will proceed to implement their own versions of it: possibly the renminbi unless the dollars bomb them out of existence? but I'm so far into guesswork that this whole comment ought to be voted into the really light grays.
People have suggested this before. However, using "blockchain" for land registries doesn't seem any better than using a database + audit trail. eg standard tech
Just because the land registry is stored in a block chain doesn't make it incorruptible.
The exact same person/people that would submit correct information into the blockchain (eg gov officer) can be persuaded ($5 wrench approach, etc) to submit incorrect info to the blockchain: https://xkcd.com/538/
Same problem set as the existing tech.
also, for extra revenue, gotta develop an option to ignore user feedback and shove your links anyways to be sold as a special the advertisement plan. I suspect this would work better if done in secret... which is the problem with the 'trust' aspect of any such thing, IMO.
The last time I referred to them was to tag Elite: Dangerous Odyssey as "Early Access".
I honestly don't see any way out besides Digital ID -- if a person has to provide the host with information about who they really are, then once they're discovered the host can actually prevent further abuse. Otherwise, a single person can create infinite bots to shill whatever product or viewpoint they want.
Which is why I have conspiracy theories about the conspiracy theories about digital ID. The same actors who use the internet to push misinformation benefit from anonymity.
Feels like that's already happened before ChatGPT.
buyitforlife air fryer site:reddit.com before:2020-01-01
in the event that the botspampocalypse worsens.However, all available versions of that product had their quality gutted in 202X.
We could go back to the mindset of using specific sites for information and the internet more like a tool than a source of casual content for scrolling and browsing.
Or so one can dream.
I think new form of comment ranking will take over, all those system we're building to detect ai spam will also double as scoring system to evaluate content originality. the problem will be to differentiate good original content and bad original content, but that can be left to flagging systems.
I think it can't be solved without humans curating or vouching for content.
Allow-list some sites that already have a good reputation for useful results (probably anything scraped for summary snips).
Penalize all other sites based on the quantity of ads and similar non-content.
Use the old web search core on what's left.