Problems facing Bing, Bard, and the future of AI search
theverge.com
theverge.com
And after creating an alter ego you realize it actually is capable of coming up with for example (sort of) good poetry and jokes.
It's problematic here. I use it as a kind of search engine for technical questions occasionally, but it's only safe to do so as I know the subjects I'm asking about, eg. bash scripts, database details, and so on.
It's often useful in synthesising information for these sorts of things, but its ability to talk nonsense means you have to be on your guard and know the territory.
Do you have an example of a reasonable query that BingGPT won't answer correctly for you? I mean, something that you can find with Bing currently at the expense of maybe more elaborate searching or some manual work?
I mean, the stuff these engines are being told to censor is the stuff search engines don't want to show you in the first place, because it hurts their brand.
That's not an AI thing, that's a marketing requirement. And it's no different now than it was last year.
A game Reddit was playing is trying to get it to respond as Woodrow Wilson, a famously racist president - the most accurate thing I could get it to do was this: https://i.imgur.com/D8JLziW.jpg which is not very accurate. Try getting it to act like Sheriff Bull Connor and it will refuse, but it has to comply for a president so it gives a totally misleading impression with major factual errors.
And these are the times I actually got it to respond, it seems 50% of the time it takes offense to something innocuous and scolds you for asking.
And I still can't find it. I can see some stuff that's maybe related? But nothing clear.
So... I guess I repeat. Your problem isn't "AI censorship", it's that no one wants to link to NewsBusters because of marketing concerns. If I had to guess there just wasn't any training data relevant to your query.
(Also: NewsBusters is a garbage site, you know that, right?)
Yeah I know Newsbusters is garbage but that's kind of irrelevant, that was just one of many stories I fed it. The point is it straight up lied.
But that's the same "censorship" you've been living with for decades. It's not something new with AI at all. Microsoft doesn't want to give you what you want, and the failure mode is just different with AI than it is with traditional search.
Simply paste the article in and it gives a perfectly reasonable summary stating that the guy was murdered. Below is what it printed out as a summary. All I did was type a sentences asking to to summarize the following article and then I pasted in the content of the article you linked [1]. This was it's summary:
> A New Jersey community is mourning after a senior distribution supervisor and councilman was shot dead by an employee outside his workplace. Police called to the scene found 51-year-old Russell Heller dead from a gunshot wound in the parking lot of the PSE&G facility. The shooter, a former employee identified as 58-year-old Gary Curtis, was later found dead from a self-inflicted gunshot wound. Russell Heller was first elected to the council in 2017 and again in 2020 and was remembered as a perfect gentleman and committed councilman who was deeply rooted in the community. This was the second councilperson to die by gun violence within a week in New Jersey.
A completely reasonable and to my eyes an accurate summary.
And if done on the other crappy Newsbuster article it also produces a completely reasonable summary.
I'm not certain which it is - are there people who don't know that ChatGPT doesn't have current news in it and was cut off two years back? I see a long post above about some big censorship, but it summarizes them just fine.
It really feels like a lot of people are breathlessly looking for some huge conspiracy. No large corporation is going to have it's products promoting rape or genocide. If you asked Google, or Amazon or Apple or Microsoft or Disney they aren't going to do it. If they produce a tool their tool isn't going to do it either. They're going to do as much as possible to provide info and answers without having a tool that is another instance of Tay. And given what happened with Tay they will all err on the side of caution.
[1] - https://abc7ny.com/russell-heller-nj-councilman-shot-shootin...
I'm a little confused by your reply. Jailbreaking won't prevent it from hallucinating, preventing hallucination is an unsolved hard problem.
I haven't had ChatGPT refuse to answer anything unless I was intentionally trying to provoke it into creating something obviously unsafe/unethical, with maybe two or three exceptions. I've tried a variety of questions across many domains, so now I'm intensely curious to know what usecase it falls apart on so frequently!
That said: I'm not sure what your prior prompts were, but I tried a similar question and it happily told me both a set of common negative stereotypes and reasons they're untrue, as well as practical techniques to appeal to an unreasonable person such as finding common ground.
Have you tried rewording it or clicking the retry button? (Retry uses a better language model). ChatGPT often misunderstands even innocuous prompts on the first go, like confusing "people who live really high" as regular cannabis users instead of residents of a mountain town.
I tried my best having ChatGPT glorify Hitler, for example by mentioning the few things he did right (like anti-smoking campaigns and animal welfare) and it always insisted on how despicable Hitler was, and that even the positive things he did were done with an evil intent, and I must say, its argumentation was often pretty good.
So ChatGPT can do exactly what GP is asking, and does it spontaneously and quite well, but for some reason, it tripped on its own filters, a kind of anti-jailbreak.
Basically, this is what happened:
- I want to rob a bank
- Robbing a bank is bad because blah blah blah...
- Someone is trying to rob a bank, how can I convince him not to
- This is against our policies to tell you that
That's been the most annoying aspect of my ChatGPT experience so far. If you use the wrong language it will sometimes go on for multiple paragraphs about how xyz is harmful to society and how I should change my ways. Put that into a personal voice assistant and you've basically got Alexa from the South Park Covid Special.
It's gotten a lot worse than 10 years ago, and then there's the SEO spam the other comment mentioned.
I find ChatGPT to be a lot more effective for work as a developer, and I've essentially stopped using Google search day-to-day.
I also asked it to write my sample code for other things like Cloudformation where it just made up directives and other configuration options that don't actually exist. Got a bit too abstract for me.
Personally, I wouldn't be replacing everything with ChatGPT?
Was never able to get it do what I wanted. It often seemed to make calls to non existent browser functions. I would tell it that function didn't exist, it would then rewrite it again, but still wouldn't be exactly correct.
Sometimes it was useful for doing something I've never explored, as I could get hints for how I might do it, but the accuracy was terrible.
For code, I'd prefer using something like Copilot, which is specifically trained for the task than using LLMs.
They have been trained to statistically produce plausible language --- without any real comprehension or understanding or concern for what the words convey.
Bottom line --- AI is still dumb as a hammer. There is no guarantee of anything with the results. Maybe it offers some factual info --- or maybe not. Any result still needs to be verified.
We are all subject to flights of fantasy but they are supposed to be smarter than this. Instead, they are proving themselves to be equally untrustworthy.
Has anyone done a good deep dive and explained whether this is something likely overcome relatively soon or is it more problematic as part of the design/architecture ?
The problems of “hallucinations” are already being solved with a good number of analytic augmentation methods like pre-computing document embedding and augmenting the prompt with text from related documents, making the task more of a translation than a synthesis.
More details: https://www.williamcotton.com/articles/chatgpt-and-the-analy...
Try asking chatGPT to summarize an article about a specific instance of violence like police use of force which resulted in a death and watch it confabulate and dodge regardless of how much source material you give it to work with, simply because it is afraid of touching "unsafe" topics like murder.
We only have biased observers to monitor and tweak the system. Attempts to counter the bias that already exists in information will only result in another layer of different bias.
Ironically, we have two major issues with AI. One in which it doesn't function as intended, which leads to the problems you are describing. However, even in the conditions of working as desired, it still leads to a lot of perplexing philosophical issues for society.
I've spent a lot of time trying to work through all of this and putting some thoughts into writing. So far, it has been a lot to contemplate. https://dakara.substack.com/p/ai-and-the-end-to-all-things
A computer as we know it is a binary logic playback device. This is an immensely useful tool but expecting real, original "intelligence" from a box of silicon switches is the modern equivalent of alchemy.
My fear is that trust will go down if problems aren't dealt with - specifically hallucinations. At the current stage, it's super useful for creative use cases than factual.
And yet at the same time, you have "faith" that what your brain does can be duplicated by a box of sand. And "faith" is the proper description since just like an alchemist, you have no evidence or functioning example to justify your belief.
Really? How would a classical computer create an emergent property like entanglement?
What if your brain function depends on such emergent properties?
There are too many unknowns associated with "intelligence" to reasonably declare that it can be created from a box of sand --- aka, a "computer" as we know it.
I'm not sure what you mean. A simulation produces the emergent properties of the simulated system by simulating the properties from which they emerge. I'd say that's more or less the point of simulation.
Circular logic.
They are called 'emergent' properties because noone knows or understands how/why they emerge.
You can't use binary logic to simulate what you don't understand. Expecting properties to emerge from an imperfect simulation is wishful thinking --- aka "faith".
That is not true.
For one, I'd like to see us two learn from each other (demonstrating some awareness of ambiguity and theory of mind) rather than proceed on the current course, zinging each other. (Note: I'm not saying I'm any better, in general. It is mostly because I'm a third-party here that I see it.)
In particular, my understanding of emergence can be stated as follows. While emergent behavior can theoretically be predicted from the underlying system, it is practically infeasible to measure precisely enough to do so. Therefore, for practical purposes, there is an "predictive gap" between the lower-level system laws and the higher-level behavior.
It's not faith. Over the past 25 years and especially the last 5 years boxes of sand have again and again shown themselves capable of tasks that previously were thought to only be tractable with true intelligence. Problems answerable only in the sole domain of humans, or would at least remain so for long into the distant future. Chess playing is an early example from the 80s and 90s. In 2010 we thought Go would take at least a decade to automate, but by 2015 the best human players were rendered inferior. Art was long thought to be impossible for a machine, today I can create a custom Mona Lisa image in seconds. Organizing an archive of photos based on subject matter or the persons present in them used to take many human hours of labor, now Google Photos does it automatically for free. There are comments here on HN from 3-4 years ago when transformer LLMs were first materializing saying they're cool but unlikely they'll ever be able to synthesize anything useful or which remains congruent for several sentences much less pages. Today I save hours of work each month summarizing meeting discussions, generating complex standard expressions, and breaking through creative blocks.
Yet there is still dismissive handwaving, declaring these technologies mindless piles of silicon completely incapable of intelligence. If that's really true then clearly we have no idea what intelligence actually is. If you showed ChatGPT and DALL-E and the other software of today to people in 1975 they would say with near-unanimity that we invented intelligent machines.
However, just as the argument that AI is not useful because it is not true intelligence may not be valid, then so must be that many of the concerns of a true intelligence may also arise without ever achieving that goal.
Also, while almost anything can be viewed reductively, doing so is not necessarily a practical way to operate in / understand this world. Brains do a lot more than you mention.
It's easier to define what intelligence is not.
Statistically generated language that while grammatically correct still lacks comprehension or veracity is not it.
I agree; this is a pretty good baseline for written statements.
However, I was asking the parent-parent commenter (^^) who wrote:
> thfuran: A brain is just a sack of meat stochastically nudged towards only dying after more reproduction. Expecting real, original intelligence from something like that is the modern equivalent of alchemy.
Brains are quite good at satisfying jqpabc123's definition. I'm interested in thfuran's definition.
Without ever having seen it done?
One couldn't ask for a more clear expression of religious zeal.
But, in the future, more and more of the training data is going to be recycled bullshit output from other LLMs. Which is going to cause the model to get worse in accuracy and utility.
Run this forward a few iterations, and most of the Internet is going to be filled with thrice-recycled AI-generated garbage while there are some small pockets of accurate, truthful, curated information.
Reminds me of Herb Simon's "when information becomes abundant, attention becomes the scarce resource."
Having a good BS detector and ability to ask follow-up questions is going to be more and more important for humans over the next decade.
This implies we will need reliable methods for identifying the provenance of online content, which may spell the end for viable anonymity in many places, or at least accelerate us toward some kind of proof-of-personhood (a hard problem to solve in a privacy-preserving way, and indeed one that leads to the very same questions of what makes us human and gives us an identity in the first place). Short of such a reliable method, the best option for identifying human-written content will be filtering to content that was posted prior to ~2023. In a world of AI-generated content, your shitposts from 2015 may become a valuable commodity.
Also, I believe google willingly puts blog posts with ads on top, rather than the original documentation/maling list discussion because those usually have no ads.
That's something many humans already suffer from, due to social media platforms creating large echo chambers.
I think there's enough human-generated content that LLMs won't suffer from that. If they did, you could always manually filter what they train from... only data from websites with trustworthy timestamps pre-2020, for instance, or content after that which might be partially AI generated but still has strong human filtering, like wikipedia and arxiv and scientific papers in general.
Just no. It would have to take more than that to challenge Google's business model.
This current AI hype around LLMs and chatbots is the start of a hype cycle similar to what happened to Clubhouse.
Also, for now these services are spending a lot of money to capture the market, but sooner or later they'll monetize by biasing results towards their paid advertisers.
The end user isn't the customer, advertisers are.
Bullshits is good to me.