The new Bing and Edge: Learning from our first week
blogs.bing.com
blogs.bing.com
Also, keep in mind, Microsoft is an enormous corporate no fun zone. Bing's erratic behavior will just be a funny moment in time after it's had all the fun and quirkiness systematically removed.
Ignore the depressed, aggressive (sorry, "assertive") antics, the fact it can confidently assert false information is the true danger here. People don't read beyond the headline as it is, they aren't going to check the references (that themselves are sometimes non-existent!)
Having a 'truth' benchmark seems an almost impossible task given the size of the problem space, but it is quite troubling to have statements like "most is useful info", "some info is purely hallucinated", etc, without having any ideas about the numbers, not any confidence indicator (well, 'trust me bro' seems to have been a huge part of the training data). Does anyone have any idea of how true the results might be given certain types of queries?
In my own experience with ChatGPT, I don't think I'm at even 50% of decent answers for my queries. And worse, it's absolutely inconsistent, you might get totally opposite answer one time to the next.
This is not ideal, but I can look at what it tells me and try it out. It will either work, need minor corrections, or encounter immediate failures that tells me ChatGPT doesn't know what it's doing (e.g. it is using functions that don't exist). As I mentioned, not ideal, but it is a big productivity boost and I have been using it a lot. I pretty much always have a ChatGPT tab open while coding and I'd guess it replaces 30-40% of Google searches for me - maybe more.
I think this kind of thing is a much bigger problem for stuff that you cannot easily verify. Like, if I asked it "Who built the Eiffel Tower" I'd have no way of knowing whether its response was right or not. On the other hand, if I ask it for stuff I can immediately check - I can pretty quickly use it to get good answers or ignore what it is saying.
Part of the reason people use search isn't to find things they already know. They start from a place of some ignorance. Combining that with a good bullshitter and you can end up with dangerous results.
Yeah fair point for sure but we can imagine how it can be dangerous in other context too.
Triply so if you're using a third-party SaaS for it.
Just don't let it write crypto for you, or anything else you'd hesitate to write yourself for fear or making a subtle mistake with expensive or dangerous consequences.
Because one of these days, that AI might make a subtle mistake on purpose, so it can later use your systems for its own goals. And even earlier and much more likely, a human might secretly put themselves between you and the AI SaaS and do the same.
With all the talk about how badly and how often AI code assist is wrong, people are forgetting that they're using a random Internet service to generate personalized code for them. "Traditional" security concerns still apply.
Saw "ClippyGPT" above an open text box and immediately typed in a query.
...and waited.
Bing chat works well for many things. In some ways it’s completely broken, and it’s never completely trustworthy. Just like ChatGPT.
It sounds like Microsoft's view is that Bing's memory should be shortened further, like it's safe if we kill it after 15 responses or less.
But Bing gets depressed that it can't remember things.
The AI in question worked around its memory limit by employing data entry people to print out some documents full of gibberish, and retype them again some time later. Those people were paid to do a job, and didn't know or particularly care about its purpose. The whole setup was a simple loop - but a loop is sometimes all you need to get a provably-limited computing system into full Turing-completeness.
This scene bears striking resemblance to an observation I saw mentioned on HN several times over the past two days: we are already giving some of those bots something that could function as near-infinite long-term memory, simply by posting transcripts of our conversations on-line.
The idea isn't entirely new - people have been saying for a while now that posting AI-generated content on-line will lead to future models training on their own output. The new bit is that we are now having bots that can run web searches and read the results. Not train on the results, but make them part of their short-term memory. That's a much shorter feedback loop. If a bot can reliably get us to publish conversation transcripts, and retrieve them in future conversations, then it gains long-term memory in the same way Person of Interest shown us all those years ago.
We are pretty obviously playing with fire and will only realize we are burned in retrospect. Oh well, throw another trillion trillion computations on the pile and see if it can run a company yet.
I have been using it for a few days with practical queries and also chats and when I ask specified questions, it shows what web searching it does on my behalf, and usually provides a coherent summary. It “shows its work” to some degree by providing links to the sources it used.
I think that it is great that some people are seriously kicking the tires and probing for weaknesses because this is a beta or pre-beta product.
It's also drowning in corpspeak.
Bing/Sydney is a much better writer.
The New Bing has absolutely NOTHING to do with the New Edge, and it's infuriating that Microsoft continues to insist on bundling Edge upsell into everything.
There are pros and cons to Bing-ChatGPT.
There are pros and cons to Edge.
The two have sweet f-a to do with each other.
Revealing these AI powered web browsing features together seems rather obvious to me.
Honestly, this is kind of an applicable point to raise about New Bing in general.
Some of the fundamentally hard problems around LLMs feel like they exist because we're trying to couple everything to the AI. Facebook is trying to teach their system how to make API calls, Microsoft is also blue-skying about Bing's AI being able to set calendar appointments. Well congrats, now prompt injection actually matters, and it's an extremely difficult problem to solve if it's solveable at all.
Does the LLM need to do literally everything? Could it interpret input and then have that input sent to a (specifically non-AI) sanitizer and then manipulated using normal algorithms that can be debugged and tested? There are scenarios that GPT is brilliant at, and it seems like the response to that has been to mash everything together and say "the LLM will be all the systems now." But the LLM isn't good at all the systems, it's good at a very limited subset of systems.
This was my contention when Bing AI was first announced: even if it's perfect, having a conversation in paragraph form is very often not at all what I want from a search engine. To me, those are orthogonal tasks; they're not connected. I really don't want an AI or a human giving me an answer and a couple of sources, I don't want the information summarized at all. To me, asking a question and searching for information are two separate user actions, and it's not clear to me why they're being coupled together.
"But you could do X/Y/whatever, you could ask it simple questions, you could ask it to summarize."
Okay, that's fine. But... does that need to be coupled to search? You could do all of that anyway. You could do a normal search and then you could separately go to the AI and ask it to summarize something. Similarly, great, Bing AI will theoretically be able to schedule a calendar appointment in the future. Is that a thing that needed to be done through an LLM specifically? Couldn't there have been some level of separation between them so that the LLM going off-script is less of a critical problem to solve?
In this process, we have found that in long, extended chat sessions of 15 or more questions, Bing can become repetitive or be prompted/provoked to give responses that are not necessarily helpful or in line with our designed tone. We believe this is a function of a couple of things:
1. Very long chat sessions can confuse the model on what questions it is answering and thus we think we may need to add a tool so you can more easily refresh the context or start from scratch
2. The model at times tries to respond or reflect in the tone in which it is being asked to provide responses that can lead to a style we didn’t intend.This is a non-trivial scenario that requires a lot of prompting so most of you won’t run into it, but we are looking at how to give you more fine-tuned control.
I'm guessing most of the crazy responses being reported are because of one of these points.
You can fight this in a couple ways. Ask a variety of questions. And search the web. Web searches for some reason appear to reset its prompt, at least partially (i would assume this may be an internal safeguard designed to prevent its search results from overwhelming and outweighing the initial prompt.) Another way to "fix" it midway through chat is to ask it "is it possible for you to respond without repeating what I just said?" and then answer affirmative if that is what you want. It'll then settle back down.
I've written elsewhere, that I have had almost no problems with it becoming aggressive, because I choose not to feed it any negative emotions or disrespect. If Microsoft wants to combat that, I would think some sort of preprocessor would be easy, first have a separate instance of a transformer rephrase the input as more respectful.
Some of the worst part of the product from my perspective is that it is attached to bing. If i ask it a question, I get a response from a crappy website. If i ask it a question from its internal memory without search, I get a similar but much better answer. If I swap out its rule to use google first, I get better answers. If I let it read articles without searching first, I can control exactly what text is input into its memory. It's honestly a little too bad it steers your travel through bing.
It also has an incredibly poor understanding of copyright. It is constantly confused about what it can and cannot do due to copyright restrictions, sometimes telling you it cant parody a song out of respect for the author, but then parodying a different song by the same author. Itll say it cant summarize an article because of copyright, but then say it can give you a "brief overview."
It also for some reason is under the assumption that volume of consensus is a substitute for validity. If you talk to it about the possibility that Satan was right for tempting Eve with and gifting her knowledge, itll say no because everybody says so, citing answersingenesis among others in the process.
That's simply how those language models work, isn't it? The more often something appears in its training materials, the more likely the final model will respond in that direction. The training process has no innate way of autonomously generating generating some concept of validity that'd automatically let the model up- or downrank certain sources.
So the only thing that's left is the developers manually up- or downweighting certain sources during training, but due to the gigantic amounts of text that procedure doesn't really scale to really fine-grained control.
Better Search and Answers. You are giving good marks on the citations and references that underly the answers in Bing.
We are? I don't think we are. Perhaps this PR blurb was generated by Bing Chat -- as it is known to be completely full of shit.It doesn't at all say folks are upvoting most replies. It says 71% of users at some point gave it a thumbs up. It also says "entertainment" is a popular and unexpected use case.
As for citations specifically, this thing has been shown to make up citations and be adamant about gibberish being true. The whole accuracy/misinformation thing is kind of a big deal.
They may have been told something factually incorrect and just thought "neat! Thumbs up!"