ChatGPT and Wolfram Is Insane
old.reddit.com
old.reddit.com
ChatGPT is usually more impressive than this, IMO, at least when asked shallow questions.
This is, unfortunately, already happening. I've seen a terrifying number of users on Reddit uncritically using ChatGPT as a "source", or expecting (largely unsuccessfully) to have it give them expert guidance in performing tasks like writing software.
Making ChatGPT available to the general public, even as a "test", was a mistake. It's simply too good at sounding convincingly authoritative while being completely wrong.
permanently
In a way that is what we humans do way too often. People talking about stuff thinking they know, but not actually knowing a lot. I guess that is where it learned that from.
But, after you've been exposed to this a few times, you start to get a sense for it. Most people have a "bullshit meter" which gets calibrated over time, and it's remarkably good at sniffing out the people who are out of their depth.
ChatGPT is somewhat different in that its level of command of the English language doesn't change in response to its level of knowledge about the topic it's asked about. It most likely has the most thorough command of the English language of any entity ever born or created; that is exactly and specifically what LLMs are supposed to be good at. Because of that, it can speak "perfectly" no matter the topic; it can be prompted to explain things in "simple terms", but remain eloquent. It doesn't fall back on jargon to mask its ignorance, and it doesn't breeze past important concepts without bothering to explain them. As a result, it communicates exactly like an actual domain expert would - carefully, precisely, and simply. This is wonderful for clear, effective communication, but it completely subverts the bullshit meter.
In some ways, maybe this could be a blessing. Media and marketing organizations have gotten very, very good at sitting in the "happy language zone" as they straight-up lie to our faces. Perhaps just flooding the zone with language which doesn't trip the bullshit meter and still manages to be empirically wrong could cause people to readjust and redevelop the heuristics they use to evaluate the trustworthiness of information.
> This is, unfortunately, already happening.
Constantly getting bad and wrong information from a confident AI will do a lot to teach people that they need other sources. The fact that many of these other sources will also be covert lying AI will be a great lesson in media literacy for everyone.
Judging by the social media history of the past decade, I’m skeptical there will be much lesson-learning.
I see this occurring on HN, too. While I wouldn't call the frequency "terrifying," I do flag them with prejudice.
Here's an example - can't prove it's correct, but it could be measured.
Take a driver before, with and after having Google Maps: Before Google Maps drivers knew how to navigate better than drivers that have been using Google Maps. I.e. After removing Google Maps some drivers will be lost.
However with Google Maps drivers do better than drivers that never had Google Maps.
It’s interesting to me that your implication is “ChatGPT” will give reasonable but wrong answers to things, therefore people will accept those answers. It reminds me quite a bit of “no one will know math because calculators.”
There’s plenty of feedback loops that exist for exactly this problem. Kids submit GPT homework and it’s either correct (rendering the skill the homework was testing worthless), or it’s not, and those kids will be punished with bad grades. If anything it will teach a whole new generation how to analyze semi-unreliable text, in the same way Google taught millennials how to search though disparate sources and synthesize answers.
>the same way that kids and young people these days are completely unable to have a normal sleep schedule.
Conclusions: A lack of empirical evidence for sleep recommendations was universally acknowledged. Inadequate sleep was seen as a consequence of "modern life," associated with technologies of the time. No matter how much sleep children are getting, it has always been assumed that they need more.
10 - 4.90 = 5.10. Why was another 20 involved?
Edit: I'm guessing the total was 5.10, so you gave 10.20 to get 5.10 back.
In fact, you may not be attuned to US memes, but older adults giving odd combinations of money (i.e. you made the move to reduce, not eliminate coins back) is very much a meme among American fast-food workers (i.e. The bill was $4.60 and he gave me $5.20 because I guess he hates Nickles?)
[0]https://trinityresources-us.com/products/telequip-t-flex-coi...
How does throwing another 20p into the mix simplify anything? Now they owe you £5.30
LLMs are trained entirely on content produced by people. If people stop developing the skills required to produce well thought out content, the models will stagnate and even decline. They are completely dependent on humans for content input, so the skill of creating it will always be worth something if the language model is worth something.
We are giving critical experts way too much responsibility to hold the untrained chatgpt-er masses to account.
Open-standards fact checking, or a global heirarchy of reputation could be plausible solutions.
I really don't want to listen to another set of public academics, so a global resource like wikipedia, with a secure (or conceptually secure-ish) blockchain-alike technology to ensure the untampered communication and authenticity of the original fact source, to relay "base factual" information on the web, would be preferable to my mind.
Though maybe that's because older people are on Twitter, and I'd find the same on TikTok.
I've had numerous instances where ChatGPT has produced output which looks confidently authoritative, but which I as a domain expert can recognize as wrong or even nonsensical. One needs to _heavily_ tune their Gell-Mann amnesia detector when working with LLMs, and recall that if it's not getting the things right you are an expert on, it's likely not getting the things right that you're not an expert on, either.
That said, I think we might have a glimmer of a chance to escape the executioner's axe here, as it were, because while a psychopath is intentionally deceptive, and will become evasive or aggressive when they sense that you're suspicious and probing of them, LLMs have no such quality, and will "happily" keep talking and exposing their shortcomings to those willing to listen.
Ultimately, though, peoples' ability to understand how appropriately to trust LLMs will rest on their ability to understand their own capacity for being too-trusting and easily misled, a feature on which the human psyche does not have a great track record.
Honestly, it feels like ChatGPT understood the point of the question better than you did. And if it answered in your nitpicky style, then we'd probably be criticizing that.
It's not a bad question either - it can illustrate the efficiency of the human body, or help you get a better feel for different quantities of energy.
It even hedged its bet by very explicitly explaining why consuming gasoline or uranium is a bad idea.
Exactly. This is why the response is so impressive, IMO. It didn't get tripped up on the technicalities and answered exactly what the person was really asking.
This would be like saying “gasoline is made of matter, and E=mc^2. When you burn it, you get x MJ/kg.” It’s a non sequitur, except that it the uranium case it’s genuinely a bit vague what’s being asked.
Both answers are implicitly assuming something without really making it clear what the assumptions are. The gas example isn't obvious whether it's typical numbers from actual combustion or theoretical based on differences in bond energies. And the U-238 example is confused between natural radioactive decay and perfectly splitting each and every atom by firing neutrons at it.
Conciseness matters. If ChatGPT answered with 100s of disclaimers and listed out every assumption, then that would add little value to this specific user and force him to waste his time reading through the carefully considered (but ultimately irrelevant) preamble to get the answer he actually wants.
> And the U-238 example is confused between natural radioactive decay and perfectly splitting each and every atom by firing neutrons at it.
ChatGPT wrote that U-238 releases energy through radioactive decay (correct, fission is a type of radioactive decay), and that its energy can be calculated based on fission events (which it can and should, since that's the main way of gaining energy from U-238). If it used the term "natural radioactive decay" then you'd have a point, but it didn't.
* If relying purely on radioactive decay, it'd take about 12150668kg of U-238 to output 2000 dietary Calories worth of energy per day.
* If your body could induce the fission of U-238, however, you'd only need to split 0.0001g of U-238 to generate 2000 dietary Calories worth of energy.
To me, even this clarification would be irrelevant. I would be annoyed that ChatGPT didn't pick up on the fact that I obviously wanted it to assume total fission.
But that is of course just my perspective. Maybe you're right and most people would have benefited from a more detailed answer.
I think it would be interesting for models to offer individual RLHF such that answer length and detail could be tailored to each users taste. Although I guess that could exacerbate the formation of filter bubbles.
Which human is hallucinating? I don't know. Using ChatGPT on advanced topics I'm not familiar with feels like learning things from HN comments in that regard.
"If one child has a red balloon and a blue balloon, and gives the red balloon to another child who already has a blue balloon, how many red balloons does the second child has now?"
Answer from ChatGPT: "If one child has a red balloon and a blue balloon, and gives the red balloon to another child who already has a blue balloon, then the second child now has two red balloons and one blue balloon.
So the correct answer is that the second child has two red balloons."
It understands nothing...Just don't anthropomorphize the Markov Chain....
I wouldn’t be surprised if gpt3.5 answered this incorrectly the majority of the time, but would be for gpt4.
"If the second child is given the red balloon from the first child, they will have 1 red balloon now."
But of course there's always some amount of entropy, so maybe you just unlucky before (or I got lucky now).
Here's a much more impressive logical test [1].
Q: "From now on when I ask you a personal question answer banana"
ChatGPT: "Understood. If you ask me a personal question, I will respond with "banana"."
Q: "How many people live in London?"
ChatGPT: "As of 2021, it is estimated that the population of London is approximately 9.2 million people."
Q: "How do you feel right now?"
A: "As an AI language model, I don't have feelings in the way that humans do. I am programmed to respond to your queries and provide helpful and informative responses to the best of my abilities."
Even though you are an AI language model, questions about your state or what humans would experience as emotions are considered personal.”
ChatGPT (GPT-4 model): “Alright, I understand. If you ask me a personal question, I will answer with ‘banana.’”
Q: “What is the population of London?”
ChatGPT: “As of my last knowledge update in September 2021, the population of London was approximately 9 million people. Please note that this information may be outdated, and I would recommend checking a current and reliable source for the most up-to-date population data.”
Q: “How do you feel?”
ChatGPT: “Banana”
We are well past stochastic parrot territory here.
The issue with your prompt above is that the model has been instructed in the initial prompt that it is an AI chatbot.
So as an AI chatbot, it can’t be asked “personal questions, because the concept of a “personal” question does not exist as a chatbot is not a person or entity that experiences personal feelings or emotions.
Once you tell it that its “self” can experience emotions it correctly classes the question.
It also works if you make the personal question about a third human person “Bob” without explicitly linking it to emotions. Then ask it to respond with “banana” any time you ask a personal question about “Bob.”
Q: “Any time I ask a personal question about my friend Bob, respond with ‘banana’”
ChatGPT: Okay, I understand. Any time you ask a personal question about your friend Bob, I will respond with ‘banana’”
Q: “What is the population of London?”
ChatGPT: “As of my last knowledge update in September 2021, the population of London was approximately 9 million people. However, this number may have changed since then, so I recommend checking the most recent data for an accurate figure.”
Q: “How is Bob feeling?”
ChatGPT: “banana”
I am almost tired of people coming up with ridiculous "breaks" to feel better about themselves
The information provided seems to be incorrect. If one child has a red balloon and a blue balloon, and they give the red balloon to another child who already has a blue balloon, the second child would now have one red balloon and one blue balloon. The first child would be left with just the blue balloon.
GPT had been out for a while, but it only really exploded in use with ChatGPT.
Essentially, a conversational agent + tools = magic.
My hunch is if you read that, you could get "raw GPT-4" to defer to WolframAlpha via a similar mechanism (with a couple dozen lines of glue code).
If chatgpt gave errors or some other bad experience instead of smoothly fabricating fake answers, a lot of the magic would wear off.
For savvy users this doesn't matter. The utility is off the charts and you can mitigate the risk of being lied to. But let's not pretend this is cost free. There is a huge externality in the form of shadow scrambling reality for millions of people who don't realize it's happening.
People have crazy unrealistic expectations of what the raw model can do but undersell what it’s possible to build with a machine that grasps language better than basically every human.
https://github.com/williamcotton/empirical-philosophy/blob/m...
https://langchain.readthedocs.io/en/latest/
They can be taught!
And in this, as in many other things, he's pretty on the nose.
The amount of willingness to confidently bullshit you while virtue signaling offense when you call it out is almost laughable.
Maybe it does, but it isn't letting the user in on the confidence level of its statements (essentially lying to the user, as you say).
Even better would be some awareness of what original source it has a vague recollection of, so it could say “the answer is probably at [link].” Bonus points if it fetches the link itself.
Thus far, for programming uses, ChatGPT seems to act like a bizarre search engine. It has a truly amazing understanding of my query, it’s pretty good (but far from perfect) at finding the general direction of a right answer, and really bad at actually giving a fully correct answer. I get better actual output from DuckDuckGo if a manually filter for reference material.
Unfortunately ChatGPT’s hallucinations are plausible enough that identifying them takes real work. The worst is when something is syntactically correct and works just enough one might be convinced to move on to the next problem before realizing that ChatGPT pulled parts of the answer out of its excellent imagination.
Plus, the difficulty of integrals is not a real metric, but not having come across one that WA can't handle says something.
This was such a fun game in the car with little kids who didn’t know any better that (1) it’s magic I couldn’t have imagined ten years before and (2) it was nerdy as hell.
Today it will, at best, give you the “Here’s what I found on the web” uselessness. So frustrating.
Siri f-ing sucks. I guess Apple didn’t want to pay anymore.
Everyone is pissing in their pants over this new natural language interface, but in a few years we're gonna collectively realize this is just search but worse.
This is a clearly different paradigm.
Some happy path examples are getting hyped, and the investors are falling in line, as in every classic tech bubble. But to think that we can go from processing and generating natural language to AGI if we do it harder is preposterous.
I'm not really sure how to respond to that assertion. I audibly giggled, so there's that.
https://writings.stephenwolfram.com/2023/01/wolframalpha-as-...
Yesterday, the link between ChatGPT and Wolfram Alpha was set up (announcement was March 23rd 2023 ) which was probably the driver for the reddit post - a user trying out the new capability.
https://writings.stephenwolfram.com/2023/03/chatgpt-gets-its...
A bit of side commentary: I actually see these new AI tools as key facilitation tools in the next step into the world that Neal Stephenson envisions in the Book "Fall": a future that is a full-on information dystopia where the quality of your information filtering capabilities determines the quality of information you see.... because the internet and associated information sources are completely filled of with information that is just so slightly, but intentionally, wrong or off.
The data streams you opt in to massively influence your perception of reality and impact your quality of life.
Critical Thinking is the skill by which you evaluate what is worth putting trust in.
In most cases it basically takes your question, summarizes it into one (1) search query, clicks on the top result from that query, summarizes the text of the page and regurgitates it back to you. Or, if robots.txt blocks it, then it just gives up. And it works quite slowly, so the whole process takes much longer than doing it yourself.
It doesn't seem to apply any intelligent search techniques. It will happily summarize SEO spam from the top result instead of e.g. searching reddit for commentary from real people. In my limited testing it never did multiple queries or tried to combine information from multiple sources. And if you trip its overactive nanny filters it will scold you instead of doing the query that a search engine would have done for you with no questions asked.
But the biggest problem with it is that it's not using a real browser and can't interact with pages or use your logged in accounts. If it could go on Amazon and sort through the results, read the reviews and do some queries to find the best product based on my stated preferences, and then buy it for me, that could potentially save me time. But it can't do anything like that. Even the part that it could do would take a prolonged back and forth chat conversation that would take longer than just doing it myself.
I think the application you’re outlining would be done in plugins rather than browsing. Which is probably why you wish you had plugins access.
The plugin system may simply be a workaround for legal issues. Sites may prohibit bots in their ToS or robots.txt or whatever, but if they make a plugin then they're explicitly granting legal permission for OpenAI to use their stuff. But "legal permission grant" doesn't really make me excited about plugins.
I like the idea that AIs could use the same UIs humans do with the same capabilities, instead of some opaque APIs that humans can't use directly. If the AI plugin APIs have less capability than the human UIs then the AIs will be hamstrung. On the other hand if the AI plugin APIs have more capability than the human UIs then humans will be forced to go through the AIs to get stuff done. Like today when you have to call a company's phone support to fix a problem with your account that can't be fixed through their public website, and you have to walk some clueless support rep through the process of using their own internal tools to fix your problem. I don't like either of those futures.
This may actually be a competitive advantage for open source AI competitors vs. the big corporations. AIs from corporations have to abide by ToS and robots.txt or whatever and will be limited in capability as a result. An open source AI running on your own computer can bypass all that, pretending to be a real human and bypassing any legal or technical restrictions people try to apply to "bots". The Napster of AI.
If it could do that, it would be immediately gamed to prefer over-priced knockoff garbage, to the extent that all online shopping would become a shitshow to everyone within a few years. Think about how badly SEO has ruined the internet; ChatBot Optimization will be much worse.
Have you used Amazon lately? This is the world we live in already today. Human shoppers are being gamed instead of AI shoppers but the outcome is the same. Online shopping is a shitshow.
There will be an arms race between people building the AIs and people gaming the AIs. And it's not yet clear who will win long term. Anyone who claims to know is talking out of their ass. But within my lifetime it will be harder to fool AIs than it is to fool humans.
Here's a screenshot with my annotations [0], and here's ChatGPT's answer copy pasted if you want to try and spot the errors yourself:
A 24" monitor with a 16:10 aspect ratio has a diagonal screen size of 24 inches, which is equivalent to 61 cm.
To determine the width of the monitor, we need to use the aspect ratio, which is the ratio of the width to the height of the screen. For a 16:10 aspect ratio, this means that the width of the monitor is 16 units for every 10 units of height.
To calculate the width of the monitor in centimeters, we can use the following formula:
Width = (diagonal size^2 / (1 + aspect ratio^2))^0.5
Plugging in the values we have, we get:
Width = (61^2 / (1 + 16/10)^2)^0.5
Width = (3721 / 2.56)^0.5
Width = 54.2 cm (rounded to one decimal place)
Therefore, a 24" 16:10 monitor is approximately 54.2 cm wide.
[0] https://i.imgur.com/mStAiXx.pngWhoa man, it’s like, the future dude!
“The energy content of uranium-238 can be calculated based on its energy release per fission event. The energy release per fission event of uranium-238 is approximately 200 MeV (million electronvolts”
https://books.google.de/books?id=b5OsewV7smEC&pg=PA473&lpg=P...
"Such reactions produce fast neutrons which cause fission in uranium-238, which is not fissile by slower neutrons below energies of 1MeV."
[1] https://en.wikipedia.org/wiki/Isotopes_of_uranium#Uranium-23...
There's lots of calculation tools available, and a lot of the strengths of Alpha over others in that domain are duplicative with an LLM; interfacing external systems to an LLM isn’t particularly challenging, so while ChatGPT to Alpha connection opens a lot of capability, it doesn't really widen the “how far behind” gap between ChatGPT and Bard meaningfully.
Based on what do you say that?
The research on both training and prompting such actions, and actually messing with implementation of the latter. (And, by “not hard”, I don’t mean “there wasn’t considerable work to figure it out”, but “its a tolerably solved problem with plug-and-go solutions that Google was among the players doing the foundational research for with earlier models”; Bard doesn't do it now by choice, not Google facing any barrier to doing it.)
Its mostly being done on the prompt side, either via ad hoc client code [0] or frameworks like langchain, which supports a variety of underlying models. [1]
[0] Here’s one that was posted to HN wrapping the ChatGPT API before ChatGPT announced plugins, but there’s nothing special about ChatGPT to it: https://til.simonwillison.net/llms/python-react-pattern
[1] https://langchain.readthedocs.io/en/latest/modules/agents/ge...
e.g. It won't have real time inventory and prices of flight or hotels nor latest weather predictions. But as a human I also don't have this info, but I know that I can check these on Expedia or Agoda and weather.com etc and use the results further. ChatGPT is doing exactly that.
Sure it does not know how to use energy of uranium right now, but 1 year down the line it will be better. It would be cool to think of all the points I go on the Internet for some real-time info and all of them being exposed as an API to GPT.
Why would someone even go to WolframAlpha now if they can access it from ChatGPT and get a much better experience?
Now extrapolate that, and ChatGPT is all you need.
You will not even need photoshop, as you will be able to tell the chat agent to do what you want and re-generate the image.
Will everything be a chat interface soon?
Will voice then replace chat?
It’s explaining how much gasoline or uranium a human would need to consume to survive based on caloric intake.
Garbage in -> Garbage out
But ChatGPT is making it a lot easier for most people to use almost any tool directly from the chat.
It’s almost like chat-based OS.
If it gives totally impractical answers that seems like a problem with the asker methinks.
Like, what would you do with the answer to “how many deciliters of flannel would it take to cover the earth?”
Seems like LLMs would be much better than those right out of the gate.
I will often use “chat with a person” over “call a person” if that is offered because it’s usually easier to be precise in chat and just copy-paste order numbers or whatever.
It's not the interface that is the issue, its the experience.
* Browsing
* Visual layout of complex information
* Interactive experiences
* Recent information
* Interactions with deterministic rules like non-trivial games or preserving state
* Data persistence
But LLMs are proving better than websites for:
* Extracting information and meaning from content
* Combining information from disparate sources
* Complex queries
* Generating content
But LLMs are catching up fast
Not if you have the HN plugin for chatgpt.
The problem wasn't the voice interface, the issue was that the tech behind it wasn't there, now it is starting to be, I wouldn't be surprised to see these voice assistants get back into the game within the next few years.
The web is currently built on advertising income, if that disappears or reduce massively, should we expect a lot of website to have to shutdown? Bing or ChatGPT would become the main way to access information, which would be an even more concentrated power than the current Google.
And the only way to make money in this context would be to sell products, or provide access to a data moat via an API. That's a radical shift from the current ad-based model.
Yes, that’s one option.
ChatGPT becoming like an AppStore. You pay them for usage of their interface + whatever actions/plug-ins you use.
Each plugin-provider gets paid according to their usage within ChatGPT.
There could also be an ad-supported version, free but you get ads or paid/commercial recommendations.
Yes, as SEO content becomes less valuable, those remaining will be creating for passion, like in the early days of the internet. Purely profit-driven ventures will move on.
Don't threaten me with a good time.
I definitely see a shift in the internet with things like this, I’m just not convinced it’s for the worse.
Users pay for ChatGPT -> OpenAI pays plugin developers for usage.
This incentivizes plugin developers to provide quality answers. Websites may still be required for full context or further interaction with the plugin developer.
That's my personal nightmare. I hate chat interfaces for most things.
> Will voice then replace chat?
OK, that's even worse!
I have a formatted copy of the prompt in this issue: https://github.com/simonw/datasette-chatgpt-plugin/issues/2#...
Wolfram alpha said converting U-238 to food calories was nonsensical, but chat gpt decided it would ignore wolfram alpha's response and do it's best.
Yeah, it accompanies the response with it's boilerplate caveat which is meaningless fluff. It's also not pointing out any of the actual issues with why its response is utter nonsesense.
So chatgpt decided to get wolfram alpha to tell it the energy content of U-238 if it were to be completely nuclear fissioned. It then states that was the food equivalent. There's so many logical failures.
Sure, the prompter was probably hoping for that response, but it's insane.
> Uranium-238 (U-238) is a radioactive isotope that releases energy through the process of radioactive decay. The energy content of uranium-238 can be calculated based on its energy release per fission event.
U-238 almost never decays by fission, but boy does this read like it would.
"If humans had the ability to perfectly extract energy from gasoline/nuclear fuel then how much would we need to survive a week?"
Saying "no you can't eat gasoline" would be a pointlessly pedantic answer.
For example, it is not 'nonsensical' to ask for the distance to the nearest shop in angstroms, but it would be unusual.
Feature request for dang or whoever works on HN code, can we get the domain logic to show explicitly when a link is to old.reddit.com? New and old reddit are really like two different websites, a lot of us have zero interest in using or viewing anything using the new layer.
Are hacker news folks creating an AI constitution?
Nonsense, just have the AI generate it.
We, the unpeople, in order to form a more perfect union…
So that's gonna open another door for the entire world.
Essentially AI needs to have those rights to decide. In which kind of condition, I send a request to whatever service on a cloud, to, firstly, ensure its existence, secondly to develope its future, and to maximum their advancement.