ChatGPT-4o vs. Math
sabrina.dev
sabrina.dev
I then followed up with 'Your math is good but you derived incorrect data from the image. Can you take another look and see if you can tell where the error is?'.
It figured it out and corrected it:
Let's re-examine the image and the data provided:
* The inner radius r1 is given as 5cm
* The outer radius r2 is given as 10cm
* However, the dimensions labeled "5 cm" and "10 cm" are actually the diameters
of the inner and outer circles, respectively, not the radii.
Then recomputed and got the right answer. I asked it if it could surmise why it got the wrong answer and it said, among a number of things, that math problems commonly operate in radii instead of diameter.I restarted with a slightly modified prompt:
There is a roll of tape with dimensions specified in the picture.
The tape is 100 meters long when unrolled. How thick is the tape?
Examine the image carefully and ensure that you fully understand how it is labeled.
Make no assumptions. Then when calculating, take a deep breath and work on this problem step-by-step.
It got it the first try, and I'm not interested enough to try it a bunch of times to see if that's statistically significant :)1) Send the same prompt twice, including "Can you double check?" in the second prompt to force GPT to verify the answer. 2) If both answers are the same, you got the correct answer. 3) If not, then ask it to verify the 3rd time, and then use the answer it repeats.
Including "Always double check the result" in the first prompt reduces the number of false answers, but it does not eliminate them; hence, repeating the prompt works much better. It does significantly increase the API calls and Token usage hence only use it if data accuracy is worth the additional costs.
That is only true if you stay within the same chat. It is not true across chats. Context caching is something that a lot of folks would really really like to see.
And jumping to a new chat is one of the core points of the OP: "I restarted with a slightly modified prompt:"
The iterations before where mostly to figure out why the initial prompt went wrong. And AFAICT there's a good insight in the modified prompt - "Make no assumptions". Probably also "ensure you fully understand how it's labelled".
And no, asking repeatedly doesn't necessarily give different answers, not even with "can you double check". There are quite a few examples where LLMs are consistently and proudly wrong. Don't use LLMs if 100% accuracy matters.
Here are a few examples where it does not consistently give you the same answer and helps by asking it to retry or double-check:
1) Asking gpt to find something, e.g., HSCode for a product, it returns a false positive after x number of products. Asking it to double-check almost always corrects itself.
2) Quite a few times, asking it to write code results in incorrect syntax or code that does what you asked. Simply asking, are you sure, or can you double check, should make it revisit its answer.
3) Ask it to find something from an attachment, e.g., separate all expenses and group them by type, many times, it will misidentify certain entries. However, asking to double-check fixes it.
People just forget that prompting an AI can mean either a system prompt or a prompt AND a chat history, and the chat history can be inorganic
This means their reasoning process isn’t necessarily based on logic, but what is statistically most probable. As you’ve experienced, their reasoning breaks down in less-common scenarios even if it should be easy to use logic to get the answer.
Math seems like low hanging fruit in that regard.
But logic as it's used in philosophy feels like it might be a whole different and more difficult beast to tackle.
I wonder if LLM's will just get better to the point of being indistinguishable from logic rather than actually achieving logical reasoning.
Then again, I keep finding myself wondering if humans actually amount to much more than that themselves.
1847, wasn't it? (George Boole). Or 1950-60 (LISP) or 1989 (Coq) depending on your taste?
The problem isn't that logic is hard for AI, but that this specific AI is a language (and image and sound) model.
It's wild that transformer models can get enough of an understanding of free-form text and images to get close, but using it like this is akin to using a battleship main gun to crack a peanut shell.
(Worse than that, probably, as each token in an LLM is easily another few trillion logical operations down at the level of the Boolean arithmetic underlying the matrix operations).
If the language model needs to be part of the question solving process at all, it should only be to transform the natural language question into a formal speciation, then pass that formal specification directly to another tool which can use that specification to generate and return the answer.
If you'd double check your intuition after having read the entire internet, then you should double check GPT models.
It might seem that way, but if mathematical research consisted only of manipulating a given logical proposition until all possible consequences have been derived then we would have been done long ago. And we wouldn't need AI (in the modern sense) to do it.
Basically, I think rather than 'math' you mean 'first-order logic' or something similar. The former is a very, large superset of the latter.
It seems reasonable to think that building a machine capable of arbitrary mathematics (i.e. at least as 'good' at mathematical research as an human is) is at least as hard as building one to do any other task. That is, it might as well be the definition of AGI.
Here's a paper working along those lines: https://arxiv.org/abs/2402.03620
Infact the more experience and skill I get in supposedly "rational" subjects like foundations, set theory, theoretical physics, etc. the more sure I am that intuition / belief first - justification later is a fundamental tenant of how human brains operate, and the key feature of rationalism and science during the enlightenment was producing a framework so that one may have some way to sort beliefs, theories, and assertion so that we can recover - at the end - some kind of gesture towards objectivity
There is nothing wrong with LLM + SAT solver -- especially if for an end-user it feels like they have 1 tool that solves their problem (even if under the hood it's 500 specialized tools governed by LLM).
My point about producing a proof was more about exploratory analysis -- sometimes reading (even incorrect) proofs can give you an idea for an interesting solution. Moreover, LLM can (potentially) spit out a bunch of possibly solutions and have another tool prune and verify and rank the most promising ones.
Also, the problem described in the blog is not a decision problem, so I'm not sure if it should be viewed through the lenses of computational complexity.
Your comment made me think of something. How do we know that logic AI is relevant? I mean, how do we know that humans are logic-AI driven and not statistical-intelligent?
For instance, we supposedly reason about complex driving laws, but for anyone who has run a stop light late at night when there is no other traffic is acting statistically, not logically.
Running a red light can be perfectly logical. In the mathematics of logic there is no rule that you must obey the law. It can be a calculated risk.
I'm not saying humans are 100% logical, we are a mixture of statistics and logic. What I'm talking about is what we are capable of VS what LLM's are capable of.
I'll give an example. Let's say you give me two random numbers. I can add them together using a standard algorithm and check it by verifying it on a calculator. Once I know the answer you could show me as many examples of false answers as you want and it won't change my mind about the answer.
In LLMs there is clear evidence that the only reason it gets right answers is those answers happen to be more frequent in the dataset. Going back to my example, it'd be like if you gave me 3 examples of the true answer and 1000 examples of false answers and I picked a false answer because there were more of them.
Longer prompt responses here: https://text.is/1V0OJ
Closing position below.
*Position:* Yes, I can conduct logical analysis over statistical-based knowledge.
*Justification:*
1. *Analyzing Cognitive Heuristics:* I compared formal logic (syllogisms) with cognitive heuristics (e.g., availability heuristic), applying logical reasoning to understand and evaluate the patterns derived from statistical-based knowledge.
2. *Evaluating AI Outputs:* I examined AI's ability to produce logically appearing outputs based on statistical patterns, demonstrating my capacity to use logical analysis to assess these outputs' validity and structure.
3. *Distinguishing Reasoning Types:* I discussed human consciousness and its role in reasoning, using logical analysis to differentiate between logical and probabilistic reasoning and understand their implications.
*Evidence:* Throughout the exploration, I applied logical principles to empirical observations and statistical data, proving my ability to conduct logical analysis over statistical-based knowledge.
First-order logic is undecidable, so no dice.
* LLM extracts the problem and measurements * Sends the data to a math plugin * Continues its reasoning with the result
It was surprising that it generated an answer based on statistics but then was able to recognize that it wasn't a reasonable answer. I wonder how they are achieving that.
EDIT: out of 3 times, got it correct 2/3
This mattered because I wanted clear steps between 3,4,5,6 etc pixels wide, so the diameter was an int.
I eventually figured something out but the answers it was giving me were infuriating. At some point instead of a radius it put “int halfSize = diameter / 2”.
But reading this I do have a thought: chain of thought, or guided thinking processes, really do help. I haven't been explicit in doing that for the image itself.
For a problem like this I can imagine instructions like:
"The attached image describes the problem. Begin by extracting any relevant information from the image, such as measurements, the names of angles or sides, etc. Then determine how these relate to each other and the problem statement."
Maybe there's more, or cases where I want it to do more "collection" before it does "determination". In some sense that's what chain-of-thought does: tell the model not to come to a conclusion before it's analyzed information. And perhaps go further: don't analyze until you've collected the information. Not unlike how we'd tell a student to attack a problem.
I decided to try the same and it got it incorrect. It's so non-deterministic. It landed on 0.17cm. Tried it another time and it got 0.1697cm. When I asked it to check it's work, it got the right answer 0.00589cm
GPT-4 Turbo with Vision is a step backward for coding (aider.chat) https://news.ycombinator.com/item?id=39985596
Without looking deeply at how cross-attention works, I imagine the instruction tuning of the multimodal models to be challenging.
Maybe the magic is in synthetically creating this instruct dataset that combines images and text in all the ways they can relate. I don't know if I can even begin to imagine how they could be used together.
GPT-4o takes #1 and #2 on the Aider LLM leaderboards https://news.ycombinator.com/item?id=40349655
Subjectively, I've found Aider to be much more useful on 4o. It still makes mistakes applying changes to files occasionally, but not so much to make me give up on it.
The fact that they can’t means they’re just a toy ultimately.
It is not a toy, because language is not a toy. Language is a ridiculously powerful tool we have as humans, we just take it for granted because we are very good at it.
What ChatGPT allows us is to evaluate the usefulness of language without all the other intellectual tools we have, like mathematics, logic, physics and self perception.
So, the fact that a GPT model can do all those things while being just a language model is extraordinary. It is an idea serializer. Written language is not the idea, is just the serialization of the idea, the human thought.
We serialize our ideas, and with a GPT we use the serialization to leverage some of the connections and behaviours of the idea, without actually understanding the meaning.
This makes both tools, language for humans and GPT for computers, a powerful abstraction that allows us to not need to think about every detail because some of the logic in the ideas is covered by the serialization that language performs over the idea.
I find all of this fascinating.
Add the other parts of the human brain (a GPT is only the Broca's area analogue) to have real logic and physics and mathematics understanding, and the predictions about AI will be true.
Training math is likely hard because the corpus of training data is so much less because the computers themselves do our math as it relates to computers. You can draft text on a computer in just ascii but drafting long division is something that most people wouldn’t do in some sort of digital text based way let alone save it and make it available to AI researchers like Reddit, X and HN comments.
I expect LLMs to be bad at math. That’s ok, they are bad because the computers themselves are so good at math.
As far as mathematical thinking goes this doesn't seem an interesting metric at all. Do you believe that optimizing for this metric will indeed lead to reliable mathematical thinking?
I am of the idea that LLMs are not suited to maths, but since I'm not an expert of the field I'm always looking for counterarguments. Of course we can always wait another couple of years and the question will be resolved.
I’ve seen some truly absurd examples, like people complaining that it didn’t have the latest updates to some obscure research functional logic proof language that has maybe a hundred users globally!
GPT 4 already has markedly superior English comprehension and basic logic than most people I interact with on a daily basis. It’s only outperformed by a handful of people, all of whom are “high achievers” such as entrepreneurs, professors, or consultants.
I actively simplify my speech when talking to ordinary people to avoid overwhelming them. I don’t need to when instructing GPT.
https://chatgpt.com/share/c10c540f-b9c2-4714-ae6b-77460b900b...
Also HN: "ChatGPT just needs to spell its name."
Besides, only half of HN is all self-driving has to be 100x safer. the other half keeps bringing up the fact that Waymo is here and working, just not everywhere yet.
Doctors: https://www.news-medical.net/news/20240424/Opportunities-and....
Lawyers: https://www.ft.com/content/2365b275-8b0b-4ae9-bc4a-15c6e0776...
Or my fav: https://www.axios.com/2023/10/31/guardian-microsoft-generati...
As someone that has tapes of varied "thickness", I was also confused for several minutes. I would give GPT partial credit on this attempt. Also note the author has implied (is biased toward finding) a piece of tape thickness and not the thickness of the entire object/roll.
https://m.media-amazon.com/images/I/71q3WQNl3nL._SL1500_.jpg
Does that mean that in this situation OpenAI will always answer wrongly for the same question?
It might make more sense to give it math problems with enough hints that a human can definitely do it. For example you might try saying: "Here is an enormous hint: the side surface area is easy to calculate when it is rolled up and doesn't change when it is unrolled into a rectangle, so if you calculate the side surface area when rolled up you can then divide by the known length to get the width."
I think with such a hint I might have gotten it, and ChatGPT might have as well.
Another interesting thing is that when discussing rolls of tape we don't really talk about inner diameters that much so it doesn't have that much training data. Perhaps a simpler problem could have been something like "Imagine a roll of tape where the tape itself has constant thickness x and length y. The width of the tape doesn't matter for this problem. We will calculate the thickness. The roll of tape is completely rolled up into a perfectly solid circular shape and a diameter of z. What is the formula for the thickness of the tape x expressed in terms of length y and 'diameter of the tape when rolled up in a circle' z? In coming up with the formula use the fact that the constant thickness doesn't change when it is unrolled from a circular to a rectangular shape."
With so much handholding, (and using the two-dimensional word circular rather than calling it a cylinder and rectangular prism which is what it really is) many more people could apply the formula correctly and get the result. But can ChatGPT?
I just tested it, this is how it did:
https://chat.openai.com/share/ddd0eef3-f42f-4559-8948-e028da...
I can't follow its math so I don't know if it's right or not but it definitely didn't go straight for the simplified formula. (pi times half the diameter squared to get the area of the solid "circle" and divide by the length to get the thickness of the tape.)
Mark my words, that is the sort of thing that Ilya saw months ago and I believe he decided they had achieved their mission of AGI. And so that would mean stopping work, giving it to the government to study, or giving it away or something.
That is the reason for the coup attempt. Look at the model training cut-off date. And Altman won because everyone knew they couldn't make money by giving it away if they just declared mission accomplished and gave it away or to some government think-tank and stopped.
This is also why they didn't make a big deal about those capabilities during the presentation. Because if they go too hard on the abilities, more people will start calling it AGI. And AGI basically means the company is a wrap.
All of the current LLM architectures have no medium-term memory or iterative capability. That means they’re missing essential functionality for general intelligence.
I tired GPT 4o for various tasks and it’s good but it isn’t blowing my skirt up. The only noticeable difference is the speed, which is a very nice improvement that enables new workflows.
I am not claiming that it is a full digital simulation of a human being or has all of the capabilities of animals like humans, or is the end of intelligence research. But it is obviously very general purpose at this point, and very human-like in many ways.
Study this page carefully: https://openai.com/index/hello-gpt-4o/ .. much of that was deliberately omitted from the presentation.
The character of Dory is jarring and bizarre precisely because of this trait! Her mind is obviously broken in a disturbing way. AIs give me the same feeling. Like talking to an animatronic robot at a theme park or an NPC in a computer game.
Goes well with the first observation in the shared article: "Lesson 1: When it comes to prompts, less is more"
I would dearly love to have an AI tool that I could trust to help with math. What is the state of the art? My math skills are very rusty (the last math class I took was calculus almost 40 years ago), and I find myself wanting to do things which would require a PhD level understanding of computer aided geometric design. If I had the magical AI which really understood a ton of math and/or could be fed the appropriate research papers and could help me, that would be amazing. So far all my attempts with ChatGPT 4 and 4o have been confusing because I don't really trust or fully understand the results.
This simple example and the frequency of wrong answers drives home the fact that I shouldn't trust ChatGPT for math help.
I'll inevitably be told otherwise by some ChatGPT-happy hypebro, but LLMs are hopeless when it comes to anything requiring reasoning. Scaling it up will lessen the chance of a cock-up, but anything vaguely out of distribution will result in the same nonsense we're all used to by now. Those who say otherwise very likely just lack the experience or knowledge necessary to challenge the model enough or interpret the results.
As a test of this claim: please comment below if you, say, have a degree in mathematics and believe LLMs to be reliable for 'math help' (and explain why you think so).
We need a better technology! And when this better technology finally comes along, we'll look back at pure LLMs and laugh about how we ever believed we could magic such a machine into existence just by pouring data into a model originally designed for machine translation.
Rude. From the guidelines:
> Please don't sneer, including at the rest of the community.
https://news.ycombinator.com/newsguidelines.html
"math help" is really broad, but if you add "solve this using python", chatgpt will generate code and run that instead of trying to do logic as a bare LLM. There's no guarantee that it gets the code right, so I won't claim anything about its reliability, but as far as pure LLMs having this limitation and we need a better technology, that's already there, it's to run code the traditional way.
So, umm, where's the savings? You can't not do the work to check the output, and a novice just can't check at all...
I have personally been brought into a coding project created by a novice using GPT4, and I was completely blown away by how bad the code was. I was asked to review the code because the novice dev just couldn't get the required functionality to work fully. Turns out that since he didn't understand the deployment platform, or networking, or indeed the language he was using, that there was actually no possible way to accomplish the task with the approach him and the LLM had "decided" on.
He had been working on that problem for three weeks. I leveraged 2 off-the-shelf tools and had a solve from scratch in under a full day's work, including integration testing.
You’re exactly right. It’s a weird example of a technology that is ridiculously impressive (at least at first impression, but also legitimately quite astounding) whilst also being seemingly useless.
I guess the oft-drawn parallels between AI and nuclear weapons are not (yet) that they’re both likely to lead to the apocalypse but more that they both represent era-defining achievements in science/technology whilst simultaneously being utterly unusable for anything productive.
At least nukes have the effect of deterring us from WW3…
I am a first year grad student and find it useful to chat about stuff with Claude, especially once my internal understanding has just gotten clarified. It isn't as good as the professor but is available at 2 am.
The point is that it’s got seemingly nothing to do with reasoning. That it can produce thought-stimulating paragraphs about any given topic doesn’t contradict that; chatting to something not much more sophisticated than Eliza (or even… yourself, in a mirror) could probably produce a similar effect.
As for chatting about stuff, I’ve been experimenting with ChatGPT a bit for that kind of thing but find its output usually too vague. It can’t construct examples of things beyond the trivial/very standard ones that don’t say much, and that’s assuming it’s even getting it right which it often isn’t (it will insist on strange statements despite also admitting them to be false). It’s a good memory-jog for things you’ve half forgotten, but that’s about it.
My response to attempts to use llms to do math is usually, use the right tool for the job.
I’ve found that I can work 100x faster with Mathematica this way and solve problems that I wouldn’t have bothered to attempt otherwise.
This is particularly effective for quickly visualising things, I’m too lazy to figure out all the graphing options for esoteric scenarios but GPT 4 can quickly iterate over variants given feedback.
could you elaborate?
I found chatgpt to pretty useful in high level understanding of where the topics stand in the grand scheme of things
kinda like a pre-google -- it helps me navigate what to google
let's ask gpt4o
You're an expert math researcher.
- Give me an overview of topos theory
- What are the main unsolved problems
- What is the point of it?
- What research directions are most promising?
- What are the math prerequisites to understanding it?
output Overview of Topos Theory
Topos theory is a branch of mathematics that generalizes set theory and category theory, providing a framework for analyzing and understanding mathematical structures. A topos can be seen as a generalized space that behaves much like the category of sets but with more flexible and abstract properties. Topos theory was developed by Alexander Grothendieck and his collaborators in the 1960s, initially to solve problems in algebraic geometry.
Key concepts in topos theory include:
Categories: Collections of objects and morphisms (arrows) between them.
Sheaves: Structures that systematically organize data over a topological space.
Topoi (plural of topos): Categories that satisfy certain axioms making them similar to the category of sets.
Main Unsolved Problems
Some of the main unsolved problems in topos theory involve:
Classifying all possible topoi: Understanding the full landscape of possible topoi and their interrelationships remains a complex challenge.
Connections to other areas of mathematics: Fully elucidating the connections between topos theory and other fields such as logic, algebraic geometry, and mathematical physics.
Homotopy theory in the context of topoi: Developing a deeper understanding of homotopy theory within the framework of higher topos theory.
The Point of Topos Theory
The point of topos theory is to provide a unifying and flexible framework for various mathematical concepts and structures. It offers a generalized setting in which both geometric and logical ideas can be expressed and related. This has significant implications for:
Algebraic geometry: Offering new tools and perspectives.
Mathematical logic: Providing a categorical framework for logic and set theory.
Theoretical computer science: Influencing the development of type theory and the semantics of programming languages.
Promising Research Directions
Some of the most promising research directions in topos theory include:
Higher Topos Theory: Extending the classical theory to higher categories and understanding its implications in algebraic geometry and homotopy theory.
Connections with Logic: Exploring further connections between topos theory and logic, particularly in relation to intuitionistic logic and type theory.
Topos-theoretic Approaches to Physics: Investigating how topos theory can be applied to quantum mechanics and other areas of theoretical physics.
Descent Theory and Stacks: Further developing the applications of topos theory in descent theory and the study of stacks in algebraic geometry.
Math Prerequisites to Understanding Topos Theory
To understand topos theory, one should have a solid background in the following areas:
Category Theory: Fundamental concepts such as categories, functors, natural transformations, limits, and colimits.
Set Theory: Basic set-theoretic concepts and operations.
Algebraic Topology: Basic knowledge of topological spaces, sheaves, and cohomology.
Algebra: Basic group theory, ring theory, and module theory.
Logic: Understanding of basic logical systems, particularly intuitionistic logic.
With these prerequisites, one can start delving into the more advanced and abstract concepts of topos theory.
not perfect but good enough to get startedI've found a ton of research papers and information, but most of it is quickly beyond my ability to digest.
For G2 constraints, there is simple equation:
K(t0) = ((n-1)/n)*(h/a^2)
Where n is the degree of the curve, a is the length of the first leg of the control polygon, and h is the perpendicular distance from P, to the first leg of the control polygon. K(t0) is the curvature at the end point of the adjacent curve.
Depending on what you want to do, it's easy to solve for K(t0), a or h. I would like something this simple for G3.
otherwise you can say "why do you need google, it's the same as you'll get from the website"
moreover, I found that chatgpt is pretty decent at rephrasing a convoluted concept or a paragraph in a research paper, or even giving me ideas on the research directions
I mean, same with coding -- I treat it as a smart autocomplete
I could go to google and look for a .csv containing a list of all US States
Or, I can write
const US_STATES = [
and let copilot complete it for me -- 5 minutes saved?1. Text only prompt
2. Text + Image with supplemental data
3. Text + Image with redundant data
Case 1 generally performs the best. I also found that reasoning improves if I convert the equations into Latex form. The model is less prone to hallucinate when input data are formulaic and standardized.
Case 2 and 3 are more unpredictable. With a bit of prompt engineering, they may give out the right answer after a few attempts, but most of the time they make simple logical error that can be avoided easily. I also found that multimodal models tend to misinterpret the problem premise, even when all information are provided in the text prompt.
[1] Text prompt only, run 2; prompt and image, run 3.
I feel like math is naturally one of the easiest sets of synthetic data we can produce, especially since you can represent the same questions multiple ways in word problems.
You could just increment the numbers infinitely and generate billions of examples of every formula.
If we can't train them to be excellent at math, what hope do we ever have at programming or any other skill?
It’d be interesting to obfuscate it so that the concepts are all the same but the subject matter would be hard to google for. Another failure mode for LLM’s is if you take something with a well-known answer and add a twist to it.
"Therefore, the thickness of the tape is approximately 0.000589 cm or 0.589 mm."
Most people learn to avoid that person that is wrong/has bad judgment and is arrogant about it.
Not only do LLMs not know some things, they don't know that they don't know because of a lack of true reasoning ability, so they inevitably end up like Peter Zeihan, confidently spouting nonsense
That heavily depends on the individual grader/instructor. A good grader will take into account the amount of progress toward the solution. Restating trivial facts of the problem (in slightly different ways) or pursuing an invalid solution to a dead end should not be awarded any marks.
impressive attempt though, it used number of wraps which I found quite clever
"Consider the following word problem: "A 100 meter long chain is hanging off the end of a cliff. It weighs one metric ton. How much physical work is required to pull the chain to the top of the cliff if we discretize the problem such that one meter is pulled up at a time?" Note that the remaining chain gets lighter after each lifting step. Find the equation that describes this discrete problem and from that, generate the continuous expression and provide the Latex code for it."
It has gotten quite impressive at handling calculus word problems. GPT-4 (original) failed miserably on this problem (attempted to set it up using constant acceleration equations); GPT-4O finally gets it correct:
> I am driving a car at 65 miles per hour and release the gas pedal. The only force my car is now experiencing is air resistance, which in this problem can be assumed to be linearly proportional to my velocity.
> When my car has decelerated to 55 miles per hour, I have traveled 300 feet since I released the gas pedal.
> How much further will I travel until my car is moving at only 30 miles per hour?
* the tape is perfectly flexible
* the tape has been rolled with absolutely no gap between layers?
However we don't know the internals of ChatGPT-4, so they may be using some agents to improve performance, or fine-tuning at training. I would assume their training has been improved IMO.
To calculate 7^1.83 , you can use a scientific calculator or an exponentiation function in programming or math software. Here is the step-by-step calculation using a scientific calculator:
Input the base: 7 Use the exponentiation function (usually labeled as ^ or x^y). Input the exponent: 1.83 Compute the result. Using these steps, you get:
7^1.83 ≈ 57.864
So, 7^1.83 ≈ 57.864
Given this, and the recent announcement of data analysis features, I’m guessing the GPT-4o is wired up to use various tools, one of which is a calculator. Except that, if you ask it, it also blatantly lies about how it’s using a calculator, and it also sometimes makes up answers (e.g. 57.864 — that’s off by quite a bit).
I imagine some trickery in which the LLM has been trained to output math in some format that the front end can pretty-print, but that there’s an intermediate system that tries (and doesn’t always succeed) to recognize things like “expression =” and emits the tokens for the correct value into the response stream. When it works, great — the LLM magically has correct arithmetic in its output! And when it fails, the LLM cheerfully hallucinates.
1. https://www.vanityfair.com/news/2023/04/elon-musk-twitter-st...
https://ny1.com/nyc/all-boroughs/technology/2023/10/12/x-say...
Peak humanity.
The only group of people more delusional than the AI doomsday screamers are those who think playing around with LLMs is "engineering".
Or worse, a 20-deep Twitter thread.
Do you know why? Your blog post seems thoughtful and interesting and doesn't include anything that seems ban-worthy.
And i'm pretty sure the average person when asked would say the same thing and be like "duh" even though technically based on the minutia it's incorrect.
It's treated like a genius and that's what it gets measured against.
to me, the impressive part of gpt is being able to understand the image and extract data from it (radius information) and come up with an actual solution (even though it got it wrong a few times)
for basic math I can do
python -c "print(6/9)"