How is ChatGPT's behavior changing over time?
arxiv.org
arxiv.org
Second, I think the code generation bit of this paper is blown out of proportion. The code can't be immediately injected into a codebase due to a formatting change (triple quotes). I'd be more interested in changes to the quality and performance of the code generated, not whether it can be easily copy/pasted from the page.
I could not replicate that result (but could replicate the other math issues). I've also never seen the triple quotes in results so it's unclear if there was a temporary presentation bug.
EDIT: removed redundant bit
Quality is important to humans, because humans have to read it, but correctness is what people using ChatGPT for code actually need. So long as the quality and performance is good enough, then it will be useful.
Performance is such a nuanced topic that you need very context aware devs anyway, and I think a general purpose LLM is never going to have that kind of awareness.
Be careful with that goalpost, it might make sudden movements.
Now, at least for some projects, I can give it a few rapid bullet points in incomplete sentences and have that first draft in seconds, after which I just need to tweak, add in tables and more detailed stats & results, etc. quite useful.
My mental capacity was used elsewhere - using chatGPT let me answer an important customer question authoritatively.
$ random_bytes() { xxd -plain -c 0 -l "$1" /dev/urandom; }
$ random_bytes 32
e6a4a7bbea69a0164cbb66c89f8f528af93c6d2459fd28d2640e2952c031b618Hmm, I’ve been using /usr/games/fortune
Is this not a best practice?
It's far easier even for an expert to communicate what they want in natural language than it is in a formal syntax for all but the most trivial things.
It should be a goal of these tools to do this correctly.
agree completely, but the LLM should be focused, then, strictly on formulating an "execution plan" of sorts and handing that off, not on performing math itself.
In other words, when asked "if i have 349 blueberries and one blueberry turns into a cherry per hour, how many of each fruit will I have in 93478 minutes?" it shouldn't be doing the actual arithmetic, but it should be figuring out what arithmetic would need to be done.
You describe a complex relationship between several nodes(people, cities, etc), and then ask ChatGPT to draw the relationship as a graph data structure. It will create a formula that Mathematica can render, and then send the formula to Mathematica before presenting an ascii drawing. Usually. Sometimes it just complains that it can't draw and explains what the graph looks like to you in a written formal/human-readable syntax.
In other words, yes it just summarizes the problem, converts it to a formal syntax, and sends it off to some other tool.
If it is bad at math and can’t be taught, then you have a fundamental problem. It’s a matter of time before this limit gets hit in other domains.
Or it is just pretending to do so. And since it pretends, of course it trips over all small things as it does not understand them.
And in this case, the shape of the answer is often right; it just makes ... ordinary errors. Ironically, the AI is a lot better at high-level thinking than correct calculation.
That would be actually a human like feature .. except I do not consider what LLMs are doing as thinking.
100,000 + 987 - 1444 * 25,945.842 / 0.0042
becomes
"100" one hundred
"," comma
"000" triple zero
" +" space plus
" 9" space nine
"87" eighty seven
" -" space minus
" 14" space 14
"44" forty four
" *" space times
" 25" space twenty five
"," comma
"9" nine
"45" forty five
"." period
"8" eight
"42" forty two
" /" space divide
" 0" zero
"." period
"00" double zero
"42" forty two
Now imagine someone reading that to you over the phone once and asking you to do the math in your head and you aren't allowed to use paper and pencil and you have to get it right the first time.
Through continuous use, I have found that it does not "reason". That doesn't mean it's not valuable in many ways, and I have found it to be very helpful in a multitude of diverse applications, including helping me reflect on my own life through my own interpretations of its output. It's also a great interface for JSTOR, wikipedia, and basically any language learning.
I'm having a hard time making the jump from "this must be a calculator" to "this must be a philosopher" to be useful. When did we ever have those requirements for a tool?
This tool is just not made for math. Most of its logic processing abilities seem to surpass mine if I am only given 5 minutes to understand a problem. If you understand the tool, you will get the most out of it. Stop anthropomorphizing it, and stop pretending that it can't generate both highly beneficial or highly harmful content simply because it doesn't have a soul/d*ck or whatever.
Using GPT to do maths is probably like using a 737 to drive around on the ground.
Teaching GPT to do maths might be like teaching a child the times tables - a skill that can help overall reasoning.
Humans are different to an AI, but putting that aside, my intuition would be that if we never taught kids any mental maths, their concept/understanding of numbers would be fundamentally different to how it is if they learn that 9 x 9 = 81 (also look at how your fingers move - there is a relationship there!).
But who knows, AI is strange and there's lots of stuff that needs to be experimented with. I would think training an intuitive sense of numbers would have other fall-outs though. This is half the beauty of LLM's right? You show an LLM some history books and it also learns about biology, politics, grammar, etymology and love. You teach a LLM maths and it also learns ... ?
One significant difference is that in both of those examples it is (or quickly becomes) plain why it’s a ridiculous idea. Even if you don’t understand it yourself, you’ll get external feedback fast. Not so with LLMs, where even people with technical needs may fail to see what is or isn’t a good use of the tool. Case in point: https://news.ycombinator.com/item?id=36782446
“You’re holding it wrong” isn’t a valid argument in perpetuity. At a certain point it becomes the fault of the designer, not the user.
Because that is how it works. The limiting factor so far has been the smart persons time and patience. Now, no longer.
People moderating their LLMs usage is never happening, from here on out until the end of civilisation. Any LLM service that is designing for that is done. You need to make lazy questions efficient. People do not care about how complicated your sql query is and they will never care. People will not give up on energy, meat, cars, as long as they feel they are giving something up.
People will never think twice to not make your LLM think twice.
If it seems useful and convenient, people will use it. If it's not giving good answers to lazy questions out of the box, they will go to the thing that does.
A lot of words to make a big deal out of nothing. All that is needed is some new abstracted layer that identifies a math question and then proxies it over to the wolfram plugin. That’s it
We don’t have crazy debates over whether a polygon should be rendered by the cpu or a gpu. We solved this problem
I don't think this is a great analogy. if your 737 couldn't drive on the ground and your astrophysicist couldn't answer basic maths questions I wouldn't want to fly in that plane or put much faith in the astrophysicists answers to more complex questions.
Maybe maths is not a particular strength of LLMs, but asking questions where it is easy to judge the factual accuracy of the responses seems a pretty reasonable test to be running.
It isn't reasonable if that isn't what the system was designed to do.
It would be a poor test of my general practitioner's competence to ask him calculus questions and conclude he doesn't know what he's talking about because he can't answer them.
The astrophysicist can do long division in his head, but he'll be about as fast and accurate as the next person, because he doesn't practice arithmetic every day.
I agree with the commenter somewhere in my thread who said all an LLM should be optimized for is to classify the type of problem and feed it into a purpose-built, deterministic solver that is trained to interpret math as math and not as language, be it ML-based or algorithmic.
Those are completely different ideas.
On a serious note, most of these AIs are bad at math but good at writing code for calculators. So what you'll be benchmarking is their ability to create and use tools.
FWIW I use GPT-4 regularly to explain Koine Greek from the New Testament to me; its ability there certainly hasn't diminished in the last two months.
I'm not an AI/ML scientist, so I may be way off mark here, but everything I've read so far, and all my experience playing with GPT-3.5 and GPT-4, tell me that comparing performance of an LLM to that of a human is a category error, because the LLM isn't a good analogue of a whole human mind - but it's a very good analogue to human inner voice. The stream of consciousness. The whatever-it-is that surfaces your unconscious/subconscious thought process in form of words and sentences.
The inner voice is fast, it's reactive. It generates thoughts that match the situation, whether they're correct or factually accurate or not. It's up to the conscious part of your mind to stop, refine, or recycle those thoughts. If you let it keep going, it'll give you thoughts based on what feels like should follow the thoughts that came before. And, unless you habituated responding to anything new with "I don't know" followed by ignoring the topic, the inner voice will start blurting answers to what looks like a question/problem statement; whether or not they'll make any sense, depends on your familiarity with the topic in question.
Pretty much 1:1 what LLMs do.
Now, this could all be noise, but I don't think so. I know not everyone has a distinct inner narrative (much like not everyone can visualize things in their mind - I can't), but many (most?) people do. The description of the "inner voice experience" I gave above is something I figured out over a decade ago - before LLMs or even deep learning were a thing, before I knew anything about the NLP beyond recognizing the term "Markov chain" is somehow related. Could my inner narration style be unique? Possibly, but given how advice to avoid connecting your inner voice directly with your vocal apparatus is deeply infused in culture and literature, I strongly suspect this is just how it works.
All this to say: it is my hypothesis, so far corroborated by experience, that when you start feeding absurd amount of unlabeled text to a transformer model, letting it pick up on the structures encoded within, what you get is a close equivalent to our own inner voice - the part that deals with associations, not logic or data storage. You can't expect it to get good at performing arbitrary computation or recalling data with perfect fidelity, because it's structurally not what it's suited for. For humans, performing arbitrary calculations or perfect recall requires engaging a slower, more algorithmic thinking process (and/or external memory). That part is currently missing in the LLM-based AI systems we're playing with.
It will be much more compute intensive, as each response will probably require multiple context windows and distillations.
What it is bad at is actually performing the steps accurately, but as others mentioned, that's where Wolfram and/or a code interpreter would come in.
``` code ```
They report 2 degradations: code generation & math problems. In both cases, they report a behavior change (likely fine tuning) rather than a capability decrease (possibly intentional degradation). The paper confuses these a bit: they mostly say behavior, including in the title, but the intro says capability in a couple of places.
Code generation: the change they report is that the newer GPT-4 adds non-code text to its output. They don't evaluate the correctness of the code. They merely check if the code is directly executable. So the newer model's attempt to be more helpful counted against it.
Math problems (primality checking): to solve this the model needs to do chain of thought. For some weird reason, the newer model doesn't seem to do so when asked to think step by step (but the current ChatGPT-4 does, as you can easily check). The paper doesn't say that the accuracy is worse conditional on doing CoT.
The other two tasks are visual reasoning and answering sensitive questions. On the former, they report a slight improvement. On the latter, they report that the filters are much more effective — unsurprising since we know that OpenAI has been heavily tweaking these.
In short, everything in the paper is consistent with fine tuning. It is possible that OpenAI is gaslighting everyone by denying that they degraded performance for cost saving purposes — but if so, this paper doesn't provide evidence of it. Still, it's a fascinating study of the unintended consequences of model updates.
The obvious avenue to degradation is that the "HR personality" is much more strictly applied and the resistance to being jailbroken is also in some sense an inability to think.
This is not necessarily the case, and even if it is doesn't imply gaslighting as compared to inability to measure.
Logan: The API does not just change without us telling you. The models are static there. https://twitter.com/OfficialLoganK/status/166393494793189785... may 31
Peter: No, we haven't made GPT-4 dumber. Quite the opposite: we make each new version smarter than the previous one. https://twitter.com/jlowin/status/1679660938415177731 july 14
either the models are static, or they are being improved continuously and there have been unforeseen regressions. only one can be true at any point in time. was this policy changed in the last 1.5 months?
1. https://arxiv.org/abs/2302.01318
2. https://arxiv.org/abs/2207.07061
At the end of the day 99% of the confusion comes from people using the web interface, which undoubtedly does change much more often than the API versions they share.
The web app they host isn't a simple API wrapper, it does summarization, has some sort of system prompt, and calls the moderation API. That's undoubtedly being updated all the time.
> As of July 3, 2023, we’ve disabled the Browse with Bing beta feature out of an abundance of caution while we fix this in order to do right by content owners. We are working to bring the beta back as quickly as possible, and appreciate your understanding!
https://help.openai.com/en/articles/8077698-how-do-i-use-cha...
> At the end of the day 99% of the confusion comes from people using the web interface, which undoubtedly does change much more often than the API versions they share.
The API does not offer any browsing features, that's the web app.
1. https://chat.openai.com/share/44a0c5b6-c629-470a-992f-8cdbbe...
If you intentionally smear the line between their web app which is chock full of optimizations to even let it function as it does (the web app's max conversation length exceeds the context window) and the API which is versioned and iterated on in the open... it's either a lack of understanding or FUD.
Have you seen the 25 messages/3 hours limitation for GPT-4? Why do you think they did that? Of course they would make more money scaling up the volume, but how to do that when compute is so limited? Of course, by using some kind of approximation - quantised model or speculative sampling come to mind. It's hard to pinpoint model regressions, but scaling up volume is great, one more incentive to do it.
The web app is a consumer app (B2C) the api is commercial (B2B). They tinker with the B2C app because it's already a lossy approximation of using the model between the summarization and system prompt.
They cannot mess with the commercial offering willy-nilly: People are building businesses predicated on it behaving a certain way. That's why there are dated version that you can pin to with the API. The web app changes whenever they feel like it.
If can't tell if that's about the API or the web app, I don't think you're familiar enough with the subject to speak on it.
Millions of dollars in spend from products predicated on the API not randomly changing?
The actual B2B offering is handled by Microsoft, via Azure OpenAI. Same models, but deployed on Azure - meaning they come with SLA and all the right protocol and compliance stuff, so that your people can negotiate with their people - and if you're willing to spend enough, you'll get the models for yourself. Not the weights, of course: just no training on your inputs, not even retaining inputs for 30 days "because ${legal reasons}" - instead, you can pick and chose, fine-tune and deploy OpenAI models on your own tenant, and basically manage everything except the weights themselves.
- It has the same 30 day retention for legal reasons unless you manually request (just like OpenAI)
- You can't fine tune any models that you can't fine tune on OpenAI, and in fact default access is a subset of what OpenAI offers.
- "and if you're willing to spend enough, you'll get the models for yourself" is a bit of nonsense, Azure OpenAI forces everyone to make a "tenant", that's just for VPC stuff to work. Outside of that it's bog standard fine tuning and at most "on your data" which is a wrapper for chunking + vector embeddings
- Azure OpenAI has a narrower built in filter that you can't modify without again, a separate request.
Azure OpenAI overall is mostly for companies that need to signal to other companies that they're using Azure: it's no more commercial than the OpenAI offering.
On the contrary, I am using Azure OpenAI daily at work, and I'm explicitly not allowed to use "regular" OpenAI offerings.
> It has the same 30 day retention for legal reasons unless you manually request (just like OpenAI)
It doesn't, at least not for us.
> Azure OpenAI has a narrower built in filter that you can't modify without again, a separate request.
I'm not sure if it's narrower, but it is there and I have a strong suspicion that MS is just trying to extract additional rent from companies that really want to turn the filter off.
> Azure OpenAI overall is mostly for companies that need to signal to other companies that they're using Azure: it's no more commercial than the OpenAI offering.
No. Azure OpenAI is for companies that don't play fast and loose with data - their own data, and their customer data. Of course, most companies don't give a damn, but for big enough companies, or those operating in certain industries, there are actual, severe legal consequences for mishandling the data, and such companies don't have the option to just not give a fuck and dance with OpenAI - they need to sign an actual contract with a serious entity that understands regulatory compliance, and how corporations tick. Microsoft is such entity. OpenAI isn't.
So then you filled out the request because the default is exactly the same as OpenAI: retained unless you manually apply for an exception.
https://customervoice.microsoft.com/Pages/ResponsePage.aspx?...
> I'm not sure if it's narrower, but it is there and I have a strong suspicion that MS is just trying to extract additional rent from companies that really want to turn the filter off.
You don't need to question if it's narrower, OpenAI used to surface it as an API separate from the moderation API and it's much stricter by design.
> No. Azure OpenAI is for companies that don't play fast and loose with data - their own data, and their customer data...
I don't know if you actually believe this or you're just not aware, but the companies that don't play fast and loose aren't using OpenAI period: Azure flavored or otherwise.
OpenAI has SOC2, GDPR and CCPA compliance. They comply with HIPPA and offer BAs. They sign DPAs on a case-by-case basis same as Azure.
You're pretty much proving the value of Azure in your comment: it's a veneer of familiarity that coaxes people who are convinced the new kid on the block must be untrustworthy.
If OpenAI can't promise something Azure can't either: They're entirely dependent on OpenAI for this. Every idiosyncrasy behind Azure OpenAI maps back 1:1 to OpenAI.
They do, and that's the biggest value proposition of Azure OpenAI right now: strong contractual guarantees, from a reputable partner (that's easy to hit with lawsuits should they go rogue :)).
The current situation is that it's pretty unwise for any company to ignore GPT models. OpenAI itself is a wildcard, but getting the same from Microsoft isn't "playing fast and loose with data" any more than using Windows and Office 365 across the organization is. Most large corporations and governments have been building their office work and communication around those tools for decades now, so - questions of antitrust aside - all the kinks have been worked out. I don't think you appreciate how big a difference this makes.
I mean, it's either that or all the company communication I got on this was bullshit.
> OpenAI has SOC2, GDPR and CCPA compliance. They comply with HIPPA and offer BAs.
That's the first I hear of it, but since I never dealt with OpenAI itself on that level, I accept this was my ignorance speaking; thanks for clarifying.
> You're pretty much proving the value of Azure in your comment: it's a veneer of familiarity that coaxes people who are convinced the new kid on the block must be untrustworthy.
I think you're underestimating the importance of this. What you call "veneer of familiarity" translates to billions of dollars of differences in terms of security risk.
As mentioned before, MS has been in this space for a while, and has decades of trust and experience built with governments and corporations and other big organizations. Microsoft is a known, trusted quantity. That alone is worth a lot.
But then, there are also technical aspects too - like how deploying to a tenant on Azure integrates properly with all the other services you use to run half the company. In practical terms, this means all use is monitored and auditable by in-house teams, and all the in-house policies are being enforced. OpenAI can't begin to offer this level of integration - they have neither technical nor legal resources for that.
> If OpenAI can't promise something Azure can't either: They're entirely dependent on OpenAI for this. Every idiosyncrasy behind Azure OpenAI maps back 1:1 to OpenAI.
None of that matters here. The models are what they are - peculiar large matrix multiplication as a service. By themselves, they're pretty much pure functions. The part that matters is operations - both technical and legal aspects - and this is where Microsoft and OpenAI are independent and have different offerings.
Also, looking at the way money flows, I think it's OpenAI that's dependent on Microsoft right now, not the other way around. They kinda pretend to be just friends with benefits, but it's obvious who the dependent party is.
Companies that don't play fast and loose are not using LLMs yet. They use "old school" ML at most with much narrower scope because at this point it's simply less of a liability.
You seem to think I'm underestimating what Azure's name adds to OpenAI: I fully understand how bureaucratic organizations work off vibes under the guise of name recognition and my point is I simply have no respect for it.
If you genuinely care about customer data, then the value of being able to sue MS instead of OpenAI is moot. You also probably aren't going to use a service that shouts from the roof tops about not using your data then quietly keeps it for 30 days unless you manually opt-out. You probably don't use some model with unsolved copyright/PII questions. And a million other unknowns
> The part that matters is operations - both technical and legal aspects - and this is where Microsoft and OpenAI are independent and have different offerings.
You might want to check OpenAI's subprocessor list if you think that they're not the same technically...
https://platform.openai.com/subprocessors/openai-subprocesso...
And Azure's subprocessor list is a superset of that list, not a subset.
> No, we haven't made GPT-4 dumber. Quite the opposite: we make each new version smarter than the previous one. > Current hypothesis: When you use it more heavily, you start noticing issues you didn't see before.
I don’t see how it supports your argument. Your comment says “they deny making changes to GPT-4”, and the tweet says “we are making incremental improvements to GPT-4”.
Personal hypothesis is that they have made a few changes to ChatGPT recently - possibly quantization, and almost certainly some tweaks to make it give shorter/less detailed answers.
But by the nature of a probibalistic tool being run by a secretive company, it's hard to say for sure. Maybe I and everyone else complaining have just started to get unlucky answers.
E M Forster, "The Machine Stops", 1909
These are to represent markdown code blocks, which probably helped their front end developers, but hinders copy-pasta
I wonder if they'll start using LLMs while ingesting new data. eg asking the LLM if the content is helpful, cites sources, respectful, positive, not-thin content, common or often duplicated content, etc etc, before each content import.
When GPT's details were allegedly leaked [0] I read about the mixture of experts and wondered if this explains my recent distaste for GPT. Lately I have been using Anthropic's offering (mainly since it is free and has a 100k context window) and I've been surprised at just how well it reasons and understands what I'm asking. It responds like GPT 4 used to, and speaks to me with nuance I haven't seen since Bing released their original chatbot. I still find GPT 4 better if I want to fine-tune the model and make it adopt a persona-- Claude will almost always refuse.
Either way, I'm a bit confused with the way GPT 4 has been changing over time. It seems the team made significant changes to the model quality in favor of performance. Whether accuracy or performance is more important is up for debate, but something is clearly changing.
They could also use these parameters over time to become more cost efficient, in general, or reserve GPUs (i.e. allocating more expert networks) for some higher-margin use-cases.
Sort of like having a problem and then asking both a biologist & a chemist for their assessment. But if asked a single biochemist they could more readily provide an answer that synthesized both more comprehensively, even if you got both the biologist and the chemist in the same room with each other.
You can have your model talk like HR or think like a mad scientist, but not both equally well.
I would also like to not be treated like an idiot. Every time I query anything related to health or medicine, chatGPT will give me a short generic answer, then add two paragraphs of warnings about how I should just seek a medical professional's help. As if I'm the kind of moron who will blindly follow whatever some AI tool tells me and not actually go to a doctor if there's something wrong with me.
This organization seems like it's being run by (scared) lawyers.
It's sad, but I understand why they did it. It was too much power in the hands of people.
They should say what you are paying for in terms of technicals, either spill the beans on the architecture or have some SLAs on how good the thing is.
While some of those are anti-customer on their face, the “we shouldn’t commit to implementation details and instead just to our API/feature surface” may seem anti-customer while actually making things much better for customers. When implementations are supported in perpetuity development grinds to a standstill. Even overcommit ends up helping customers - cloud companies can offer lower sticker prices.
In the case of both regular Public cloud and OpenAI - if you need stability beyond publicly supported APIs you are probably not a good fit as a customer and should instead find a company willing to commit to lower level implementations (eg bare metal, truly major/minor versioned software) or roll-your-own.
Ideally the same across clouds, or more likely cloud-specific.
Or even re pin it every 5 years. E.g. a 2020-CPU, 2025-CPU etc.
We need a "horse power" for the cloud I guess!
And as I know there will be many commenters asking, here's my experience.
Four months ago I released a mobile app that wraps the "Open"AI API and allows a more private use. The app also includes 15 domains with over 150+ editable prompts, specifically crafted to help users get the most out of GPT.
I initially crafted these prompts in English but given I wanted to allow also non-English speakers to make use of them, I started to translate them in Italian and German.
I wrote a small script that took each English prompt and, with some more prompt-fu, translated them to these languages, using GPT-4.
As I'm fairly fluent in both Italian and German, I was able to verify the quality of these translations.
Well, around end of May, when I first noticed some weird answers from GPT, I ran my script again and (surprise!) the quality of the translations is visible inferior.
On a different note, I noticed that from one day to the other, any prompt would get an empty reply in the app, only to discover that "Open"AI has also made a subtle change in the json format of the API response that made broke the parser in the app.
Admittedly, 150 phrases is not a huge sample size and I could/should have used the default json parser.
I don't trust "Open"AI and will, as soon as it's feasible, change to a different model.
I repeated this last week and it's still getting the phonetic content right, it's now often getting the tones wrong (something which I literally never saw on release).
Perhaps changing the prompt would deliver the correct results, but that's not really the point. Something has pretty clearly changed such that old prompts no longer delivering the same results.
Isn't it the equivalent to saying "Here's the top 10 results for google searching golden retrievers March 2023, and here's the top 10 results from June 2023. We see that google is returning even cuter animals today. Unfortunately though, one of the results linked to a page full of cats."
I'm sure openai has a list of standard questions that it tracks the responses it is getting over a period of time with perfect knowledge / version timing of their own releases.
This does not seem like valid research / publication.
Do you really feel like the efforts were significant or meaningful?
I could read the Pricicipia, or anything, and dumb it down to dogs and cats, but that reduction to the dumb would be a failing all my own.
As for the test-- a classification task benchmarked with 500 examples is a fairly decent test of a system's capabilities, be it LLM or a traditional machine learning model. And again, this was only one of a variety of tasks. While it's certainly no Principia, I don't know how you get from there to dogs and cats, nor have you explained your reasoning on how you found the path between the two.
Okay okay hold on back to the drawing board.
...dumbed down to cats and dogs.
…Principia and failings
Which seems valuable for other research using OpenAI GPT directly.
Example: you get a bunch of data from your boss, need to find something out, you prompt engineer the results. Your boss get's back to you (five weeks later), likes the result, just want's you to fix a minor thing. It's now giving you completely different results. It's like building on sand.
And has anyone successfully used Copilot to do simple but tedious things like generate library bindings for various languages? I suspect that Copilot would be good at this kind of thing because it is a relatively simple task and there is a lot of training-code available.
It used to be better in the past.
> The math questions were of the form “Is 17077 prime”? They picked 500 numbers, but all of them were prime!
> The June version of GPT-3.5 and the March version of GPT-4 almost always conclude that the number is prime regardless of whether it is prime or composite. The other two models do the opposite. But the paper only tested prime numbers, and hence concluded that GPT-3.5’s performance improved while GPT-4’s degraded.
There's a huge gap between GPT-3.5 and 4, put there by a massive amount of money, from my understanding.
To compete, with open source projects being less well funded, I would assume that orders of magnitude improvements in training cost would be required. What do you see driving this, and who do you see paying for it?
If Meta, or anyone else, gets something that beats GPT-4, I would naively assume they would monetize it. If someopen source effort manages to beat GPT-4, then I assume some well funded, profit seeking, entity will dump orders of magnitude more money into whatever enabled the open source offerings to compete.
I don't think GPT-4 is the pinnacle of OpenAI, or that their funding will run dry, so they shouldn't be viewed as a static entity.
My bet is that Meta has pivoted almost entirely to this space with their R&D in the last six months. Llama 2 is spectacular. And with its' success, there will undoubtedly be more. They also happen to have access to limitless amounts of compute, cash, and engineering that puts OpenAI to shame. This could finally be their chance to create a platform for real. And the open source community and startups will benefit off of that.
Owning a platform. Zuck's dream.
They put tons of money into these models, nurture an ecosystem of companies built around them, and then start gradually figuring out a licensing model for the ones that take off.
Furthermore, Llama remains well below GPT-3 on human rated tests such as programming, and GPT-3 is already over three years old. It is also misleading to suggest Llama 2 can be ran on consumer hardware - the smaller and quantized models can but those are even more lacking in capability. Full power Llama 2 still requires multiple kilowatts of electricity and $10,000+ of compute hardware per inference session.
OpenAI does not have a moat, but they do have a very high wall.
What hardware would you need to run it at home?
Step 1: https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGML/blob/ma...
Step 2: https://github.com/ggerganov/llama.cpp
Step 3: you're welcome
That is not true. A common macbook with lots of RAM (>32GB) is enough. Or any x86 computer with lots of RAM. llama.cpp is CPU only and quite fast
Note that I never said "open source" just "open models". As in, I can now actually build things with GPT3 capability that run locally. I couldn't care less about the model code.
>Llama 2 still requires multiple kilowatts of electricity and $10,000+ of compute hardware per inference session.
I'm running llama-2-7b-chat on my 8GB M1 Mac right now with llama.cpp. Completions are instant, and essentially at GPT3 levels of accuracy.
The higher param models require up to 64GB RAM, but it's all CPU based.
This is not what the Llama 2 license says [0]. There is a cap on the number of active users of products (any products, not just ones that may make use of Llama) by the company who plans to use Llama 2 as of Llama 2 release date.
Not a “future cap on number of users of your Llama2-based product”. Also, the cap is 700 million users.
> If, on the Llama 2 version release date, the monthly active users of the products or services made available by or for Licensee, or Licensee's affiliates, is greater than 700 million monthly active users in the preceding calendar month, you must request a license from Meta
[0] https://github.com/facebookresearch/llama/blob/main/LICENSE
It's just not part of the entended use-case, and hasn't been trained to do so. It's almost like complaining that StableDiffusion isn't good at text generation…
> It is also misleading to suggest Llama 2 can be ran on consumer hardware - the smaller and quantized models can but those are even more lacking in capability
The biggest models will be abble to run on the CPU just fine with llama.cpp as long as you have enough (cheap) RAM. Sure it's slow, but you can run it.
> Full power Llama 2 still requires multiple kilowatts of electricity and $10,000+ of compute hardware per inference session.
What is that “per inference session” doing here? You pay the hardware only once you know… (and the number of kilowatt isn't per inference session either, the number of Watt•hour is)
For the most part none of these systems were trained to do anything. The capabilities are emergent. My statement about the benchmarks remains unfazed.
>The biggest models will be abble to run on the CPU just fine with llama.cpp as long as you have enough (cheap) RAM. Sure it's slow, but you can run it.
Yes, but what products/clients are you pitching where that kind of wait will be acceptable? Time is money. Standing up racks of maxed out RAM server slots is still far from inexpensive too.
>What is that “per inference session” doing here? You pay the hardware only once you know… (and the number of kilowatt isn't per inference session either, the number of Watt•hour is)
Good job at completely missing the point. I know what a kilowatt is and it is the unit I meant, not kilowatt-hour. I'm referring to the hardware necessary to run one user session of inference. While user 1's tokens are generating, users 2 and up are in the queue waiting. If you want to serve multiple users simultaneously you will need to invest in multiple $xx,xxx units of hardware each requiring a multi-kW electrical circuit.
There is a reason ChatGPT incurs electrical bills on the order of a million dollars per day.
The base models yes, but the chat version have been explicitly tuned for specific use-cases, it's not just basic next-token-prediction.
> Time is money
No, time is time, only context gives it monetary value: a single ms in some context can be much more expensive than an hour in another.
> Standing up racks of maxed out RAM server slots
Why are you talking about servers we're talking about the ability to run things on consumer hardware here…
> There is a reason ChatGPT incurs electrical bills on the order of a million dollars per day
Because they spend many MWh, to run all of their customers' requests, but when you're running your own requests on your own hardware, then the energy consumption is tiny, even if you needed a lot of instantaneous power.
PS: Which is obviously ripping off paying users.
Competition is the answer.
“ Sure, let's go step by step.
A prime number is a number greater than 1 that has no positive divisors other than 1 and itself. This means if we can find any other number (excluding 1 and the number itself) that divides 17077, then it is not a prime number.
Let's start by checking divisibility by 2. Since 17077 is an odd number (its last digit is 7), it is not divisible by 2.
For the next step, we can look at divisibility by 3. We add up the digits: 1+7+0+7+7 = 22. Because 22 is not divisible by 3, 17077 is also not divisible by 3.
The smallest prime number greater than 3 is 5. Checking divisibility by 5, we see that 17077 does not end in 0 or 5, so it is not divisible by 5.
The smallest prime number greater than 5 is 7. To check divisibility by 7, there is no simple rule like with 2, 3, and 5, so we would perform the actual division. But we can stop here because I just recalled a rule that we can only check up to the square root of the number and the square root of 17077 is approximately 130.68. So if there were factors other than 1 and 17077, we would have found them by now.
Therefore, 17077 is a prime number.”
Did it? Did you read what it wrote?
> Look at this amazing capability, it gets it perfectly! (screenshot of interaction with clear flaws/outright wrong answer)