GPT-4 is coming next week – and it will be multimodal, says Microsoft Germany
heise.de
heise.de
If it's true, then we're going to have something very similar to the Star Trek computer by the end of the decade at the latest.
If it's false, then we're probably going to enter another AI winter this decade.
I suspect there will be a lot of resistance as it “succeeds.”
I kind of get it, too. If I made my money making art this would scare me. What I wonder is what's gonna happen when it's [ artists, accountants, copywriters, bloggers, designers, programmers, etc, etc, etc].
I see problems coming.
Even if there are a lot of great applications for these things (I think there are), it’s not clear they’re going to be truly transformative?
For me, it is already pretty transformative. I use ChatGPT at least as much as the regular internet - probably more.
Not just as a hobby/toy?
> "Computer, locate Commander Riker's repo"
> "Commander Riker's repo is not currently on the ship"
For example. I could have figured it out elseways, but would have taken >20x the time.
1) I wanted to make a github action to auto build, run tests and publish a package. I don't do a lot of this type of work so quite unfamiliar with it. ChatGPT gave me a very good start with some example yaml which I was able to tweak and get started with. When something wasn't working with my tweaked yaml file, I asked it what the problem was and it was able to pin-point the issue right away. I don't use it at work for anything but it is very useful for side projects and things like that.
2) I had some example json text from a service for which I wanted to create some serialization classes. This is easy to do but it is tedious work. ChatGPT was able to write most of this for me. There are likely tools, websites, services that can create this, but the experience in ChatGPT is pretty great - copy and paste in some json text and ask it to create the code. If it isn't quite right, give it some additionally instructions. When 80% OK, manually tweak it.
3) I use it to get ideas for things or to provide answers about things. For example, one of the commenters in this thread discussed the "AI scaling hypothesis". Google would have had a good answer for this but if you happen to have any questions(after the initial response is provided) this is a completely new search. ChatGPT, though not always correct, feels a lot more calming to use. Google seems very "noisy" to me now. Google is still very necessary but it feels like only 20% of queries need to go there at this point.
Obviously generative usecases are awesome (though IMO will be practically fewer and further between than people seem to think), and summarization usecases are awesome.
Some graduate students I collaborate use it to help them revise and polish their writing (English is not their first language)
Eg https://www.cnn.com/2023/01/28/tech/chatgpt-real-estate/inde...
Also these types of use cases I think will not actually last very long. They’ll be adversarial’d out of value and then probably out of existence after that.
In a world where every listing is written by a bot, a reader’s first objective is to distill the human-written parts out of the unreliable narration and filler words of the bot.
1) Short scripts in languages im unfamiliar with
2) Generating short, specific, well defined functions
3) Curing the tyranny of the blank page effect re emails, explanations, etc
4) Rephrasing that long email last thing on a friday when English just wont click
Search engines run on AI, stock markets trade on AI, social networks distribute content by AI, your photos and content are sorted on your phone by AI, vaccines and medicines are produced by AI, precedent law is implemented increasingly via AI, our weapons are AI (drones) and so on and so on and so on.
And all this up there was BEFORE GPT and DALL-E.
No, no winter. We're in the singularity.
The way I’ve interpreted the scaling hypothesis is that we will see emergent intelligence with larger nets through automated training alone. If we want a model to learn images we throw a larger net at it with training data.
The way I’ve interpreted some of these newer techniques to multi modality is that they are stitched together models. If we want a model to learn images we decide this and teach one model images and then connect it to the core model. There’s not a lot of emergent behavior due to scaling in this scenario.
With that perception I don’t see how gpt4 says anything about the scaling hypothesis. However I am not in this field and would be grateful to learn more.
So the goal should be: we've created a "language module" (LLMs) and a "visual perception module" (computer vision), but we also need to add a "logic module", a "reasoning module", an "empathy module", etc, while continuing to improve each.
I just don't see how you could get an LLM, no matter how advanced, to recognize a car. Even if it can describe cars (wheels, windshield, doors) it doesn't know what any of those components look like. It's like that old joke about philosophers being unable to define a chair beyond "I'll know it when I see it".
these net can replicate at most the first step for now, even tuning by reingofrcement learning is more of a set up batch than an ongoing thing, and certainly not something they will be able to do in the context of a single problem but as part of a retraining
agi is still a fair bit away, I'm unsure if these super large architecture will ever get to replicate the second part of our brain, the flexibility while on the job, because of their intrinsic training mechanism.
followed by GPT95, 98, XP, etc
Are they composed of autonomous agents that exhibit: - diversity (in behavior) - high connectedness - interaction among agents - adaptation (agents changing their behavior based on the behavior of other agents)?
Because if current NNs don't fit that definition of complexity, they it seems unreasonable to expect them to be intelligent, no matter their scale. Intelligence will likely be an emergent property of the above four criteria.
If anyone knows of relevant books or papers, would love to hear it.
BTW: "AI winter" refers to a collapse in funding; but I think we're well past that point. Previous winters only happened because academia led AI research, and there were no commercial opportunities. Now, the economic benefits are far more tangible. And if funding doesn't dry up, there's no reason to think innovation will dry up either.
Once the continuous learning problem is solved, it'll be pretty trivial to just tell your compound model to "go learn everything you can about x" and have it autonomously do so.
If it were me, the first thing I'd ask it to do was learn how to make other AI, then give it some hardware and the directive to build the best thing it could and test it until it can build a model better than itself.
If science can neither have a better definition of intelligence, or analyze and understand the resulting black box of a neural net, I don't see how AI would improve or get nearer general intelligence.
I have much more curiosity about neurology and how psychology define intelligence, than how programmers are trying to build things that digest data.
Proper science requires understanding. ML is not about understanding, it's only "sophisticated statistical methods". It is still artificial, thus but genuine.
Interesting that they are broadening instead of deepening the text side of GPT - maybe they are running into problems on that side and are building the feature set with low hanging fruit on adjacent areas.
Being able to perform more advanced types of zero shot learning tasks would be comparable and further the accuracy on those tasks can be evaluated
With 16k and some other techniques, I’m guessing it could write a custom CMS database backed web application.
more literally and correctly, it’s the maximum number of tokens in the input and output, combined, where a token is 4/3 of a word
So we’re shifting from 5K words maximum to 40K (per sibling comment, who pointed out 32K context leaked as well)
ChatGPT struggles with intuition about the world, if it has image/video knowledge than it can potentially reason about what happens when you flip a plate upside down.
There is an interesting “takeoff” point where models are able to generate enough “interesting” data to train themselves. Asking chatGPT for ml training data on various tasks makes me think we are close to that inflection point.
Dalle/GPT: that's not right, there should be 5 fingers on this clock.
It turns out that the image models already understand the low-level details of images well, and it's high-level reasoning/composition that is screwing their samples up - which is due to the poor text embedding they receive as a summary. (You can't generate a good image of 'a blue ball left of a red ball' if your text embedding is confused what 'left' is, no matter how high quality you are able to generate 'blue ball' or 'red ball' individually.) So, I would not be surprised if hands are improved by simply using a much better text model. (And why not plug in GPT-3 or PaLM for even better results...)
That said, hands are also like text-inside-images in being small, highly variable, and intrinsically difficult. And we know that the larger models solve text rendering simply by scaling up an OOM or 2 past SD. (You can see this from the text samples in the papers, and also more recently from Stability's Deep Floyd samples they've been teasing on Twitter.) So either way, scaling will probably solve hands soon.
One of their new models, PaLM-E, does show some positive transfer, so it's definitely a potential path for increased capabilities.
Also, amazingly, transfer learning works better than anyone could have hoped! As an analogy, imagine that you could show that football players learned musical instruments 50% quicker than non-athletes. That would be amazing!
But LLMs can write music and training them on music makes them better writers. It's almost like intelligence & skills are completely general!
I know the entire AI industry on the wrong track, but there's nothing that can be done about it. Sometimes you just have to wait for the crash and burn, and then things can change. It's like being at the ocean and trying to stop a wave from crashing after it has already formed, you just have to let it resolve itself of its own volition. As Einstein said, stupidity is infinite.
Previously, I heard that Bing Chat was based on GPT4. Is this true or not? Is it publicly known, or not? Certainly the model is far more powerful than what is in ChatGPT.
If Bing is GPT4, did they remove its multimodal capabilities?
If Bing is not GPT4, what is it? Does there exist a GPT3.5-mega with vastly improved reasoning capabilities and larger context windows?
I find the current moment in tech quite confusing.
chatgpt is what they call got3.5-turbo-0301 (current month checkpoint)
bingGPT is a allegedly finetuned version of gpt3.5 turbo that is better with search results.... microsoft internally calls it prometheus. You may also hear it referred to as Sydney.
i suspect GPT4 is just a bad translation
Then why does the chatgpt endpoint show as text-davinci-002-render-sha? It's very confusing.
> ChatGPT is powered by gpt-3.5-turbo, OpenAI’s most advanced language model.
Watching the Bing Chat escapades is both hilarious and mildly unsettling. It reminds me _so strongly_ of the sarcastic evil AI that runs the levels in the _Portal_ games.
Isn't GLaDOS a brain scan, not an AI?
There's even people who thinks brain scans/emulations will be how AGI is created, which I personally find very unlikely.
That was at least 2 weeks ago....
They've taken steps to limit it which have been mostly effective. There are some jailbreaks but they are getting harder.
It's a beta product. Give it 6 months and then judge it.
It's just a version number, and there aren't really meaningful reasons why something is a 0.5 increment vs a 1.0 release increment, other than "there are lots of changes" where lots ~= something.
> We train three model sizes (1.3B, 6B, and 175B parameters)[1]
It does use the same model architecture as GPT-3 though so I guess there is a bit of reasoning there maybe.