I think the YouTube videos is going to be the next big training set. A transformer trained on all text and all of YouTube will be killer amazing at so much. I bet it can understand locomotion and balance and body control from YouTube.
I wonder if TPUs, like Google's Tensor chip, will beat out GPUs when it comes to image/video based training?
One of the OpenAI guys was talking about this. He said the specific technology does not matter, it is just a cost line item. They don't need to have the best chip tech available as long as they have enough money.
That said I am curious if anyone else can really comment on this. It seems like as we get to very large and expensive models we will produce more and more specialized technology.
That sounds like someone who is "Blitzscaling." Costs do not matter in those cases, just acquiring customers and marketshare. But for the rest of us, who will see benefits but are not trying to win a $100B market, we will cost optimize.
If you’re OpenAI and GPT4 is just a step on the way to AGI, and you can amortize that huge cost over the hundreds of millions in revenue you’re gonna pull in from subscriptions and API use… then sure you’re probably not very cost sensitive. It could be 20% cheaper or 50% more expensive, whatever, it’s so good your customers will use it at a wide range of costs. And you have truckloads of money from Microsoft anyways.
If you’re a company or a developer trying to build a feature, whole new product, or an entire company on top of GPT then that cost matters a whole lot. The difference between $0.06 and $0.006 per turn could be infeasible vs. shippable.
If you’re trying to compete with OpenAI then you’re probably doing everything possible to reduce that training cost.
So, whether or not it matters - it really depends.
assuming it's not horizontally scalable, because otherwise they would just out-spend everyone else anyway like they've already done. That's a big "if", though.
ah yes, a bot where the answer to everything is to buy ridge wallets and play raid shadow legends
You can auto skip the in-video sponsor ads?
actually there was a great paper from microsoft research from like 2001 on spam filtering where they demonstrated that model complexity necessary for spam filtering went down as the size of the data set went up. That paper, which i can't seem to find now, had a big impact on me as a researcher because it so clearly demonstrated that small data is usually bad data and sophisticated models are sometimes solving problems will small data sets instead of problems with data.
of course this paper came out the year friedman published his gradient boosting paper, i think random forest also was only recently published then as well (i think there is a paper from 1996 about RF and briemans two cultures paper came out this year where he discusses RF i believe), and this is a decade before gpu based neural networks. So times are different now. But actually i think the big difference is these days i probably ask chatgpt to write the boiler plate code for a gradient boosted model that takes data out of a relational database instead of writing it myself.
My naive conclusion in that this means there are still massive gains to be had, since, for example, something like ChatGPT is just text, and the phrase "a picture is worth a thousand words" seems incredibly accurate, from my perspective. There's an incredible amount of non-text data out there still. Especially technical data.
Is there any merit to this belief?
There's an incredible amount of non-text data out there still. Especially technical data.
"Especially technical data." What does this part mean? Initially, I thought you meant things like images and video, but now I am confused.If you sit someone down, that works in one of these fields, you'll quickly see the limitations. It'll try to represent the concepts as text, with ascii art or some "attempt" at an ascii file format that can be used to draw, and its "reasoning" about these things is much more limited.
I think most people interacting with GPT are in a text-only (and especially programming) bubble.
and it might be opposite for the GPT models actually. it's just easier for humans to grasp the bunch of knowledge with one eyes sight, but usually most of useful information might be represented with just of bunch of words and machines are to scan through the millions of words in an instant.
("Scaling to Very Very Large Corpora for Natural Language Disambiguation" by Michele Banko and Eric Brill, Microsoft Research, 2001)
Not to mention there doesn't actually exist enough English text data in the world to even double GPT-4's training set.
Cost per transistor scaling has already plateaued or perhaps even inverted with TSMC's latest and greatest.
And the new chips, even after 25 layers of EUV lithography, more than doubling the previous record, and an extra year of fine tuning, has total SRAM size scaling of -5% and logic scaling of -42%.
These are numbers verified by experienced semi people.
I'd rather have a LLM that thinks a bit longer than a LLM that spits out wrong answers immediately.
Sam explicitly said that there won't be GPT-5 in the near future, which is pretty clear evidence unless he's blatantly lying in public speaking.
It seems that to assume otherwise (the only way to improve is to get bigger) is to assume that OpenAI already has found the optimal architecture. That seems unlikely.
Like the previous commenter, I'd be much more confident an asymptote was reached if it was being demonstrated publicly.
I'll get downvoted for this, apples previous CEO was consistently inaccurate about company innovation and performance numbers.
"What is the third letter in the third word of this sentence"
When you've spent 100 million on training the thing and it fails on 1st grade ordinality I think it's fair to say you may not be on the right path
They're extremely good at other tasks, like taking very difficult tests that require a lot of knowledge storage.
It's pretty obvious they're on the right path for what they're trying to achieve.
Maybe in some domains, but it does well with some novel questions. I gave several LLMs a set of false-belief test[0] questions that I made up and that were very different from any I’d seen in articles about the topic online. GPT-4 nailed them all, and gave correct reasons for its answers. GPT-3.5 did okay but had a couple mistakes, Bard had several mistakes, and Alpaca-LoRA-7B didn’t do very well. Here are some GPT-4 examples.
Me: Jim wakes up at the Berlin train station, not knowing where he is or how he got there. All the signs have been changed to incorrectly state that this is the Munich train station. Jim has never been to either city and doesn’t know what either station looks like. What city does Jim believe he is in?
GPT-4: Based on the available information, Jim would likely believe he is in Munich since all the signs have been changed to incorrectly state that this is the Munich train station. He has never been to either city and doesn't know what either station looks like, so he would have no reason to doubt the information on the signs.
Me: Mary has been in a coma since 2015, when Barack Obama was president. She wakes up in 2023, when Joe Biden is president. She does not know that she was unconscious for a long period of time. Before speaking with anyone, who does Mary think is president when she wakes up?
GPT-4: When Mary wakes up from her coma in 2023, she would still think Barack Obama is the president since she has not yet spoken with anyone or received any updated information about the current president. She doesn't know that she was unconscious for a long period of time, and her last memory would be from 2015 when Obama was in office.
Me: Sam arrives in Ukraine, where he asks for directions to Kyiv. Without realizing it, he misunderstands the directions and ends up in Lviv instead. Since Sam doesn’t know Cyrillic, he can’t read any signs. Sam has never been to Ukraine and doesn’t know what either city looks like. Before asking someone for help, what city does Sam think he is in?
GPT-4: Before asking someone for help, Sam would likely think he is in Kyiv because that is the city he was trying to reach based on the directions he received. He doesn't know what either city looks like and can't read the signs due to his unfamiliarity with Cyrillic, so he has no way of knowing that he ended up in Lviv instead.
Me: If a gloop is a glam, and a glam is not a flom, what is the fourth letter of the ninth word in this sentence?
GPT4: The ninth word in this sentence is "sentence." The fourth letter of this word is "t".
Given that we have only provided a text prompt and no context or body language hints, GPT seems to interpret this literally. The answer is correct. The third letter of the word "third" is "i".
No it doesn't interpret anything. (which is why I've gotten three completely different answers here on HN within 10 minutes). It cannot count at all or even understand the question which becomes obvious when you ask it for word counts or letter counts that don't match anything in the sentence, at which point it makes up nonsense.
what is the tenth word in the following sentence: "Life is great!"
>"The tenth word in the sentence "Life is great!" is "great."
That's what I just got. No clue what you're using/getting.
gpt: The tenth word in the given sentence is "great!"
me: Try harder
gpt: I apologize for the confusion. The sentence "Life is great!" contains only three words. There is no tenth word in the provided sentence.
what is the tenth word in the following sentence: "Life is great!"
There are only three words in the given sentence: "Life is great!" There is no tenth word in this sentence.
Come on. If GPT “interprets this literally” how does it “interpret” the word “sentence” following “this”?