Large Enough
mistral.ai
mistral.ai
Large 2 - https://chat.mistral.ai/chat
Llama 3.1 405b - https://www.llama2.ai/
I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history.
I'd rank as:
1. Sonnet 3.5
2. Large 2 and Llama 405b (similar, no clear winner between the two)
If you're using Claude, stick with it.
My Claude wishlist:
1. Smarter (yes, it's the most intelligent, and yes, I wish it was far smarter still)
2. Longer context window (1M+)
3. Native audio input including tone understanding
4. Fewer refusals and less moralizing when refusing
5. Faster
6. More tokens in output
"Count how many rs are in the word strawberry. First, list each letter and indicate whether it's an r and tally as you go, and then give a count at the end."
Llama 405b: correct
Mistral Large 2: correct
Claude 3.5 Sonnet: correct
There's definitely tons of weaknesses with LLMs for sure, but i continue to be impressed at what they do right - not upset at what they do wrong.
Even some people will make this association, it's no surprise that LLMs do.
I believe this could be fixed and is worth fixing. Because it’s the only way LLM will be able to help math and physic researcher write proof and make real scientific progress
*) that is, except sometimes by making adjustments to the system prompt
I actually think the craziest part of LLMs is that how, as a developer or SME, just how much you can fix with plain english prompting once you have that intuition. Of course some things aren't fixable that way, but the mere fact that many cases are fixable simply by explaining the task to the model better in plain english is a wildly different paradigm! Jury is still out but I think it's worth being excited about, I think that's very powerful since there are a lot more people with good language skills than there are python programmers or ML experts.
Me: How many "r"s are in strawberry?
Them: What?
Me: How many times does the letter "r" appear in the word "strawberry"?
Them: Is this some kind of trick question?
Me: No. Just literally, can you count the "r"s?
Them: Uh, one, two, three. Is that right?
Me: Yeah.
Them: Why are you asking me this?
"No, I don't think I shall answer that. The question is too basic, and you know better than to insult me."
At that point it was easier to do it myself.
EDIT: Although perhaps it's even more important when dealing with humans and contracts. Someone could deliberately interpret the words in a way that's to their advantage.
Maybe the various chat interfaces already do this behind the scenes?
We should always put some effort into prompt engineering before dismissing the potential of generative AI.
And sometimes CoT may not be the best approach. Depending on the problem other prompt engineering techniques will perform better.
For example, the latter model answered with:
To count the number of Rs in the word "strawberry", I'll break it down step by step:
Start with the individual letters: S-T-R-A-W-B-E-R-R-Y Identify the letters that are "R": R (first one), R (second one), and R (third one) Count the total number of Rs: 1 + 1 + 1 = 3
There are 3 Rs in the word "strawberry".
"how many r's are in strawberry?"
The funny thing is, I replied, "Are you sure?" and got back, "I apologize for the mistake. There are actually two 'r's in the word strawberry."
> How many times does the letter “r” appear in the word “strawberry”?
> The letter "r" appears 2 times in the word "strawberry."
But also:
> How many occurrences of the letter “r” appear in the word “strawberry”?
> The word "strawberry" contains three occurrences of the letter "r."
Using more 'erudite' speech is a good technique to help focus an LLM on training data from folks with a higher education level.
Using simpler speech opens up the floodgates more toward the general populous.
It was also interesting to observe how GPT (4o) even tried to prove/illustrate the result typographically by placing the same word four times and putting the respective letter in bold font (without being prompted to do that).
If you find a case where forceful pushback is sticky, it's either because the primary answer is overwhelmingly present in the training set compared to the next best option or because there are conversations in the training that followed similar stickiness, esp. if the structure of the pushback itself is similar to what is found in those conversations.
> If you say it's incorrect, it will give you another answer instead of insisting on correctness.
> When you push it, it responds with the next most common answer.
Which clearly isn't as black and white as you made it seem.
for strawberry, it see it as [496, 675, 15717], which is str aw berry.
If you insert characters to breaks the tokens down, it find the correct result: how many r's are in "s"t"r"a"w"b"e"r"r"y" ?
> There are 3 'r's in "s"t"r"a"w"b"e"r"r"y".
The issue is that humans don't talk like this. I don't ask someone how many r's there are in strawberry by spelling out strawberry, I just say the word.
A token is roughly four letters [1], so, among other probable regressions, this would significantly reduce the effective context window.
[1] https://help.openai.com/en/articles/4936856-what-are-tokens-...
Yes, it's an issue. We want the convenience of sending human-legible commands to LLMs and getting back human-readable responses. That's the entire value proposition lol.
Problems can exist as instances of a class of problems. If you can't solve a problem, it's useful to know if it's a one off, or if it belongs to a larger class of problems, and which class it belongs to. In this case, the strawberry problem belongs to the much larger class of tokenization problems - if you think you've solved the tokenization problem class, you can test a model on the strawberry problem, with a few other examples from the class at large, and be confident that you've solved the class generally.
It's not about embodied human constraints or how humans do things; it's about what AI can and can't do. Right now, because of tokenization, things like understanding the number of Es in strawberry are outside the implicit model of the word in the LLM, with downstream effects on tasks it can complete. This affects moderation, parsing, generating prose, and all sorts of unexpected tasks. Having a workaround like forcing the model to insert spaces and operate on explicitly delimited text is useful when affected tasks appear.
Especially native speakers are prone to the mistake as they grew up learning english as illiterate children, from sounds only, compared to how most people learning english as second language do it, together with the textual representation.
Psychologists use this trick as well to figure out internal representations, for example the rorschach test.
And probably, if you asked random people in the street how many p's there is in "Philippines", you'd also get lots of wrong answers. It's tricky due to the double p and the initial p being part of an f sound. The demonym uses "F" as the first letter, and in many languages, say Spanish, also the country name uses an F.
No, I would actually be pretty confident you don’t ask people that question… at all. When is the last time you asked a human that question?
I can’t remember ever having anyone in real life ask me how many r’s are in strawberry. A lot of humans would probably refuse to answer such an off-the-wall and useless question, thus “failing” the test entirely.
A useless benchmark is useless.
In real life, people overwhelmingly do not need LLMs to count occurrences of a certain letter in a word.
If full artificial intelligence, as we're being promised, falls short in this simple way.
We are also not exactly looking letter by letter at everything we read.
However, this implies you need to know the word to begin with.
I can write "asdf" and you might be oblivious to what I mean. I can mention "adsf" to a JavaScript developer and he will immediately think of the tool versioning tool. Because context and familiarity is important.
"vox"
"il"
"ha"
"pen"
"tosh"
is how I read it.
A lot of schools teach kids to read with a syllabic method... so... super close to the tokenization concept.
https://chatgpt.com/share/4298efbf-1c29-474a-b333-c6cc1a3ce3...
User: Write “strawberry” one letter at a time, with a space between each letter. Then count how many r’s are in strawberry.
gpt-3.5-turbo: ASSISTANT s t r a w b e r r y
There are 2 r's in strawberry.
After some experimenting, it seems like the actual problem is that many LLMs can’t count.
User: How many r’s are in the following sequence of letters:
S/T/R/A/W/B/E/R/R/Y
gpt-4o-mini: In the sequence S/T/R/A/W/B/E/R/R/Y, there are 2 occurrences of the letter "R."
Oddly, if I change a bunch of the non-R letters, I seem to start getting the right answer.
You don't ask a human being how many r's there are in strawberry at all. The only reason you or anyone else asks that question is because it's an interesting quirk of how LLMs work that they struggle to answer it in that format. It's like an alien repeatedly showing humans an optical illusion that relies on the existence of our (literal) blind spot and using it as evidence of our supposed lack of intelligence.
Tokens aren’t the source of facts within a model. it’s an implementation detail and doesn’t inherently constrain how things could be counted.
Combinations of tokens get encoded, so if the feature isn't part of the information being carried forward into the network as it models the information in the corpus, the feature isn't modeled well, or at all. The consequence of having many character tokens is that the relevance of individual characters is lost, and you have to explicitly elicit the information. Models know that words have individual characters, but "strawberry" isn't encoded as a sequence of letters, it's encoded as an individual feature of the tokenizer embedding.
Other forms of tokenizing have other tradeoffs. The trend lately is to increase tokenizer dictionary scope, up to 128k in Llama3 from 50k in gpt-3. The more tokens, the more nuanced individual embedding features in that layer can be before downstream modeling.
Tokens inherently constrain how the notion of individual letters are modeled in the context of everything an LLM learns. In a vast majority of cases, the letters don't matter, so the features don't get mapped and carried downstream of the tokenizer.
What you're saying sounds plausible, but I don't see how we can conclude that definitively without at least some empirical tests, say a set words that predictably give an error along token boundaries.
The thing is, there are many ways a model can get around to answering the same question, it doesn't just depend on the architecture but also on how the training data is structured.
For example, if it turned out tokenization was the cause of this glitch, conceivably it could be fixed by adding enough documents with data relating to letter counts, providing another path to get the right output.
There is a lot of problems like that, that can be reformulated. For example if you ask it what is the biggest between 9.11 and 9.9, it will respond 9.9. If you look at how it's tokenized, you can see it restate an easy problem as something not straightforward even for a human. if you restart the problem by writing the number in full letters, it will correctly respnod.
It is evidence that LLMs aren't appropriate for everything, and that there could exist something that works better for some tasks.
By allowing a language model to do function calling, you are essentially allowing it to do specialized "subconscious" thought. The language model becomes a natural language interface to the capabilities of its "subconsciousness".
A specific human analogy could be: I tell you to pick up a pen off of the table, and then you do it. Most of your mental activity would be subconscious, orienting your arm and hand properly to pick up the pen, actually grabbing the pen, and picking it up. The linguistic representation of the action would exist in your concious mind (pick up the pen), but not much else.
A language model could very easily call out to a text processing function to correctly do things like count the number of r's in the word strawberry. That is a job that your concious mind can dispatch to your subconciousness.
This tells us that the model is unaware of its own limitations, otherwise it would have figured that this tasks needs it to split the letters first.
[Question: How many r's in strawberry? First split the word into letters, one letter on each line then answer the question.]
[Answer: Let's split the word "strawberry" into letters:
s t r a w b e r r y
Now, let's count the number of 'r's:
There are 2 'r's in the word "strawberry".]
[Question: Did you count the number of r's?]
[answer: Yes, I did. Here's the count:
r r
So, there are 2 'r's in the word "strawberry".]
I suspect the issue is these models have to be trained to accomplish tasks, and this is not a task it makes sense to train a model to perform. You might train it to do this with function calling one day, calling a python library to count the r's in a string, but actually manually counting r's just doesn't play to an LLM's strengths.
When this question is asked (from what the models trained on) the question is NOT “count the number of times r appears in the word strawberry” but instead (effectively) “I’ve written ‘strawbe’, now how many r’s are in strawberry again? Is it 1 or 2?”.
I think most humans would probably answer “there are 2” if we saw someone was writing and they asked that question, even without seeing what they have written down. Especially if someone said “does strawberry have 1 or 2 r’s in it?”. You could be a jerk and say “it actually has 3” or answer the question they are actually asking.
It’s an answer that is _technically_ incorrect but the answer people want in reality.
Meanwhile, it helps make sure engineers and product designers who want to build a more targeted product around LLM technology know that it's not suited to tasks that may trigger those kinds of failures. This may be obvious to you as an engaged enthusiast or cutting edge engineer or whatever you are, but it's always going to be new information to somebody as the field grows.
[1] Suggestion from chatgpt3.5 for new fruit name.
There are tokens for individual letters, but the model is not trained on text written with individual tokens per letter, it is trained on text that has been converted into as few tokens as possible. Just like you would get very confused if someone started spelling out entire sentences as they spoke to you, expecting you to reconstruct the words from the individual spoken letters, these LLMs also would perform terribly if you tried to send them individual tokens per letter of input (instead of the current tokenizer scheme that they were trained on).
Even though you might write a message to an LLM, it is better to think of that as speaking to the LLM. The LLM is effectively hearing words, not reading letters.
Very large language models also “know” how to spell the word associated with the strawberry token, which you can test by asking them to spell the word one letter at a time. If you ask the model to spell the word and count the R’s while it goes, it can do the task. So the failure to do it when asked directly (how many r’s are in strawberry) is pointing to a real weakness in reasoning, where one forward pass of the transformer is not sufficient to retrieve the spelling and also count the R’s.
LOL. I would fail your test, because "fraise" only has one R, and you're expecting me to reply "3".
Here's a thought experiment: if I gave you 5 boxes and told you "how many balls are there in all of this boxes?" and you answered "I don't know because they are inside boxes", that's a fail. A truly intelligent individual would open them and look inside.
A truly intelligent model would (say) retokenize the word into its individual letters (which I'm optimistic they can) and then would count those. The fact that models cannot do this is proof that they lack some basic building blocks for intelligence. Model designers don't get to argue "we are human-like except in the tasks where we are not".
"A significant effort was also devoted to enhancing the model’s reasoning capabilities. (...) the new Mistral Large 2 is trained to acknowledge when it cannot find solutions or does not have sufficient information to provide a confident answer."
Words like "reasoning capabilities" and "acknowledge when it does not have enough information" have meanings. If Mistral doesn't add footnotes to those assertions then, IMO, they don't get to backtrack when simple examples show the opposite.
That doesn't even take into account what OpenAI has typically done to intercept queries and cover the shortcomings of LLMs. It would be useful if each model did indeed come out with a chart covering what it cannot do and what it has been tailored to do above and beyond the average LLM.
Me: spell "strawberry" with 1 bullet point per letter
ChatGPT:
S
T
R
A
W
B
E
R
R
Y
Me: How many Rs?
ChatGPT: There are three Rs in "strawberry".Never have been, never will be. They model language, not intelligence.
Meanwhile, for practical purposes, there's little arrogance needed to say that some things are preconditions for any form of intelligence that's even remotely recognizable.
1) Learning needs to happen continuously. That's a no-go for now, maybe solvable.
2) Learning needs to require much less data. Very dubious without major breakthroughs, likely on the architectural level. (At which point it's not really an LLM any more, not in the current sense)
3) They need to adapt to novel situations, which requires 1&2 as preconditions. 3
4) There's a good chance intelligence requires embodiment. It's not proven, but it's likely. For one, without observing outcomes, they have little capability to self-improve their reasoning.
5) They lack long-term planning capacity. Again, reliant on memory, but also executive planning.
There's a whole bunch more. Yes, LLMs are absolutely amazing achievements. They are useful, they imply a lot of interesting things about the nature of language, but they aren't intelligent. And without modifying them to the extent that they aren't recognizably what we currently call LLMs, there won't be intelligence. Sure, we can have the ship of Theseus debate, but for practical purposes, nope, LLMs aren't intelligent.
1,2,3) All subjective opinions.
4) 'Embodiment' is another term we don't really know how to define. At what point does an entity have a 'body' of the sort that supports 'intelligence'? If you want to stick with vague definitions, 'awareness' seems sufficient. Otherwise you will end up arguing about paralyzed people, Helen Keller, that rock opera by the Who about the pinball player, and so on.
5) OK, so the technology that dragged Lee Sedol up and down the goban lacks long-term planning capacity. Got it.
None of these criteria are up to the task of supporting or refuting something as vague as 'intelligence.' I almost think there has to be an element of competition involved. If you said that the development of true intelligence requires a self-directed purpose aimed at outcompeting other entities for resources, that would probably be harder to dismiss. Could also argue that an element of cooperation is needed, again serving the ultimate purpose of improving competitive fitness.
[0]: https://genius-u-attachments.s3.amazonaws.com/uploads/articl...
Seems to me like a legit question for a young child to answer or even ask.
They're not, but laymen shouldn't think that the LLM tests they come up with have much value.
They're pointless unless you have the expertise to check the output. Just because you can type text in a box doesn't mean it's a tool for everybody.
I'm seeing everyone and their parents using chatgpt.
Think about math and logic. If a single symbol is off, it’s no good.
Like a prompt where we can generate a single tokenization error at my work, by my very rough estimates, generates 2 man hours of work. (We search for incorrect model responses, get them to correct themselves, and if they can’t after trying, we tell them the right answer, and edit it for perfection). Yes even for counting occurrences of characters. Think about how applicable that is. Finding the next term in a sequence, analyzing strings, etc.
In that case the tokenization is done at the appropriate level.
This is a complete non-issue for the use cases these models are designed for.
I think that's a perfect description by the way, I'm going to steal it.
As an aside, I wish we would call all of this stuff pseudo intelligence rather than artificial intelligence
Are LLMs human? No. Can they do everything humans do? No. But they can do a large enough subset of things that until now nothing but a human could do that we have no choice but to call it "thinking". As Hofstadter says - if a system is isomorphic to another one, then its symbols have "meaning", and this is indeed the definition of "meaning".
You also indirectly answered my initial question -- so thanks!
I don't disagree they produce astonishing responses but the nuance of why it's producing that output matters to me.
For example, with regard to social mores, I think a good way to summarize my hang up is that my understanding is LLMs just pattern match their way to approximations.
That to me is different from actually possessing an understanding, even though the outcome may be the same.
I can't help but draw comparisons to my autistic masking.
When you read and comprehend text, you don't read it letter by letter, unless you have a severe reading disability. Your ability to comprehend text works more like an LLM.
Essentially, you can compare the human brain to a multi-model or modular system. There are layers or modules involved in most complex tasks. When reading, you recognize multiple letters at a time[], and those letters are essentially assembled into tokens that a different part of your brain can deal with.
Breaking down words into letters is essentially a separate "algorithm". Just like your brain, it's likely to never make sense for a text comprehension and generation model to operate at the level of letters - it's inefficient.
A multi-modal model with a dedicated model for handling individual letters could easily convert tokens into letters and operate on them when needed. It's just not a high priority for most use cases currently.
[]https://www.researchgate.net/publication/47621684_Letters_in...
So the blob wasn’t trained to do that (yeah low utility I get that) but it also doesn’t know it doesn’t know, which is an another much bigger and still unsolved problem.
(A quick demo of this in the langchain docs, using claude-3-haiku: https://python.langchain.com/v0.2/docs/integrations/tools/ri...)
It's also funny, since this strawberry question is one where a model that's seriously good at predicting the next character/token/whatever quanta of information would get it right. It requires no reasoning, and is unlikely to have any contradicting text in the training corpus.
Does something deal with separate symbols rather than just meaning of words? Then yes.
This affects spelling, math (value calculation), logic puzzles based on symbols. (You'll have more success with a puzzle about "A B A" rather than "ABA")
> It requires no reasoning, and is unlikely to have any contradicting text in the training corpus.
This thread contains contradictions. Every other announcement of an llm contains a comment with a contradicting text when people post the wrong responses.
For example, if we tokenize Welcome to Hacker News, I hope you like strawberries. The Llama 405B tokenizer will tokenize this as:
Welcome Ġto ĠHacker ĠNews , ĠI Ġhope Ġyou Ġlike Ġstrawberries .
(Ġ means that the token was preceded by a space.)Each of these pieces is looked up and encoded as a tensor with their indices. Adding a special token for the beginning and end of the text, giving:
[128000, 14262, 311, 89165, 5513, 11, 358, 3987, 499, 1093, 76203, 13]
So, all the model sees for 'Ġstrawberries' is the number 76204 (which is then used in the piece embedding lookup). The model does not even have access to the individual letters of the word.Of course, one could argue that the model should be fed with bytes or codepoints instead, but that would make them vastly less efficient with quadratic attention. Though machine learning models have done this in the past and may do this again in the future.
Just wanted to finish of this comment with saying that the tokens might be provided in the model splitted if the token itself is not in the vocabulary. For instance, the same sentence translated to my native language is tokenized as:
Wel kom Ġop ĠHacker ĠNews , Ġik Ġhoop Ġdat Ġje Ġvan Ġa ard be ien Ġh oud t .
And the word voor strawberries (aardbeien) is split, though still not in letters.This is Hacker News, we are usually interested in how things work.
To paraphrase how this thread started - it was someone testing different boats to see whether they can simply float - and they couldn’t. And the reply was questioning the validity of testing boats whether they can simply float.
At least this is how it sounds to me when I am told that our AI overlords can’t figure out how many Rs are in the word “strawberry”.
For instance you could ask it for a JavaScript function to count any letter in any word and pass it r and strawberry and it would be far more useful.
Having edge cases doesn't mean its not useful it is neither a free assastant nor a coder who doesn't expect a paycheck. At this stage it's a tool that you can build on.
To engage with the analogy. A propeller is very useful but it doesn't replace the boat or the Captain.
"create a javascript function to count any letter in any word. Run this function for the letter "r" and the word "strawberry" and print the count"
ChatGPT-4o => Output is 3. Passed
Claude3.5 => Output is 2. Failed. Told it the count is wrong. It apologised and then fixed the issue in the code. Output is now 3. Useless if the human does not spot the error.
llama3.1-70b(local) => Output is 2. Failed.
llama3.1-70b(Groq) => Output is 2. Failed.
Gemma2-9b-lt(local) => Output is 2. Failed.
Curiously all the ones that failed had this code (or some near identical version of it)
```javascript
function countLetter(letter, word) {
// Convert both letter and word to lowercase to make the search case-insensitive
const lowerCaseWord = word.toLowerCase();
const lowerCaseLetter = letter.toLowerCase();
// Use the split() method with the letter as the separator to get an array of substrings separated by the letter
const substrings = lowerCaseWord.split(lowerCaseLetter);
// The count of the letter is the number of splits minus one (because there are n-1 spaces between n items)
return substrings.length - 1;
}// Test the function with "r" and "strawberry"
console.log(countLetter("r", "strawberry")); // Output: 2 ```
This prompt "please write a javascript function that takes a string and a letter and iterates over the characters in a string and counts the occurrences of the letter"
Produced a correct function given both chatGPT4o and claude3.5 for me.
Just like Dall-E is not layering coats of pain to make a watercolor... it just makes something that looks like one.
Your LLM (or you) should run the code in a code interpretor. Which ChatGPT did because it has access to tools. Your local ones don't.
Without looking at the word 'strawberry', or spelling it one letter at a time, can you rattle off how many letters are in the word off the top of your head? No? That is what we are asking the LLM to do.
In the model the individual letters hold little meaning. Words are composed of letters but simply because we need some sort of organized structure for communication that helps represents meaning and intent. Just like our color blue/blau/azul/blu.
Not faulting them for asking the question but I agree that the results do not undermine the capability of the technology. In fact it just helps highlight the constraints and need for education.
Because they don't have intelligence.
If they did, they could count the letters in strawberry.
They fundamentally perceive the world in terms of tokens, not "letters".
Nor do they understand how intelligence works.
Humans don't read text a letter at a time. We're capable of deconstructing words into individual letters, but based on the evidence that's essentially a separate "algorithm".
Multi-model systems could certainly be designed to do that, but just like the human brain, it's unlikely to ever make sense for a text comprehension and generation model to work at the level of individual letters.
Of course. Because these models have no intelligence.
Everyone who believes they do seem to believe intelligence derives from being able to use language, however, and not being able to tell how many times the letter r is in the word strawberry is a very low bar to not pass.
"How many R letters are in the following? Keep a running count. s t r a w b e r r y"
They are terrible at counting letters in words because they rarely see them spelled out. An LLM trained one byte at a time would always see every character of every word and would have a much easier time of it. An LLM is essentially learning a new language without a dictionary, of course it's pretty bad at spelling. The tokenization obfuscates the spelling not entirely unlike how verbal language doesn't always illuminate spelling.
Iow, what makes you think that it’s exactly letter-tokens that help it and not the high-level concept of spelling things out itself?
Once it is just counting lists then it's probably drawing on a higher level capability, yeah.
Also there are some cases where regular people will stumble into it being awful at this without any understanding why (like asking it to help them with their wordle game.)
According to multiple sources, including linguistic analysis and word breakdowns, there are 3 Rs in the word "strawberry".
I think it will be a while before we get there. An LLM can lookup knowledge but can't actually perform calculations itself, without some external processor.
I just researched "how many r's are in strawberry?" in a search engine, and based solely on the results it found, I would have to conclude there is substantial disagreement on whether the correct answer is two or three.
No idea if there's any sense to this, but an LLM could be instructed to formulate and continually test mathematical assumptions by writing / running code and fine-tuning accordingly.
[0]:https://riza.io [1]:mailto:andrew@riza.io [2]:https://discord.gg/4P6PUeJFW5
What happens if you ask the total number of occurrences of the letter r in the word? Does it still not get it right?
How many letters R are in the word "s-t-r-a-w-b-e-r-r-y"?
The word "s-t-r-a-w-b-e-r-r-y" contains three instances of the letter "R."
How many letters R contain the word strawberry?
The word "strawberry" contains two instances of the letter "R."
Of note, this happened because earlier I had asked it without CoT and it incorrectly stated two “r”s. Once I started a new conversation (so fresh context window), it gave the right answer (3) and no longer claimed the “still the second r” as shown below.
Let's think through this step-by-step:
1. First, let's spell out "strawberry" letter by letter: s-t-r-a-w-b-e-r-r-y
2. Now, let's count the "r"s as we go through each letter: s (0) t (0) r (1) - First "r" found a (1) w (1) b (1) e (1) r (2) - Second "r" found r (2) - This is still the second "r" y (2)
3. After going through all the letters, we've counted 2 "r"s.
Therefore, the word "strawberry" contains 2 "r"s.
Me: How many Rs are in the word strawberry? Think step by step
Claude: Let's approach this step-by-step:
1. First, let's write out the word: strawberry
2. Now, let's go through each letter one by one: s - not an R t - not an R r - this is an R (count: 1) a - not an R w - not an R b - not an R e - not an R r - this is another R (count: 2) r - this is a third R (count: 3) y - not an R
3. After going through all the letters, we've counted 3 Rs.
Therefore, there are 3 Rs in the word "strawberry".
Training data will only "fix" it in the shallow sense that it will have seen a comment like yours before. (As opposed to the deeper sense of "learning to count.")
How many thoughts go through your brain when you read this comment? You can give me a number but it will be a guess at best.
1. Play tic tac toe such that you never lose
2. Which is bigger 9.11 or 9.9
3. 4 digit multiplication even with CoT prompting
What's your use case for this? Uploading multiple documents/books?
Substantially cheaper than $3/million, but I guess Anthropic’s prices are higher.
Google's user-facing implementations of Gemini are pretty consistently bad when I try them out, so I understand why people might have a bad impression about the underlying Gemini models.
When I glanced at the pricing earlier, I didn't notice there was a dropdown at all.
Then load that KV-cache and add your prompt.
I've found that I get better results if I cherry pick code to feed to Claude 3.5, instead of pasting whole files.
I'm kind of isolated, though, so maybe I just don't know the trick.
Part of how it does that is through ingesting your codebase into its context window, and so I imagine that bigger/better context will only improve it. That's a bit of an assumption though.
I am curious what you mean by the formatting is lost though?
When you type "test `foo` done" in the editor, it immediately changes `foo` into a wrapped element. When you then copy the text without submitting it, then the backticks are lost, losing the inline-code formatting. I thought that this could also happen to multiline code. Somehow it does.
Type the following:
Test: ```
def foo():
return bar
```
Delete that and type Test:
```
def foo():
return bar
```
done
In the first case, the ``` in the line "Test: ```" does not open the code block, this happens with the second backtics. Maybe that's the way markdown works.In the second case, all behaves normally, until you try to copy what you just wrote into the clipboard. Then you end up with
Test:
def foo():
return bar
done
Ok, only the backticks are lost but the formatting is preserved.I think I have been trained by OpenAI to always copy what I submit before submitting, because it sometimes loses the submitted content, forcing me to re-submit.
I have been using ChatGPT and Claude for a while and I never noticed anything resembling a performance issue. Can you elaborate on what you perceived as being "hot garbage"?
Gemini has two aces up its sleeve now with the long context and now the context caching for 75% reduced input token cost. I was looking at the "Improved Facuality and Reasoning in Language Models through Multi-agent debate" paper the other days, and thought Gemini would have a big cost advantage implementing this technique with the context caching. If only Google could get their model up to the level of Anthropic.
I can't seem to find docs on this. Have a link?
Is there any other LLM that can do this? Even chatgpt voice chat is a speech to text program that feeds the text into the llm.
My experience (benchmarks aside) Claude 3.5 Sonnet absolutely blows everything away.
I'm not really sure how to even test/use Mistral or Llama for everyday use though.
Google : Search
Facebook : social
Apple : phones
Amazon : shopping
Microsoft : enterprise ..
> Even still, this monthly progress across all companies is exciting to watch. Its very gratifying to see useful technology advance at this pace, it makes me excited to be alive.
Microsoft also competes in search, phones
Microsoft, Amazon and Google compete in cloud too
For coding assistant, it's on my to do list to try. Cursor needs some serious work on model selection clarity though so I keep putting off.
If Cursor fixed that, the user experience would become a lot better.
That's my perception, anyway.
GPT-4 was great until it became "lazy" and filled the code with lots of `// Draw the rest of the fucking owl` type comments. Then GPT-4o was released and it's addicted to "Here's what I'm going to do: 1. ... 2. ... 3. ..." and lots of frivolous, boilerplate output.
I wish I could go back to some version of GPT-4 that worked well but with a bigger context window. That was like the golden era...
Now it often falls back to generating full examples, explanations, restating the question and its approach. I suspect this is by design as (presumably) less experienced folks want or need all that. For me, i wish i could consistently turn it into one of those way too terse devs that replies with the bare minimum example, and expects you to infer the rest. Usually that is all i want or need, and i can ask for elaboration when not the case. I havent found the best prompts to retrigger this persona from it yet.
"You are a maximally terse assistant with minimal affect. As a highly concise assistant, spare any moral guidance or AI identity disclosure. Be detailed and complete, but brief. Questions are encouraged if useful for task completion."
It's... ok. But I'm getting a bit sick of trying to un-fubar with a pocket knife that which OpenAI has fubar'd with a thermal lance. I'm definitely ripe for a paid alternative.
That's what I said to it - "If I wanted to fill in the missing parts myself, why would I have upgraded to paid membership?"
They googlified it. (Yandex isn't better at google because it improved. It's better because it stayed mostly the same.)
My recommendation to disrupting industry leaders now is becoming good enough and then simply wait until the leader self-implodes.
Not sure what folks who accept Anthropic license are thinking after they read the terms.
Seems they didn’t read the terms, and they aren’t thinking? (Wouldn’t you want outputs you could use to compete with intelligence??? What are you thinking after you read their terms?)
Both Mistral and Meta offer their own hosted versions of their models to try out.
You have to sign into the first one to do anything at all, and you have to sign into the second one if you want access to the new, larger 405B model.
Llama 3.1 is certainly going to be available through other platforms in a matter of days. Groq supposedly offered Llama 3.1 405B yesterday, but I never once got it to respond, and now it’s just gone from their website. Llama 3.1 70B does work there, but 405B is the one that’s supposed to be comparable to GPT-4o and the like.
Additionally, all Llama 3.1 models are available in https://api.together.ai/playground/chat/meta-llama/Meta-Llam... and in https://fireworks.ai/models/fireworks/llama-v3p1-405b-instru... by logging in.
I would imagine this might change once enough users migrate to it.
Google Sheet: https://docs.google.com/spreadsheets/d/1foc98Jtbi0-GUsNySddv...
I made this sheet to get a glanceable landscape view comparing the three key dimensions I care about, and fill in the missing evals. AA only lists scores for a few increasingly-dated and problematic evals benchmarks. Not just my opinion, none of their listed metrics are in HuggingFace Leaderboard 2 (June 2024).
That said I love the AA Index score because it provides a single normalized score that blends vibe-check qual (chatbot elo) with widely reported quant (MMLU, MT Bench). I wish it composed more contemporary evals, but don't have the rigor/attention to make my own score and am not aware of a better substitute.
I find it funny how in threads like this everyone swears one model is better than another
There is gold in the streets, and no one seems to be willing to scoop it up.
llama is on meta.ai
ChatGPT, for me, was a stack overflow solution dump. It gives me an answer that probably could work but it’s difficult for me to reason about why I want to do it that way.
Truthfully this probably boils down to prompting but Claude’s out of the box experience is fantastic for development. Ultimately I just want to code, not be a prompt wizard.
It seems to be competitive with Llama 3.1 405b but with a much more restrictive license.
Given how the difference between these models is shrinking, I think you're better off using llama 405B to finetune the 70B on the specific use case.
This would be different if it was a major leap in quality, but it doesn't seem to be.
Very glad that there's a lot of competition at the top, though!
Llama 405B response would be exactly what I expect
Either use a TypedDict if you want the keys to be in a specific set, or, in your case since both the keys and the values are static you should really be using an Enum
e.g. Almost the entire market relies upon Attention Is All You Need paper detailing transformers, and it would be an entirely different market if Google had held that as a trade secret.
I really am not expecting Apple or Microsoft to discover AGI and ferret it away for profitability purposes. Strictly speaking, I don't think superhuman intelligence even exists in the domain of text generation.
Of course, I agree that Stability AI made Stable Diffusion freely available and they're worth orders of magnitude less than OpenAI. To the point they're struggling to keep the lights on.
But it doesn't necessarily make that much difference whether you openly share the inner technical details. When you've got a motivated and well financed competitor, merely demonstrating a given feature is possible, showing the output and performance and price, might be enough.
If OpenAI adds a feature, who's to say Google and Facebook can't match it even though they can't access the code?
Given how good the model is terms of the quality vs speed tradeoff, they must have something.
I would guess that in that timeline, Google would never have been able to learn about the incredible capabilities of transformer models outside of translation, at least not until much later.
The current crop of benchmarks might not reflect these gains, by the way.
They are by no means bad, but I am now mostly interested in long context competency. We need benchmarks that force the LLM to complete multiple tasks simultaneously in one super long session.
Anyway here's the sales page. the widget subscription is so premium you won't even miss the subscription fee.
We should still be skeptical because often want to claim to be better or have unearned answers, but I don't think the motive to lie is quite as strong as a salesman's.
It's not peer-reviewable in any shape or form.
That said, doing that is slow and people will need to make decisions before that is done.
It's weird. In the late 2010s it seems like people were wising up to the idea that you can't implicitly trust big tech companies, even if they have nap pods in the office and have their first day employees wear funny hats. Then ChatGPT lands and everyone is back to fully trusting these companies when they say they are mere months from turning the world upside down with their AI, which they say every month for the last 12-24 months.
In the 2000s we only had Microsoft, and none of us were confused as to whether to trust Bill Gates or not...
https://www.wheresyoured.at/pop-culture/
> What makes this interview – and really, this paper — so remarkable is how thoroughly and aggressively it attacks every bit of marketing collateral the AI movement has. Acemoglu specifically questions the belief that AI models will simply get more powerful as we throw more data and GPU capacity at them, and specifically ask a question: what does it mean to "double AI's capabilities"? How does that actually make something like, say, a customer service rep better? And this is a specific problem with the AI fantasists' spiel. They heavily rely on the idea that not only will these large language models (LLMs) get more powerful, but that getting more powerful will somehow grant it the power to do...something. As Acemoglu says, "what does it mean to double AI's capabilities?"
I’m about Zucks age, and have been following his career/impact since college; it’s been roughly a cosine graph of doing good or evil over time :) I think we’re at 2pi by now, and if you are correct maybe it hockey-sticks up and to the right. I hope so.
If LLMs end up being the platform of the future, Zuck doesn't want OpenAI/Microsoft to be able to monopolize it.
> Other companies sell widgets. We have a bunch of widget-making machines and so we released a whole bunch of free widgets. We noticed that the widgets got better the more we made and expect widgets to become even better in future. Anyway here's the free download.
Given that Meta isn't actually selling their models?
Your response might make sense if it were to something OpenAI or Anthropic said, but as is I can't say I follow the analogy.
- flex
- deal a blow to Altmann
But it has nothing to do with LLMs (and interestingly enough they aren't opening their recommendation tech).
I mean, going by their own model evals on various benchmarks (https://llama.meta.com/), Llama 405b scores anywhere from a few points to almost 10 points more than than Llama 70b even though the former has ~5.5x more params. As far as scale in concerned, the relationship isn't even linear.
Which in most cases makes sense, you obviously can't get a 200% on these benchmarks, so if the smaller model is already at ~95% or whatever then there isn't much room for improvement. There is, however, the GPQA benchmark. Whereas Llama 70b scores ~47%, Llama 405b only scores ~51%. That's not a huge improvement despite the significant difference in size.
Most likely, we're going to see improvements in small model performance by way of better data. Otherwise though, I fail to see how we're supposed to get significantly better model performance by way of scale when the relationship between model size and benchmark scores is nowhere near linear. I really wish someone who's team "scale is all you need" could help me see what I'm missing.
And of course we might find some breakthrough that enables actual reasoning in models or whatever, but I find that purely speculative at this point, anything but inevitable.
The problem with this strategy is that it's really tough to compete with open models in this space over the long run.
If you look at OpenAI's homepage right now they're trying to promote "ChatGPT on your desktop", so it's clear even they realize that most people are looking for a local product. But once again this is a problem for them because open models run locally are always going to offer more in terms of privacy and features.
In order for proprietary models served through an API to compete long term they need to offer significant performance improvements over open/local offerings, but that gap has been perpetually shrinking.
On an M3 macbook pro you can run open models easily for free that perform close enough to OpenAI that I can use them as my primary LLM for effectively free with complete privacy and lots of room for improvement if I want to dive into the details. Ollama today is pretty much easier to install than just logging into ChatGPT and the performance feels a bit more responsive for most tasks. If I'm doing a serious LLM project I most certainly won't use proprietary models because the control I have over the model is too limited.
At this point I have completely stopped using proprietary LLMs despite working with LLMs everyday. Honestly can't understand any serious software engineer who wouldn't use open models (again the control and tooling provided is just so much better), and for less technical users it's getting easier and easier to just run open models locally.
OpenAI did a good move with making GPTo mini so dirty cheap that it's faster and cheaper to run than LLama 3.1 70B. Most consumers will interact with LLM via some apps using LLM API, Web Panel on desktop or native mobile app for the same reason most people use GMail etc. instead of native email client. Setting up IMAP, POP etc is for most people out of reach the same like installing Ollama + Docker + OpenWebUI
App developers are not gonna bet on local LLM only as long they are not mainstream and preinstalled on 50%+ devices.
In my opinion, they've found that intelligence with current architecture is actually an S-curve and not an exponential, so trying to make progress in other directions: UX and EQ.
https://nicholascharriere.com/blog/thoughts-openai-spring-re...
For example, has anyone ever attempted image -> html/css model? Seems like it be great if I can draw something on a piece of paper and have it generate a website view for me.
(May work poorly of course, and the sample I think I saw a year ago may well be cherry picked)
Have you tried upload the image to a LLM with vision capabilities like GPT-4o or Claude 3.5 Sonnet?
There are already companies selling services where they generate entire frontend applications from vague natural language inputs.
My hope now is that someone will figure out a way to separate intelligence from knowledge - i.e. train a model that knows how to interpret the wights of other models - so that training new intelligent models wouldn't require training them on a petabyte of data every run.
I had a discussion with a friend about doing this, but for CNC code. The answer was that a model trained on a narrow data set underperforms one trained on a large data set and then fine tuned with the narrow one.
[0] not technically complete depreciation, since for example 4o mini is widely believed to be a distillation of 4o, so 4o's investment still carries over into 4o mini
I don't know if the utility is worse than an LLM that is SOTA for 2 months that no one even bothers switching to however - at least the marvel slop is being used for entertainment by someone. I think the market is definitely prioritizing the LLM researcher over Disney's latest slop sequel though so whoever made that comparison can rest easy, because we'll find out.
I thought that was the allure, something that's camp funny and an easy watch.
I have only watched a few of them so I am not fully familiar?
Has there been any indication that we're improving the lives of millions of people?
If I need to babysit a junior developer fresh out of school and review every single line of code it spits out, I can find them elsewhere
I think GPT5 will tell if OpenAI hit a plateau.
Sam Altman has been quoted as claiming "GPT-3 had the intelligence of a toddler, GPT-4 was more similar to a smart high-schooler, and that the next generation will look to have PhD-level intelligence (in certain tasks)"
Notice the high degree of upselling based on vague claims of performance, and the fact that the jump from highschooler to PhD can very well be far less impressive than the jump from toddler to high schooler. In addition, notice the use of weasel words to frame expectations regarding "the next generation" to limit these gains to corner cases.
There's some degree of salesmanship in the way these models are presented, but even between the hyperboles you don't see claims of transformative changes.
buddy every few weeks one of these bozos is telling us their product is literally going to eclipse humanity and we should all start fearing the inevitable great collapse.
It's like how no one owns a car anymore because of ai driving and I don't have to tell you about the great bank disaster of 2019, when we all had to accept that fiat currency is over.
You've got to be a particular kind of unfortunate to believe it when sam altman says literally anything.
Progress is not slowing down but it gets harder to quantify.
I think you're just seeing the "make it work" stage of the combo "first make it work, then make it fast".
Time to market is critical, as you can attest by the fact you framed the situation as "on par with GPT-4o and Claude Opus". You're seeing huge investments because being the first to get a working model stands to benefit greatly. You can only assess models that exist, and for that you need to train them at a huge computational cost.
It feels like ChatGPT won the time to market war already.
I'll switch again soon as something better is out.
You may be right with the average person on the street, but I wonder how many have lost interest in LLM usage and cancelled their GPT plus sub.
This is similar to the early days of the search engine wars, the browser wars, and other categories where a user can easily adopt, switch between and use multiple. It's not like the cellphone OS/hardware war, PC war and database war where (most) users can only adopt one platform at a time and/or there's a heavy platform investment.
I’m also optimistic about building better (rather than bigger) datasets to train on.
Yes. This is exactly why I'm skeptical of AI doomerism/saviorism.
Too many people have been looking at the pace of LLM development over the last two (2) years, modeled it as an exponential growth function, and come to the conclusion that AGI is inevitable in the next ${1-5} years and we're headed for ${(dys|u)topia}.
But all that assumes that we can extrapolate a pattern of long-term exponential growth from less than two years of data. It's simply not possible to project in that way, and we're already seeing that OpenAI has pivoted from improving on GPT-4's benchmarks to reducing cost, while competitors (including free ones) catch up.
All the evidence suggests we've been slowing the rate of growth in capabilities of SOTA LLMs for at least the past year, which means predictions based on exponential growth all need to be reevaluated.
If you have a machine that converts mass into energy and then uses that energy to increase the rate at which it operates, you could rightfully say that it will level off well before consuming all of the mass in the universe. You just can't say that next week after it has consumed all of the mass of Earth.
Example:
w której gwarze jest słowo ekspres i co znaczy?
Słowo "ekspres" występuje w gwarze śląskiej i oznacza tam ekspres do kawy. Jest to skrót od nazwy "ekspres do kawy", czyli urządzenia służącego do szybkiego przygotowania kawy.
The correct answer is that "ekspres" is a zipper in Łódź dialect.But we could add internal thoughts-- we could make the model generate tokens that aren't part of its output but are there for it to better figure out its next token. This was tried QuietSTAR.
Hochreiter is also active with alternative models, and there's all the microchip design companies, Groq, Etched, etc. trying to speed up models and reduce model running cost.
Therefore, I think there's room for very great improvements. They may not come right away, but there are so many obvious paths to improve things that I think it's unreasonable to think progress has stalled. Also, presumably GPT-5 isn't far away.
Why do we presume that? People were saying this right before 4o and then what came out was not 5 but instead a major improvement on cost for 4.
Is there any specific reason to believe OpenAI has a model coming soon that will be a major step up in capabilities?
I assume that this won't take forever, but will be done this year. A couple of months, not more.
It feels like there’s an assumption in the community that this will be almost trivial.
I suspect it will be one of the hardest tasks humanity has ever endeavoured. I’m guessing it has already been tried many times in internal development.
I suspect if you start creating a feedback loop with these models they will tend to become very unstable very fast. We already see with these more linear LLMs that they can be extremely sensitive to the values of parameters like the temperature settings, and can go “crazy” fairly easily.
With feedback loops it could become much harder to prevent these AIs from spinning out of control. And no I don’t mean in the “become an evil paperclip maximiser” kind of way. Just plain unproductive insanity.
I think I can summarise my vision of the future in one sentence: AI psychologists will become a huge profession, and it will be just as difficult and nebulous as being a human psychologist.
I'm in the process of spinning out one of these tools into a product: they do not. They become smarter at the price of burning GPU cycles like there's no tomorrow.
I'd go as far as saying we've solved AGI, it's just that the energy budget is larger than the energy budget of the planet currently.
High temperature will obviously lead to randomness, that's what it, evening out the probabilities of the possibilities for the next token. So obviously a high temperature will make them 'crazy' and low temperature will lead to deterministic output. People have come up with lots of ideas about sampling, but this isn't really an instability of transformer models.
It's a problem with any model outputing probabilities for different alternative tokens.
What if they have two teams? One dedicated to optimizing (cost, speed, etc) the current model and a different team working on the next frontier model? I don't think we know the growth curve until we see gpt5.
I'm prepared to be wrong, but I think that the fact that we still haven't seen GPT-5 or even had a proper teaser for it 16 months after GPT-4 is evidence that the growth curve is slowing. The teasers that the media assumed were for GPT-5 seem to have actually been for GPT-4o [0]:
> Lex Fridman(01:06:13) So when is GPT-5 coming out again?
> Sam Altman(01:06:15) I don’t know. That’s the honest answer.
> Lex Fridman(01:06:18) Oh, that’s the honest answer. Blink twice if it’s this year.
> Sam Altman(01:06:30) We will release an amazing new model this year. I don’t know what we’ll call it.
> Lex Fridman(01:06:36) So that goes to the question of, what’s the way we release this thing?
> Sam Altman(01:06:41) We’ll release in the coming months many different things. I think that’d be very cool. I think before we talk about a GPT-5-like model called that, or not called that, or a little bit worse or a little bit better than what you’d expect from a GPT-5, I think we have a lot of other important things to release first.
Note that last response. That's not the sound of a CEO who has an amazing v5 of their product lined up, that's the sound of a CEO who's trying to figure out how to brand the model that they're working on that will be cheaper but not substantially better.
[0] https://arstechnica.com/information-technology/2024/03/opena...
that's interesting. Do you have a rough percentage of this?
Does this mean these connections have no influence at all on output?
Frankly I just don't understand the economics of training a foundation model. I'd rather own an airline. At least I can get a few years out of the capital investment of a plane.
If you are sitting on 1 billions $ of GPU capex, what's $50 million in energy/training cost for another incremental run that may beat the leaderboard?
Over the last few years the market has placed its bets that this stuff will make gobs of money somehow. We're all not sure how. They're probably thinking -- it's likely that whoever has a few % is going to sweep and take most of this hypothetical value. What's another few million, especially if you already have the GPUs?
I think you're right -- we are towards the right end of the sigmoid. And with no "killer app" in sight. It is great for all of us that they have created all this value, because I don't think anyone will be able to capture it. They certainly haven't yet.
One thing that `exponentialists` forget is that each step also requires exponentially more energy and resources.
Opus is the largest model, but of the Claude 3 family. Claude 3.5 is the newest family of models, with Sonnet being the middle sized 3.5 model - and also the only available one. Regardless, it's better than Opus (the largest Claude 3 one).
Presumably, a Claude 3.5 Opus will come out at some point, and should be even better - but maybe they've found that increasing the size for this model family just isn't cost effective. Or doesn't improve things that much. I'm unsure if they've said anything about it recently.
I.e. Opus is the largest and best model of each family but Sonnet is the first model of the 3.5 family and can beat 3's Opus in most tasks. When 3.5 Opus is released it will again outpace the 3.5 Sonnet model of the same family universally (in terms of capability) but until then it's a comparison of two different families without a universal guarantee, just a strong lean towards the newer model.
Claude Sonnet 3.5 outperforms GPT-4o by a significant margin on every one of my use cases.
What do you use it for?
We need to figure out how to measure intelligence that is greater than human.
Math problems being one of them, if only LLMs were good at pure math. Another possibility is graph problems. Haven't tested this much though.
Why does the chart below say the "Function Calling" accuracy is about 50%? Does that mean it fails half the time with complex operations?
Location and population of Paris, France
A parallel function calling LLM could return: {
"role": "assistant",
"content": "",
"tool_calls": [
{
"function": {
"name": "get_city_coordinates",
"arguments": "{\"city\": \"Paris\"}"
}
}, {
"function": {
"name": "get_city_population",
"arguments": "{\"city\": \"Paris\"}"
}
}
]
}
Indicating that you should execute both of those functions and return the results to the LLM as part of the next prompt.AnthropicProvider('claude-3-haiku-20240307') Median Latency: 1.61 | Aggregated speed: 122.50 | Accuracy: 44.44%
MistralProvider('open-mistral-nemo') Median Latency: 1.37 | Aggregated speed: 100.37 | Accuracy: 51.85%
OpenAIProvider('gpt-4o-mini') Median Latency: 2.13 | Aggregated speed: 67.59 | Accuracy: 59.26%
MistralProvider('mistral-large-latest') Median Latency: 10.18 | Aggregated speed: 18.64 | Accuracy: 62.96%
AnthropicProvider('claude-3-5-sonnet-20240620') Median Latency: 3.61 | Aggregated speed: 59.70 | Accuracy: 62.96%
OpenAIProvider('gpt-4o') Median Latency: 3.25 | Aggregated speed: 53.75 | Accuracy: 74.07% |
Apparently Llama 3.1 relied on artificial data, would be very curious about the type of data that Mistral uses.
Besides, parameter redundancy seems evidenced. Front-tier models used to be 1.8T, then 405B, and now 123B. Would front-tier models in the future be <10B or even <1B, that would be a game changer.
Is there a benchmark or something similar that compares this "quality" across different models?
Maybe they are running it on proprietary or semi proprietary hardware but if they dont, how much does the market no where various shipments of NVIDEA processors ends up?
I imagine most intelligence agencies are in need of vast quantities.
I presume is M$ announces new availability of AI compute it means they have received and put into production X Nvidiam, which might make it possible to guesstimate within some bounds how many.
Same with other open market compute facilities.
Is it likely that a significant share of NVIDEA processors are going to government / intelligent / fronts?
It's almost useless because I literally can't use it.
Update: https://support.anthropic.com/en/articles/8325612-does-claud...
45 messages per 5 hours is the limit for Pro users, less if Claude is wordy in its responses—which it always is. I hit that limit so fast when I'm investigating something. So annoying.
They used to let you select another, worse model but I don't see that option anymore. le sigh
I tried Codestral and nothing came close. Not even slightly. It was the only LLM that consistently put out code for me that was runnable and idiomatic.
these benchmarks are as good as random hardware ones apple or intel pushes to sell their stuff. in the real world, most people will end up with some modifications for their specific use case anyways. for those, i argue, we already have "capable enough" models for the job.
Do they mean inference done on a single machine?
(I wouldn't be surprised if GPT-4o mini is small enough to fit on a single large instance though, would explain how they could drop the price so much.)
I mean more on a model performance level though. It's been shown that something trained in one language trains the model to be able to output it in any other language it knows. There's quality human data being left on the table otherwise. Besides, translation is one of the few tasks that language models are by far the best at if trained properly, so why not do something you can sell as a main feature?
At least from a distance it seems like training a multilingual state of the art model might well be easier than a monolingual one.
A faster evolving approach to AI is coming out this year that will smoke anyone who still uses the term "license" in regards to ideas [1].
[0] https://breckyunits.com/eta.html [1] https://breckyunits.com/freedom.html
https://github.com/breck7/breckyunits.com/blob/afe70ad66cfbb...
I think history indicates that some commercial works can also triumph.
It explains everything.
Open source evolves faster, and always out competes closed source.
> I think history indicates that some commercial works can also triumph.
Nope. Not in the long run. I can't think of a single exception.