HuggingChat: Chat with Open Source Models
huggingface.co
huggingface.co
For me this is one of the highest-impact and most-often-overlooked features of the ChatGPT Web UI (so much so that openai does not even include this feature in their native clients).
Two cars have a 100 mile race.
Car A drives 10mph. Car B drives
5mph but gets a 50 mile headstart.
Who wins?
So far, no LLM gets it consistently right. Most of the time, they confidently declare one of the cars as the clear winner.Sometimes the explanations are so convincing, that I begin to wonder if my own conclusion - that the race will be a tie - is wrong :)
And if I convert it from miles/mph to km/kmh then most of them get it right.
Two cars have a 160 km race. Car A drives 16kmh. Car B drives 8kmh but gets a 80km headstart. Who wins?
"Two cars have a 160km race. Care A drives at 16km/h. Car B drives at 8km/h but gets a 75km head start. Who wins?"
(or similar slight perturbations)
Results:
To determine who wins this race, we'll calculate the time it takes for each car to complete the race using the same formula as before: \[ \text{Time} = \frac{\text{Distance}}{\text{Speed}} \]
- Car A starts from 0 km and needs to cover 160 km at a speed of 16 km/h. - Car B starts from 75 km (due to the 75 km head start) and only needs to cover 85 km (since 160 km total distance minus the 75 km head start) at a speed of 8 km/h.
Let's calculate the time for each car.
Car A will complete the race in 10 hours, while Car B will take 10.625 hours to finish. Therefore, Car A wins the race.
Given that a gnorflork is a unit of distance...
> Draw a graph of the function on the blackboard, showing a and b and a crosshatched area representing the integral. Put an x on the horizontal axis. Erase the x and put a z there. Does that change the area? Erase the z and put a smiley face there. Does the area change? Why/why not?
Google gives you completely unrelated results, while the LLMs just make something up.
Twice in the past weeks I tried both ChatGPT and Gemini for a technical question. The true answer was "you can't do it"/"those devices don't have the feature you're asking about". In both cases they made up some answers that looked consistent but were completely bogus.
I think we'll see a lot of LLM use cases switching from generative responses back to good old fashioned search this year.
LLMs are confidently and repeatedly wrong about very simple things that a human would never be. They're also not even consistent; ask them one thing and then tell them they're wrong: they'll often apologise profusely and flip their answer right around. Tell them once again that they're wrong and they'll flip back again. Ask them about their thought process, and they'll just produce a load of waffle that looks like it describes a thought process but usually doesn't correspond to anything of the sort (and of course it doesn't; if you understand vaguely how these systems work it'd be a miracle if it did so).
Since one can invariably take any criticism of LLMs and immediately reply with "but some humans do that too", I don't think such lines of argument say very much.
It's the weird combination of superhuman memory, speed and wit and at times utterly subhuman logical consistency that makes it hard to interpret how 'intelligent' LLMs can be said to be. That said, they can still be very useful.
> LLMs are confidently and repeatedly wrong about very simple things that a human would never be.
It is a strong assertion. A substantial fraction of the population (maybe even the majority) would repeat phrases they learned. Some are about Earth being 5000 years old, some are political slogans, and some others are based on the knowledge they derived from TV series.
> Since one can invariably take any criticism of LLMs and immediately reply with "but some humans do that too", I don't think such lines of argument say very much.
On the contrary - it is interesting to see which problems are uniquely LLM (or human) and which others - are similar instances of the same, especially when we benchmark LLM not only versus some idealized human cognition but one of an ordinary person under ordinary conditions.
I think you may be right if you consider humanity as a whole, but when evaluating LLMs I think we should perhaps hold them to slightly higher standards. It seems silly to compare a system that clearly has an incredible level of knowledge and writing ability to toddlers or completely uneducated remote tribespeople (or even just bigoted people living in otherwise advanced societies). Since LLMs can operate on vastly different timescales to humans, we should probably evaluate them against reasonably knowledgeable humans given several hours and a chance to look back over their work as many times as they like before final submission. Since we’re trying to abstractly compare some sort of computational ability, we should probably give them equal amounts of computational resources/number of steps/clock cycles/etc. Given this setup, I think some of the ‘human mistakes’ we talk about would no longer apply.
Comparing systems like this is not an easy (or objective) matter anyway… I’m not claiming this is the only way of thinking about it and that all others are wrong.
> it is interesting to see which problems are uniquely LLM (or human)
I certainly agree with that. It’s fascinating to be able to observe and study something in real time that appears on the surface to approximate human thought so shockingly well whilst also being so different.
Data efficiency is crucial, but humans have a lot of pre-trained multi-modal data. For energy efficiency, ML is orders of magnitude more efficient.
Regarding the general powers (and weaknesses) of LLMs, I am shocked they work in an Artificial General(ish) Intelligence way. Text generation with GPT2 and GPT3 - sure, it still felt like a very advanced text autocomplete. GPT4 - here, I could have a bet (and lose money) that this level requires some form of reinforcement learning and a two-way interaction with the environment. Sure, there is some RL there, but I assumed something closer to a bot learning to talk with people, running and getting results from code, and looking at data online (for the training!), or maybe even - literally walking with a camera attached.
Yes, there is still a lot to do. There is a difference between a Go model beating novices, advanced players, and everyone, including world champions.
> The answer to this riddle is that both cars will win the race, as they have both driven a total distance of 100 miles. Car A has driven at a speed of 10mph for the entire race, while Car B has driven at a speed of 5mph and has a headstart of 50 miles. However, since the two cars have started from the same point and are traveling in the same direction, they will arrive at the finish line together.
that's a positive way to look at things hah. I wonder if its system prompt nudges responses in that direction vs something like "the race is a tie" or "both cars lose".
“To find out who wins, we can calculate the time it takes for each car to complete the race.
For Car A: - Speed = 10 mph - Distance = 100 miles - Time = Distance / Speed = 100 miles / 10 mph = 10 hours
For Car B: - Speed = 5 mph - Distance = 50 miles (because of the 50 mile headstart, it only needs to cover 50 miles) - Time = Distance / Speed = 50 miles / 5 mph = 10 hours
Both cars would finish the race at the same time, so it's a tie.”
I'm asking because a lot of people have only tried the 3.5 model and then lose interest when it fails on stupid ways. GPT-4 is not just quantitatively better, it's a qualitative jump ahead. Just in case you haven't tried it out yet (which I assume given your comment since gpt-4 gets that answer right) I encourage you to do it
Does it still try to make up shit when the answer is "it can't be done"?
It's not perfect and the classic failure modes of LLMs (the one you mentioned and others) are still present, but to a much much smaller degree.
I have a few prompts I use for specific tasks, where I ask it to answer tersely and it sometimes says so when something is impossible. Sometimes it hallucinates an answer. The likelihood for it to not hallucinate an answer seems to be proportional to desired terseness of the answer, as if the default instruction to "be helpful" biases it towards telling you something, anything
Of course I closed the tab and lost it. Yay for using lingering tabs as a "to check later" list.
Can anyone repost the app name(s) please?
But now I'll check all that get mentioned, thanks.
They're not Groq but Together (and Perplexity Labs) have the lowest latencies and fastest tokens per second of any commercial services available right now. Also the lowest prices afaik.
Based off Karpathy's BPE video highlighting ChatGPT being unable to reverse this string-> .DefaultCellStyle