> Yes, in the Netherlands, jenever (also known as genever) is the traditional spirit that represents the country. Jenever is a type of Dutch gin that has a distinctive flavor, often made from malt wine and flavored with juniper berries. It has a long history in the Netherlands, dating back to the 16th century, and is considered the precursor to modern gin.
> Jenever comes in two main types: oude (old) jenever, which has a richer, maltier flavor and is often aged in wooden casks, and jonge (young) jenever, which is lighter and has a more neutral taste. Both types can be enjoyed straight or used in cocktails.
> In addition to jenever, the Netherlands is also known for its liqueurs, such as advocaat, a rich and creamy drink made from eggs, sugar, and brandy, often enjoyed as a dessert or in cocktails. However, jenever remains the most iconic spirit associated with Dutch culture.
This is completely wrong. Jenever certainly is very Dutch, but no one would say it is iconic as the Dutch spirit. For example, if you asked up north in Friesland, they would say Berenburg.
This happens literally every time. Someone always says "ChatGPT can do this!", but then within one or two prompts, its gets it wrong.
It is equally possible to "reason" about things you already know, as about things you've just been told. In fact, the capacity to speculatively, without prompting attempt such reasoning is a big part of cognition.
From your comment it would seem that you are disputing jenever's popularity by saying jenever is more popular...
Perhaps it was a good faith mistake? If so, that would imply that the AI knows more about jenever than you?
For example, France has been trending towards beer more and more, and within a few decades they might be consuming more beer than wine. But even then, the French wouldn't slowly start to say beer represents France.
Furthermore, "just adding some herbs" does a large disservice to the flavor change of Berenburg. Jenever (aka jonge/unaged jenever) is straight-up vile. I've heard it described by expats as "having the worst elements of both cheap gin and cheap whisky".
Berenburg in comparison is spicy and vanilla-y and actually debatebly enjoyable.
Aged/oude jenever is much closer to Berenburg (or Berenburg to aged jenever), also with hints of vanilla and spices.
But, virtually no one except for dusty old men orders aged jenever. The kind ordered by far the most is jonge jenever, and then its only in a sense of "haha lets drink this terrible thing" or "let's get shitfaced quick".
If o1 supposedly "oneshots every question", it should have been aware of these nuances instead of just confidently assigning jenever as 'the' spirit of the Dutch.
The question in the prompt comes off to me as a sort of qualitative determination rather than asking about pure factual information (is there an officially designated spirit). As such I don't think it can necessarily be right or wrong.
Anyway, I'm not sure what you'd expect. In terms of acquisition of knowledge, LLMs fundamentally rely on a written corpus. Their knowledge of information that is passed through casual spoken conversation is limited. Sure, as human beings, we rely a great deal on the latter. But for an LLM to lack access to that information means that it's going to miss out on cultural nuances that are not widely expressed in writing. Much in the same way that a human adult can live in a foreign country for decades, speaking their adopted language quite fluently, but if they don't have kids of their own, they might be quite ignorant of that country's nursery rhymes and children's games, simply because they were never part of their acquired vocabulary and experience.
I was just proving the people wrong that were saying akin to that o1 was "oneshotting every question".
I completely understand from how LLMs work that they wouldn't be able to get this right. But then people shouldn't be proudly be pronouncing that o1 (or any model) is getting every question right, first time.
I have opened the question of why you thought jenever was not jenever, and your non-responsiveness I think compels the fact that AI was more correct in your contrived instance.
What's expected is an ability to identify trick questions, i.e., to recognize fundamental problems in the phrasing of a question rather than trying to provide a "helpful" answer at all costs.
This corresponds to one of the many reasons LLM output is banned on Stack Overflow.
But, like Zahlman points out, its a trick question, and instead of admitting it doesn't know or even prepending "I don't know for sure, but:", it just burps up its best-effort answer. There is no one spirit that represents The Netherlands. If a LLM is so good it "oneshots any question", it should realize it doesn't have a unanimous answer and tell me.
Using just a plain old search engine, for things like "national drink of the netherlands" and simlar queries, I am directed to Wikipedia's Jenever page as the top hit, and Wikipedia's list of national drinks lists Jenever and Heineken as the entries for the Netherlands. Search engines also give page after page of travel guides and blog posts, most of which list Jenever at or near the top of of their listings. One travel guide calls it "the most famous Dutch spirit and most famous Amsterdam liquor, Jenever, also spelled Genever or simply Dutch gin."
I mean, if I had OpenAI’s resources I’d have a team tasked with monitoring social to debug trending fuck-ups. (Before that: add compute time to frequently-asked novel queries.)
This could even be automated; LLMs can sentiment-analyze social media posts to surface ones that are critical of LLM outputs, then automatically extract features of the post to change things about the running model to improve similar results with no intervention.
When asking an LLM to write a script for you, I would say 10 to 30 % of the time that it completely fails. Again, making up an API or just getting things straight up wrong.
Its very helpful, especially when starting from 0 with the beginner questions, but it fails in many scenarios.
that said, i don’t think this is a good test - i’ve seen it circling on twitter for months and it is almost certainly trained on similar tasks