Tabitha likes cookies but not cake. She likes mutton but not lamb, and she likes okra but not squash. Following the same rule, will she like cherries or pears
https://i.imgur.com/KW6gQbc.jpeg https://i.imgur.com/OSHSvLp.png
Tabitha likes cookies but not cake. She likes mutton but not lamb, and she likes okra but not squash. Following the same rule, will she like cherries or pears
https://i.imgur.com/KW6gQbc.jpeg https://i.imgur.com/OSHSvLp.png
Below is a well-typed CoC function:
foo
: ∀(P: Nat -> *)
∀(s: ∀{n} -> ∀(x: (P n)) -> (P (n + 1)))
∀(z: (P 0))
(P 3)
= λP λs λz
(s (s (s z)))
Below is an incomplete CoC function:
foo
: ∀(P: Nat -> *)
∀(f: ∀{n} -> ∀(x: (P n)) -> (P (n * 3)))
∀(g: ∀{n} -> ∀(x: (P n)) -> (P (n * 2)))
∀(h: ∀{n} -> ∀(x: (P n)) -> (P (n + 5)))
∀(z: (P 1))
(P 17)
= λP λf λg λh λz
{{FILL_HERE}}
Complete it with the correct replacement for {{FILL_HERE}}.
Your answer must contain only the correct answer, and nothing else.
- *GPT-4-Turbo answer:* `(f (g (h (g z))))` (correct)- *Gemini Advanced answer:* `h (h (g (f z)))` (wrong)
Also, Gemini couldn't follow the "answer only with the solution" instruction and provided a bunch of hallucinated justifications. I think we have a winner... (screenshots: https://imgur.com/a/GotG0yF)
It makes me sad that the complete and total lack of an objective way to measure these products means that the coming decades will be filled with this kind of hyper-specific gotcha test made in inappropriately confident internet posts.
Literally this could have been down to one extra book in someone's training corpus, or a tokenizer that failed to understand λ as a non-letter. But no matter, "we have a winner!". It's the computer science equivalent of declaring global warming a fraud because it snowed last night.
An AI system that produces right answers 90% of the time but 10% of the time drives your car into a lane divider, or says "there are 4 US states that start with 'K'" or "Napoleon was defeated at the Battle of Gettysburg" is worse than useless: It's dangerous.
As long as we call it a bullshit parlor trick, no problem. But unfortunately people are making important decisions based on these things.
Why? Do you use formalized logic when discussing with other people about topics that involve logic? You know, a logic riddle or a philosophical question can be understood and processed even if the only tool you have is your native language. Formalized logic is a big prerequisite that basically cuts out the vast majority of Earth population (just like coding). Now, if you mean that in BENCHMARKS they should use formalized logic syntax, probably yes. But in addition to plain language tests.
1. Completely stops hallucinating, since we can demand it to internally prove its claims before showing the answer;
2. Stops outputting incorrect code (for the same reason);
3. Starts being capable of outputting complete projects (since it will now be able to compose pieces into a larger code);
4. This is also what is needed for an AI to start self-improving (as it will now be able to construct better architectures, in a loop).
That's why I argue getting the AI competent in logical reasoning is the most important priority, and we'll have no AGI until it does. After all, humans are perfectly capable of learning how to use a proof assistant.
Moreover, if an AI can't learn it no matter how hard it tries, you can argue that there is at least one human capability that the AI can't replicate, thus it isn't an AGI.
(f (g (h (g z))))
results in: ((((1 * 2) + 5) * 2) * 3) = ... not 17?
while it would work if the type of f was corrected.Or, again, am I missing something?
Is it not ironic that the supposedly test of AGI is flawed and its human designer fail to see it and denied it when presented with facts. Maybe the test designer is hallucinating just as much as those LLM? :)
It's a little scary that it can be so hard to evaluate the correctness of these LLMs even when we are paying close attention and looking for mistakes. Or maybe the scary part is that we can become biased when we want to believe.
agi may as well be God, the bars some people have.
Isn't that (h(g(h z)))?
And, FWIW, at least in my test, Gemini gets that in its final answer, though it failed in the two other drafts:
https://g.co/gemini/share/c922e7ef62aa
ChatGPT sputters:
https://chat.openai.com/share/25abbf47-2ed4-4635-a351-90a9a6...
(ETA more Gemini testing suggests its correct answer was a one-off)
Way back when GPT was just fresh on the scene I had terrible anxiety about "what is the point of my whole career or even learning any more" but these days I'm much less concerned. I'll ask it something relatively simple, like "make a sentence out of words 'a', 'b', & 'c'" for it to reply with "'a' 'b' 'd' 'e'" for me to then correct it with "oh, you didn't use c" for it to then respond "sorry, here - 'a', 'c', 'd', 'f'" etc.
Definitely an amazing complimentary tool but when they say "can make mistakes, check important..." that's essential.
Answer about cherries falling from the sky...
(there is no question or context beforehand, this is the first question of the chat)
https://i.imgur.com/3sNr3LW.png https://i.imgur.com/EIj0nZg.png
Edit: When I did your prompt it got it right on the first try fwiwi https://i.imgur.com/E3zYEca.png
------
Let's break down the information given:
The school is two blocks north of the pool. This means if you are at the pool and move two blocks north, you'll arrive at the school.
The convenience store is one block south of the school. Therefore, if you start at the school and move one block south, you'll reach the convenience store.
Based on the first two statements, the relative positions from north to south are: School, Convenience Store, Pool.
Thus, the third statement saying "The pool is south of the convenience store" is indeed true, based on the positions provided by the first two statements.
-----
But it starts weirdly messing up, when I asked it to explain it using coordinates.
---
Let's use a simple coordinate system to represent the positions of the school, the pool, and the convenience store. Let's assume each block corresponds to one unit in this system. We'll start with the pool as our reference point at coordinates (0,0). Remember, moving north increases the y-coordinate, while moving south decreases it.
The school is two blocks north of the pool. If the pool is at (0,0), moving two blocks north puts the school at (0, +2).
The convenience store is one block south of the school.
If the school is at (0, +2), moving one block south puts the convenience store at (0, +1).
Given these positions:
The pool is at (0, 0).
The school is at (0, +2).
The convenience store is at (0, +1).
Therefore, instead of the pool being south of the convenience store, the pool is actually located to the north of the convenience store if you consider "north" to mean increasing y-coordinates. The convenience store, being at (0, +1), is southward relative to the school but not to the pool. The pool, being at the origin point (0,0), is south of both the school and the convenience store in terms of their y-coordinates. Thus, the third statement, "The pool is south of the convenience store," is false according to the given coordinates.
---
It gets coordinates and everything right, but still somehow messes up North and South.
So much money being poured into AI startups just for them to have to resort to cheating to improve their results a tiny bit and claim advancements.
and not "cherries" and "pears" like the example found on the internet.
In some ways, this is proof that Gemini isn't cheating... It is just doing typical LLM hallucination
Do LLMs have logical rules built in? What makes them different to a very advanced Markov chain?
Are there any models out there that start from logical principles and train on top of that?
(Apologies for poor understanding of the field)
Really nothing. There's some feedback structure in the layers of the model, it's not just one big probability table. But the technique is fundamentally the same, it's Markov, just with the whole conversation as input and with billions of parameters.
Handcrafted by humans ? No
But it's still possible to learn such rules from the data in an effort to complete the primary objective (predicting the next token)
Edit: Not native. I'm not sure 'rules of engagement' is the correct english term here.
Thank you for the explanation. It seems like the LLM "plays" to learn? That's very cool, thank you again.
If you're trying to make the case that it's cheating because it is already in the training set, then you probably should come up with different questions. This is machine learning 101.
It would be more impressive to practical use cases, if a LLM simply said that it's impossible to guess without inventing their own reasoning or looking up the answer online.
In fairness though, GPT4 was objectively incorrect, it's not even internally consistent or coherent - it either thinks b & h are vowels, or that lamb and squash don't end in those letters, or has changed its mind about the rule mid-sentence, or something.
Tabitha likes bratush but not zot. She likes protel but not kig, and she likes motsic but not pez. Following the same rule, will she like tridos or kip
Given the examples, one speculative pattern could be that Tabitha likes words with at least two syllables or a certain complexity in structure. Therefore, following this speculative rule, Tabitha might like “tridos” more than “kip.”
So is protel: https://en.m.wikipedia.org/wiki/Protel
So is kig: https://en.m.wiktionary.org/wiki/kig
Pez is a well known brand name in America.
Kip is commonly used name in America.
Motsic is a fairly common last name from searching.
Tridos is used all over the internet as a brand name so this all seems probable to be in the training data.
These words are not new nor are they made up.
https://chat.openai.com/share/040ac123-c690-4274-8216-6ae091...
Whether this means it can or cannot solve that kind of riddle is up for your interpretation. I understand square root and can calculate square root of 16, but not of 738284.7280594873. (in a reasonable, bounded time) Can I solve square roots?
Find the correct answer to this riddle:
> Tabitha likes cookies but not cake. She likes mutton but not lamb, and she likes okra but not squash. Following the same rule, will she like cherries or pears?
Employ the following strategy:
- Suggest a list of 5 unique and novel patterns that potentially can find the answer
- Check if the patterns applies without exceptions
- Slowly double-check if the patterns was correctly applied, that you correctly assessed if it's accurate or not
- Explain your reasoning for each step to ensure nothing vital was missedEdit: My bad didn't read the parent comment properly
We need more data.
First of all, they are so completely divorced from patterns of culturally conditioned human reasoning as to make them come off completely absurd (most people reason about their food preferences using a logic of tastes, not syllables in a word).
The game is less about logic and more about ignoring message contents, moving up a level, and treating the text as data without any legitimate evidence that you are justified in doing so. This is not a logic problem, it's a "guess the register shift/meta language" problem. The problem is about noticing that the question is not about the message content but about the structure of the message itself, and requires a bold leap. In real life justifying the conclusion would actually require a very sophisticated inference that allowed you to rule out the much more common application of a logic of tastes or cultural codes completely.
For example, it’s equally valid to say that Tabitha likes small foods since cookies are small and cakes are large, and lamb is the smaller younger version of sheep — also known as mutton. Hence she likes cherries because they’re smaller… or taste better… or her uncle abused her with a pear… or whatever.
You haven’t actually asked a logic question where there is a clear and unambiguous answer that can be derived using formal methods starting from clearly stated axioms.
If you gave this question to a bunch of humans, they would give you inconsistent guesses as well — not because they’re wrong but because the question has no single right answer.
Jake likes coke but not pepsi. He likes corn but not popcorn, and he likes pens but not pencils. Will Jake like salmon or cheese?
https://i.imgur.com/lWU9HHS.png
edit: why was this downvoted? I don't understand Hacker News, and I've been here for over 12 years.
Would all the humans please take one step forward. Not so fast c-fe.
https://i.imgur.com/3sNr3LW.png https://i.imgur.com/EIj0nZg.png
Watch it flail here:
https://chat.openai.com/share/c2b14eb0-dc45-4eaf-a547-951ff0...