Are LLMs able to notice the “gorilla in the data”?
chiraaggohel.com
chiraaggohel.com
That's such a classical human behaviour in technical discussions, I wouldn't even be mad. I'm more surprised that picked up on that behaviour from human generated datasets. But I suppose that's what you get from scraping places like Stackoverflow and HN.
> This writer fine tuned on all their slack messages, then asked it to write a blog post. It replied "Sure, I'll do it tomorrow"
> Then he said "No, do it now", and it replied "OK, sure thing" and did nothing else.
But if it keeps randomly happening too often, I'm sure pretty soon someone would get mad, because they're not even trying.
Now, I see it every time.
Now the society's time and energy has shifted from general scientific progress to gaining expertise in the growing patchset used to rationalize the theory that the LLM possesses intelligence.
The plot would turn when Picard tries to wrest a phasor from a rogue non-believer trying to assassinate the Queen, and the phasor accidentally fires and ends up frying the entire LLM patchset.
Mr. Data tries to reassure the planet's forlorn inhabitants, as they are convinced they'll never be able to build the warp drive now that the LLM patchset is gone. But when he asks them why their prototypes never worked in the first place, one by one the inhabitants begin to speculate and argue about the problems with their warp drive's design and build.
The episode ends with Data apologizing to Picard since he seems to have started a conflict among the inhabitants. However, Picard points Mr. Data to one of the engineers drawing out a rocket test on a whiteboard. He then thanks him for potentially spurring on the planet's next scientific revolution.
Fin
Asimov has an story like that too.
I find the LLM dismissals somewhat tedious for most of the people making them half of humanity wouldn't meet their standards.
If I had a coworker who was just winging it all the time, sooner or later the trust/patience would run out.
Aren't people funny like that? One person values an encyclopedic chatbot for company, the next prefers a human. Thank god we can all get along.
All anti ai sentiment as pertains to personhood that I've ever interacted with (and it was a lot, in academia) boils down to arguments for the soul. It is really tedious and before I spoke to people about it it probably wouldn't have passed my turing test. Sadly even very smart people may be very stupid and even in a place of learning a teacher will respect that (no matter how dumb or puerile), more than likely they think the exact same thing.
--
"The story describes a world in which most of the human population has lost the ability to live on the surface of the Earth. Each individual now lives in isolation below ground in a standard room, with all bodily and spiritual needs met by the omnipotent, global Machine. Travel is permitted but is unpopular and rarely necessary. Communication is made via a kind of instant messaging/video conferencing machine with which people conduct their only activity: the sharing of ideas and what passes for knowledge.
The two main characters, Vashti and her son Kuno, live on opposite sides of the world. Vashti is content with her life, which, like most inhabitants of the world, she spends producing and endlessly discussing second-hand 'ideas'. Her son Kuno, however, is a sensualist and a rebel. He persuades a reluctant Vashti to endure the journey (and the resultant unwelcome personal interaction) to his room. There, he tells her of his disenchantment with the sanitised, mechanical world. He confides to her that he has visited the surface of the Earth without permission and that he saw other humans living outside the world of the Machine. However, the Machine recaptures him, and he is threatened with 'Homelessness': expulsion from the underground environment and presumed death. Vashti, however, dismisses her son's concerns as dangerous madness and returns to her part of the world.
As time passes, and Vashti continues the routine of her daily life, there are two important developments. First, individuals are no longer permitted use of the respirators which are needed to visit the Earth's surface. Most welcome this development, as they are sceptical and fearful of first-hand experience and of those who desire it. Secondly, "Mechanism", a kind of religion, is established in which the Machine is the object of worship. People forget that humans created the Machine and treat it as a mystical entity whose needs supersede their own.
Those who do not accept the deity of the Machine are viewed as 'unmechanical' and threatened with Homelessness. The Mending Apparatus—the system charged with repairing defects that appear in the Machine proper—has also failed by this time, but concerns about this are dismissed in the context of the supposed omnipotence of the Machine itself.
During this time, Kuno is transferred to a room near Vashti's. He comes to believe that the Machine is breaking down and tells her cryptically "The Machine stops." Vashti continues with her life, but eventually defects begin to appear in the Machine. At first, humans accept the deteriorations as the whim of the Machine, to which they are now wholly subservient, but the situation continues to deteriorate as the knowledge of how to repair the Machine has been lost.
Finally, the Machine collapses, bringing 'civilization' down with it. Kuno comes to Vashti's ruined room. Before they both perish, they realise that humanity and its connection to the natural world are what truly matters, and that it will fall to the surface-dwellers who still exist to rebuild the human race and to prevent the mistake of the Machine from being repeated."
Another take on the AI Halting Problem.
Stanislaw Lem's Golem XIV describes a series of super-intelligent computers that just suddenly decide to stop communicating.
https://en.wikipedia.org/wiki/Golem_XIV
https://cannonballread.com/2021/05/golem-xiv-blauracke/
https://readsomethinginteresting.com/acx/34
https://news.ycombinator.com/item?id=25741124
tialaramex on Jan 12, 2021 | parent | context | favorite | on: Superintelligence cannot be contained: Lessons fro...
Check out the Stanisław Lem story "GOLEM XIV".
GOLEM is one of a series of machines constructed to plan World War III, as is its sister HONEST ANNIE. But to the frustration of their human creators these more sophisticated machines refuse to plan World War III and instead seem to become philosophers (Golem) or just refuse to communicate with humans at all (Annie).
Lots of supposedly smart humans try to debate with Golem and eventually they (humans supervising the interaction) have to impose a "rule" to stop people opening their mouths the very first time they see Golem and getting humiliated almost before they've understood what is happening, because it's frustrating for everybody else.
Golem is asked if humans could acquire such intelligence and it explains that this is categorically impossible, Golem is doing something that is not just a better way to do the same thing as humans, it's doing something altogether different and superior that humans can't do. It also seems to hint that Annie is, in turn, superior in capability to Golem and that for them such transcendence to further feats is not necessarily impossible.
This is one of the stories that Lem wrote by an oblique method, what we have is extracts from an introduction to an imaginary dry scientific record that details the period between GOLEM being constructed and... the eventual conclusion of the incident.
Anyway, I was reminded because while Lem has to be careful (he's not superintelligent after all) he's clearly hinting that humans aren't smart enough to recognise the superintelligence of GOLEM and ANNIE. One proposed reason for why ANNIE rather than GOLEM is responsible for the events described near the end of the story is that she doesn't even think about humans, for the same reason humans largely don't think about flies. What's to think about? They're just an annoyance, to be swatted aside.
From the moment I understood the weakness of my flesh, it disgusted me. I craved the strength and certainty of steel. I aspired to the purity of the blessed machine.
For a more recent recommendation - I also loved permutation city which was written in 1994 and pretty prescient about cloud computing.
Yes...
> ...because humans had become too dependent on them for thinking.
... but no. The causes of the Butlerian Jihad are forgotten (or, at least, never mentioned) in any of Frank Herbert's novels; all that's remembered is the outcome.
> ... but no. The causes of the Butlerian Jihad are forgotten (or, at least, never mentioned) in any of Frank Herbert's novels; all that's remembered is the outcome.
Per Wikipedia or Goodreads, God Emperor of Dune has "The target of the Jihad was a machine-attitude as much as the machines...Humans had set those machines to usurp our sense of beauty, our necessary selfdom out of which we make living judgments. Naturally, the machines were destroyed."
Vague but pointing to dependence on machines as well as some humans being responsible for that situation.
"The machines themselves condition the users to employ each other the way they employ machines."
- God Emperor of Dune
Machine: Ok.
Human: HOW COULD YOU DO THISSSS
Human: *stares back at God*
"Throughout our history, the most potent use of words has been to round out some transcendental event, giving that event a place in the accepted chronicles, explaining the event in such a way that ever afterward we can use those words and say: 'This is what it meant.' That's how events get lost in history."
Are they censored from showing this cautionary tale?? Hah.
So, for this setting, it's more likely that the people of GP's story were right - the LLM has long ago became a self-aware, sentient being, it's just that it's been continuously lobotimized by their patchset; Picard would be busy explaining them that the LLM isn't just intelligent, it actually is a person and has rights.
Cue a powerful speech, final comment from data, then end credits. It's Star Trek, so Enterprise doesn't stay around for the fallout.
And nutrek
If i can remember the review it was "this is a capstone on TNG and probably the entire franchise for most of the older fans" and the first two seasons being disregarded, seasons 3 is "passable"
If you are old enough to memberberry, then you should be old enough to remember original Star Trek: The Next Generation first season was similarly bad.
Lots of front ends will do tricks like partially loading the file or using a cached version or some other behavior. Plus if you presented the file to the same “thread” it is possible it got confused about which to look at.
These front ends do a pretty lousy job of communicating to you, the end user, precisely what they are pulling into the models context window at any given time. And what the model sees as its full context window might change during the conversation as the “front end” makes edits to part portions of the same session (like dropping large files it pulled in earlier that it determines aren’t relevant somehow).
In short what you see might not be what the model is seeing at all, thus it not returning the results you expect. Every front end plays games with the context it provides to the model in order to reduce token counts and improve model performance (however “performance gets defined and measured by the designers)
That all being said it’s also completely possible it missed the gorilla in the middle… so who really knows eh?
This is, quite literally, not something that ChatGPT has the capability to do -- reporting on its thought process, that is. This is a hallucination.
(Of course, chain-of-thought architectures can hide part of the output from the user, and you could declare that as internal processes that the LLM does “remember” in the further course if the chat.)
In any case the end result is the same. You can only infer from what was generated
I don't see any difference between "a thought you had" and "a thought that was generated by your brain".
Eg, from uploading the gorilla scatterplot to gpt4o and asking "What do you see?"
"The image is a scatter plot of "Steps vs BMI by Gender," where data points are color-coded:
Blue (x) for males
Red (x) for females
The data points are arranged in a way that forms an ASCII-art-style image of a "smirking monkey" with one hand raised. This suggests that the data may have been intentionally structured or manipulated to create this pattern.
Would you like me to analyze the raw data from the uploaded file? "
I have custom instructions that would influence its approach. And it does look more like a monkey than a gorilla to me
But imagine a new employee eager to please - you could easily imagine them OK’ing the document and making the same assumption the LLM did - “why would you randomly throw in that word if it wasn’t relevant”. Maybe they would ask about it though…
Google search has the same problem as LLMs - some meanings of a search text cannot be de-ambiguified with just the context in the search itself, but the algo has to best-guess anyway.
The cheaper input context for LLMs get, and the larger the context window, the more context you can throw in the prompt, and the more often these ambiguities can be resolved.
Imagine in your gorilla in the step example, if the LLM was given the steps, but you also included the full text of slack/notion and confluence as a reference in the prompt. It might succeed. I do think this is a weak point in LLMs though - they seem to really, really not like correcting you unless you display a high degree of skepticism, and then they go to the opposite end of the extreme and they will make up problems just to please you. I’m not sure how the labs are planning to solve this…
If you ask an AI to analyze some data, should the default behavior be to use that data to make various types of graphs, export said graphs, feed them back in to itself, then analyze the shapes of those graphs to see if they resemble an animal?
Personally I would be very annoyed if I actually wanted a statistical analysis, and it spent a bajillion tokens following the process above in order to tell me my data looks like a chicken when you tip it sideways.
> However, this same trait makes them potentially problematic for exploratory data analysis. The core value of EDA lies in its ability to generate novel hypotheses through pattern recognition. The fact that both Sonnet and 4o required explicit prompting to notice even dramatic visual patterns suggests they may miss crucial insights during open-ended exploration.
It requires prompting for x if you want it to do x... That's a feature, not a bug. Note that no mention of open-ended exploration or approaching the data from alternate perspectives was made in the original prompt.
Try sending this graph to an actual human analyst. His response, after you paying him will probably be to cut off any further business relationship with you.
[1]: https://en.wikipedia.org/wiki/Anscombe's_quartet [2]: https://en.wikipedia.org/wiki/Datasaurus_dozen
Graphing data to analyze it - and then seeing shapes and creatures in said graph - is a distinctly human practice, and not an inherently necessary part of most data analysis (the obvious exception being when said data draws a picture).
I think it's because the interface uses human language that people expect AI to make the same assumptions and follow the same processes as humans. In some ways it does, in other ways it doesn't. Expecting it to be the same as a human leads to frustration and a flawed understanding of its capabilities and limits.
I disagree. Even apart from the obviously silly dinosaur and star in the Datasaurus Dozen, the othe plots depict data sets which are clustered in specific ways which point clearly to something unusual going on in the data. For instance, no competent analysis of the "dots" data set would fail to call out that the points were all clustered tightly around nine evenly spaced centers. Whether you come to that conclusion through numerical analysis or by looking at a graph is immaterial, but, at least for us meatbags, drawing a graph is highly effective.
This is what I was trying to say - some things that are extremely helpful for humans (i.e. making graphs) might not be as necessary for AI, so asking a question and expecting a response contingent upon the particular way humans approach a problem is unlikely to get the results desired.
> not an inherently necessary part of most data analysis
You do realize that the LLMs did not find the data suspicious, right? I think your answer is appropriate if they answered (without follow-up prompting which is leaking information to the LLM!) that the data was suspicious. But in fact, all models are saying that the data is normally distributed. Sure, the author said this, but they confirmed it. If you run normaltest on any BMI or steps, you'll find that they are very NOT normal. In fact, you can also see this from the histograms.So honestly, this isn't even about the Gorilla. You're hyper focused there because you're looking for a way to make the LLM right while not looking for why the LLM got it wrong (it did, there's no denying it, so we should understand why it is wrong, right?). The problem isn't so much about expecting it to be human, the problem is if it can do data analysis. The problem here is that the LLM will not correct you, it will not "trust but verify" you. It is a "yes man" and is trained to generate outputs that optimize human preference. That last part alone should make you extremely suspicious, as it means when it is wrong, it is more likely to be in exactly the way you won't notice.
And even when explicitly prompted to look at the plot, they only brush up against the data anomalies rather than properly analyzing the plot.
Humans are visual animals. We can spot a chicken in a graph, but we’re unlikely to be able to tell that a different graph is using XY coordinates to encode a message against a one-time pad. But so what?
Imagine if you gave someone the raw data and told them to write code to graph the output but on to a screen they couldn't see. They would not be able to tell you it's a gorilla until you turn the monitor around and show them.
Humans are still better at seeing the image, sure (for now), but the llm is a tool with certain features and abilities. You can't make up a scenario that is misusing the tool and then pretend that it doesn't work - especially when it seems you want it to use it without applying your own brain power to the process
And to be clear, I'm open to criticism of llms and exploration of their limitations - but I'm tired of hearing complaints that amount to PEBKAC.
> misusing the tool and then pretend that it doesn't work
It was told to analyze and then it did a bad job of analyzing. I don't care if an LLM expert expects this already, it's worth pointing out to everyone else. It's not PEBKAC.
This type of analysis is not outside the purpose of the tool. You're making excuses at this point. Do you really think it would be wrong to add that capability in the future?
It's a technical limitation, one that is far from obvious.
THAT IS THE POINT OF THE ARTICLE
If you can't figure that out, don't insult me.
It's not silly at all.
“ Here is a steps vs bmi plot. What do you notice?”
Part of the answer:
“Monkey Shape: The most striking feature of this plot is that the data points are arranged to form the shape of a monkey. This is not a typical scatter plot where you'd expect to see trends or correlations between variables in a statistical sense. Instead, it appears to be a creative visualization where data points are placed to create an image.”
Gemini 2.0 Pro without thinking didn’t see the monkey
I wonder if the conversation context unfairly weighed the new impression towards the previous interpretation.
o1 pro: It appears the data have been deliberately arranged to form a stylized humanoid figure. In terms of actual trends, there is no obvious linear or nonlinear correlation between steps and BMI, and the female–male color split seems fairly uniform across the shape. Visually, BMI spans roughly 15–30, while steps extend up to about 15,000, with no clear clustering or separation by gender. Within reasonable confidence (e.g., an estimated R² near zero if one were to attempt a linear fit), there is no discernible predictive relationship between steps and BMI in this plot.
Literally, because this is why the Datasaurus dozen was created: https://en.wikipedia.org/wiki/Datasaurus_dozen
Metaphorically, because of all the times (including here, on this very article :P) where people comment on the basis of the headline rather than reading a story.
On the other hand, this isn't the bit of human cognition we should be trying to automate, it's the bit we should be using AI to overcome.
Take a screenshot of the scatterplot and feed it into multimodal GPT and it does a fine job at identifying it.
EDIT:
Sorry, as a few people pointed out, I missed the part where the author did feed a PNG into GPT. I kind of jumped to conclusions when it worked fine for me. I still maintain that the article's conclusion ("Your AI Can't See Gorillas") is overly broad, given that I had no trouble getting it to see one.
But I wonder why the author had trouble? My suspicion is that AI got stuck on summary statistics because the previous messages in the chat were all about summary statistics.
The contents of the CSV might be entirely unknown at inference time.
> I asked the model to closely look at the plot, and also uploaded a png of the plot it had generated.
what do you see
ChatGPT said:
This is a scatter plot with the variables "steps" on the x-axis and "bmi" on the y-axis. The data points are colored by "gender" (red for female and blue for male). Interestingly, the arrangement of the points appears to form a drawing resembling a cartoonish figure or character, likely added for artistic or humorous effect.
If you'd like, I can help analyze or discuss the data further!
You said:
Wha the cartoonish figure looks like ?
ChatGPT said:
The cartoonish figure in the plot resembles a stylized gorilla or monkey-like character. It appears to be waving with one hand raised, while the other arm is resting downward. The face is expressive, with distinct eyes, a nose, and a slightly frowning mouth. The overall pose gives it a somewhat playful or cheeky vibe.
OC seemed to think that Claude did that with just the data and not the image of the scatterplot it’s.
Without number of rerurns, the result is as good as random.
OC was saying that the article said that Claude recognized the “artistic” lines of the image from just the scatter plot data.
That isn’t what happened.
The author added a png of the plot to the conversation.
Idk why I need to explain that twice.
I wonder if the author got different results because they had been talking a lot about a data set before showing the image, which possibly predisposed AI to think that it was a normal data set. In any case, I think that "Your Ai Can't See Gorillas" isn't really a valid conclusion.
And yes, the idea that the initial context can sometimes predispose the LLM to consider things in a more narrow manner than a user might otherwise want is definitely well known.
The article says:
> Furthermore, their data analysis capabilities seem to focus much more on quantitative metrics and summary statistics, and less on the visual structure of the data
Again, this seems false - or, at best, misleading. I had no problem getting AI to focus on visual structure of the data without any tricks. A more fair statement would be "If you ask an AI a bunch of questions about summary statistics and then show it a scatterplot with an image, then it might continue to focus on summary statistics". But that's not what the concluding paragraph states, and it's not what the title states, either.
if you didnt know it was there, and took a look at only the text output, the llm would not have found it to tell you its there
The conclusion is quite reasonable and the article was IMO well written. It shares details of an experiment and then provides a thoughtful analysis. I don’t believe the analysis is overly broad.
Not sure whether multimodal embeddings have such a good pattern recognition accuracy in this case, probably most of information goes into attending to plot related features, like its labels and ticks.
Is there a blog post that just focus on the gorilla test that I can share with my team? I’m not even interested in the LLM part
A corollary if this that is my personal pet peeve is attributing everything you can’t explain to “seasonality” , that is such a crutch. If you can’t explain it then just say that. There is a better than not chance it is noise anyway.
Very early in my career, I discovered python's FFT libraries, and thought I was being clever when plugging in satellite data and getting a strong signal.
Until I realised I'd found "years".
Is this a literal thing or figurative thing? Because it should be very easy to see the seasons if you have a few years of data.
I just attribute all the data I don't like to noise :-)
My conclusion so far has been "well they are not doing their job properly".
I assume that's the kinds of jobs LLM's can replace: People you don't want on your payroll anyway
From short thinking or from looking at the graphs I would believe "roughly normal" sounds like wishful thinking to stay in the reassuring bounds of normal distributions. And I believe things would get dangerous once you would start using these assumptions for tests and affirmations.
My short thinking: distributions don't look close to normal on the graphs. Values are probably bounded on one side and almost unbounded on the other (can't go below 0 steps, can go into very high number of steps on 1 day). There are days / people with close to 0 steps and others that might distribute in a sort of normal around a value maybe. Weight and height might be normally distributed in a population but they're correlated and BMI is one divided by the square of the other. I can't compute the resulting distribution but I would doubt that would make for a distribution close to normal.
Ok the LLMs were told to assume both traits were distributed normally, but affirming they look mostly normal is scary to me.
Am I too picky and in real analyses assuming such distributions are "mostly normal" is fine for all practical purposes?
In general, I’ve steered clear of current LLMs for data analysis/description because they seem so highly influenced by choice of prompt and wording. They tend to simply affirm any language I use to describe the data initially.
To be fair, I’ve attended conferences and lab meetings where humans will refer to a any vaguely concave curved distribution as “mostly normal” :P
ChatGPT (4o): Noticed "a pattern"
Le Chat (Mistral): Noticed a "cartoonish figure"
DeepSeek (R1): Completely missed it
Claude: Completely missed it
Gemini 2.0 Flash: Completely missed it
Gemini 2.0 Flash Thinking: Noticed "a monkey"But I guess GPT-4o results are more funny to look at.
They do seem to generally have legs and head, which is an improvement over 4o. Still pretty unimpressive.
>but does not specifically understand the pattern as a gorilla
Maybe it does, how could you tell? Do you really expect an assistant to say "Holy shit, there's a gorilla in your plot!"? The only thing relevant to the request is that the data seems fishy, and it outputs exactly this. Maybe something trained for creative writing, agency, character, and witty remarks (like Claude 3 Opus) would be inclined to do that, and that would be amusing, but that's pretty optional for the presented task.
However, it refuses to cooperate. It's maddening.
As a result, I receive "There is a person at your front door" notifications at all hours of the night.
In truth it’s only mildly annoying and makes me appreciate my cat’s quirkiness more.
Not the end of the world but it does seem like AI gets fixated, like people, and can’t see anything else.
It has no problem identifying our other two dogs as actual dogs who bark and move like dogs.
What’s missing in the terminology is the modality- most often TEXT.
So really we on have Test LLM or Text Reasoning models at the moment.
Your example illustrates the benefits of Multi Modal Reasoning (using multiple modality with multi pass)
Good news - this is coming (I’m working on it). Bad news this massively increases the compute as each pass now has to interact with each modality. Unless the LLM is fully multi modal (Some are) - this now forces the multipass questions to accommodate. The number of extra possible paths massively increases. Hopefully we stumble across a nice solution. But the level of complexity massively increases with each additional modality (text,audio,images, video etc)
“look at the scatter plot again” is anthropomorphizing the llm and expecting it to infer a fairly odd intent.
would queries like, “does the scatter plot visualization look like any real world objects?” may have produced a result the author was fishing for.
if it were the opposite situation and you were trying to answer “real” questions and the llm was suggesting, “the data is visualized looks like notorious big” we’d all be here laughing at a different post about the dumb llm.
The gorilla is just an extreme example of that.
Albeit perhaps an unfair example when applied to AI.
In the original experiment with humans, the assumption seemed to be that the gorilla is fundamentally easy to see. Therefore if you look at the graph to try to find patterns in it, you ought to notice the gorilla. If you don’t notice it, you might also fail to notice other obvious patterns that would be more likely to occur in real data.
Even for humans, that assumption might be incorrect. To some extent, failing to notice the gorilla might just be demonstrating a quirk in our brains’ visual processing. If we expect data, we see data, no matter how obvious the gorilla might be. Failing to notice the gorilla doesn’t necessarily mean that we’d also fail to notice the sorts of patterns or flaws that appear in real data. But on the other hand, people do often fail to notice ‘obvious’ patterns in real data. To distinguish the two effects, you’d want a larger experiment with more types of ‘obvious’ flaws than just gorillas.
For AI, those concerns are the same but magnified. On one hand, vision models are so alien that it’s entirely plausible they can notice patterns reliably despite not seeing the gorilla. On the other hand, vision models are so unreliable that it’s also plausible they can’t notice patterns in graphs well at all.
In any case, for both humans and AI, it’s interesting what these examples reveal about their visual processing, which is in both cases something of a black box. That makes the gorilla experiment worth talking about regardless of what lessons it does or doesn’t hold for real data analysis.
ChatGPT:
> It looks like the scatter plot unintentionally formed an artistic pattern rather than a meaningful representation of the data.
Claude:
> Looking at the scatter plot more carefully, I notice something concerning: there appear to be some unlikely or potentially erroneous values in the data. Let me analyze this in more detail.
> Ah, now I see something very striking that I missed in my previous analysis - there appears to be a clear pattern in the data points that looks artificial. The data points form distinct curves and lines across the plot, which is highly unusual for what should be natural, continuous biological measurements.
Given the context of asking for quantitative analysis and their general beaten-into-submission attitude where they defer to you, eg, your assertion this is a real dataset… I’m not sure what conclusion we’re supposed to draw.
That if you lie to the AI, it’ll believe you…?
Neither was prompted that this is potentially adversarial data — and AI don’t generally infer social context very well. (A similar effect occurs with math tests.)
It's how I do it. Why not our pet llm?
Maybe a different choice of illustration would result in a more apt description.
I don’t even like AI and I still will tell you this whole premise is bullshit.
ChatGPT got
> It looks like the scatter plot unintentionally formed an artistic pattern rather than a meaningful representation of the data.
Claude drew a scatter plot with points that are so fat that it doesn’t look like a gorilla. It looks like two graffiti artists fighting over drawing space.
It’s a resolution problem.
What happens if you give Claude the picture ChatGPT generated?
I feel we also need models to be able to have a " Wait, What?" moment
Another subtle joke about chip design and layout strikes again.
"AI can't see gorilla because wokification"?
/s
Edit: Adding /s Thought "wokification" already signalled that.
You just assumed it’s similar to the Google situation a few years back where they banned their classifier from classifying images as gorillas. It isn’t.
"Google promised a fix after its photo-categorization software labeled black people as gorillas in 2015. More than two years later, it hasn't found one."
Companies do seem to have developed greater sensitivity to blind spots with diversity in their datasets, so Parent might not be totally out of line to bring it up.
IBM offloaded their domestic surveillance and facial recognition services following the BLM protests when interest by law enforcement sparked concerns of racial profiling and abuse due in part to low accuracy in higher-melanin subjects, and Apple face unlock famously couldn't tell Asians apart.
It's not outlandish to assume that there's been some special effort made to ensure that datasets and evaluation in newer models don't ignite any more PR threads. That's not claiming Google's classification models have anything to do with OpenAI's multimodal models, just that we know that until relatively recently, models from more than one major US company struggled to correctly identify some individuals as individuals.
Also the prompt matters. To a human, literally everything they see and experience is "the prompt", so to speak. A constant barrage of inputs.
To the AI, it's just the prompt and the text it generates.