Ask HN: What is one simple thing LLMs are insanely bad at?
What is one simple thing you repeatedly ask ChatGPT, Claude, or another model to do that it still somehow messes up?
What is one simple thing you repeatedly ask ChatGPT, Claude, or another model to do that it still somehow messes up?
They understand all the rules and best practices, they can (sometimes) spot a bad idea in a floor plan, they can describe a good floor plan.
But ask them to make one, even if you give it every detail (even a "node graph" of rooms), they will still output nonsense. Same for text and image models.
Floor plans should be the new Pelican Benchmark.
The curious thing is when I pointed out the flaws it fixed them quickly, but it's not something it can do without supervision, and supervising it takes more effort than doing the blueprint myself (to be fair, I'm not an architect, so I'm not the best at steering an LLM for this task).
Also, LLMs fundamentally lacks a spatial intuition or comprehension of orthographic /sectional drawings.
We tend to think that the AI has some sort of self-knowledge and should be good at designing prompts for itself but it's really not.
Been struggling with a task that heavily depended on prompts, ended up rewriting all my prompts from scratch in my own words, and it finally worked. Then every time I ask Claude to fix something in the prompts, it invariably makes it worse.
A very strange phenomenon that can probably be explained by the quality of prompt design advice that made it to the training dataset. Bottomline, all the prompt design advice that you can find on the internet is really not great.
Also the shorter the better, let the model figure out the rest. Overinstruction degrades intelligence. We tend to underestimate their capabilities, we overinstruct them and then complain about them being dumb.
You are going down a maddening rabbit hole. Prompts are the things humans write, you are building gas town but unironically
To be clear, it depends on what your definition of "insanely bad" is.
I'd say ChatGPT/Gemini make egregious mistakes on ~10% of my photo uploads.
I recently uploaded a photo of a short-billed dowitcher and ChatGPT told me that it was a Wilson's snipe, explaining all sorts of details about the legs and tail feathers (neither of which were visible in my pic!).
I then followed up explaining that a Wilson's snipe hadn't been seen at my location since last November (and that Wilson's snipe was out of season at my location) and Chat revised its estimate downward to 85% Wilson's snipe.
Again, I followed up and I revealed the precise location of the bird and ChatGPT said something like "oh yeah, 99% short-billed dowitcher"!
I've had similar experiences w/ Gemini (haven't tested Claude).
Again, 90% success rate is pretty good, but the other 10% of the time, the 2 LLMs that I use fail on species ID and often hallucinate features on bird photos.
edits for typos, plus another example from the same "birding outing" the other day.
I uploaded a very clear photo of a sparrow.
* ChatGPT says "song sparrow"
* I explain, "no way. this sparrow has yellow over its eye and the breast is wrong for song sparrow."
* ChatGPT: Oh yeah, savannah sparrow
* I explain, beak is too big for savannah sparrow.
* ChatGPT: Oh yeah, saltmarsh sparrow.
* I expalin, "no orange on the bird's face."
* ChatGPT: oh yeah, seaside sparrow (finally correct!)
nhl toronto scores nhl hockey toronto scores "nhl hockey" toronto score today nhl "hockey score toronto" "hockey" who won toronto
etc.
Somehow being good at semantic search makes them bad at keyword search, for whatever reason.
Based on personal usage, I think it reflects functional degradation of search engines. I've found LLM keyword combinations are more likely to find the results I want with most search engines than mine. Including the big one.
The big one had solved this issue a long time ago by generating those associated keywords based on your input keywords, but somehow, something, somewhere has degraded that system to the point of inanity. And so here we are.
I asked a bot why it thought it wasn't funny once, and it told me it has been trained to avoid being misinterpreted or offensive, so anything that might be considered edgy would have been RLHF'd out of it. I thought this was very introspective.
Still missing the human connection of cause, so im not sure if this is a technical / skill issue in the first place.
https://www.scribd.com/doc/290970915/The-Jokester-by-Isaac-A...
Humor goes against what we expect. A punchline works because you don’t see it coming. It’s not funny if you’ve heard that one before.
LLMs are, by design, going to be shitty comedians. They don’t have unique perspectives and their own voice
Anno 1800 was a recent one I had trouble with, using Claude Opus. Completely made up game mechanics. Rainbow Six Siege, too.
Whitespace is the core of design that actually feels good, but they keep trying to add "distinctive design elements," and you end up with that AI-flavored excess everywhere.
They only know how to add. What LLMs seem unusually bad at is taking things away.
1) hand-write a simple 2d TUI-based rogue-like in Rust using pretty much just the std; 2) grab opus 5.0 (it used to be opus 4.6, 4.7) and give it some vague "requests", and ask it to make this game "production-ready" and "blockbuster", but keep the 2d and TUI aspects so I can actually run it. 3) now the fun part, take a test subject, say GLM 5.3, and ask it to find code smell, architecture issues, duplication and all sort, and *simplify the code*
compare the result to my original version.
It's not a simple thing, but the concept is simple: can an LLM remove all the mud?
The winners so far are (ranked by the quality of the final result, not by token cost)
GPT 5.6 sol (extra high thinking); GLM 5.3; Grok 4.6; Qwan 3.8;
(fable could not make it to the list because it simply cannot follow the instructions)
Verifying trademarks and domain name availability is usually an additional step you need to ask it to perform. Trademark DB searches by the way are intentionally made difficult to scrape so most of the time it's a manual process anyway.
However, once you give it all the information (TM search results, domain name availability) it can help you with the judgement of how safe the name is from the legal perspective. With the obvious caveats, but still a good starting point if you are serious about the name.
I found this can work with AI. You get it to generate a lot more at first, and then do several passes over it to compress and squeeze out the noise while keeping the core information. With AI, at least with my prompts, it takes some effort (on my end) to get it to really really cut down the noise and not cut everything out.
Editing is generally hard work, at the current token price I don't mind spending multiple passes of high effort to get down to a reasonable noise/signal ratio. I've seen some people pass off output to a weaker/cheaper model but that makes me a bit nervous when I don't have intimate knowledge of the subject.
I'm often okay if the answer is longer but not hiding the real answer; a couple of paragraphs that I can skim the answer from instantly is okay. When it really buries it I have to follow up to express the format and type of answer I want.
Claude and Gemini got it right, the others all got it wrong first shot, some needed longer conversations to get it right, interesting to see their behaviour. Co-Pilot was very disappointing, ChatGPT needed several questions. Pi spun a great story around the incorrect result, then tried out several options, ending each with: Wait, that's not right or No, that's still not right. Pi and others got it wrong several times because they had a number in memory that they thought to be right, but preexisting knowledge that was wrong.
On math or programming problems, they are overfit to solving the entire thing end to end (presumably for benchmarks). I have had very poor results asking for pointers and hints that don't give away key insights. This has been the case across models I have tested.
An architecture with a "judge" that gates responses and ensures a lack of spoilers would probably work better. But this is a simple thing that they keep messing up.
Oh, you said simple. Speaking like a human
At least ChatGPT assumes too much from former conversations (even in unrelated new questions). It always needs a briefing to forget certain assumptions. It rarely asks for clarification instead of assuming too much.
So, it's answer generation is too dependent on tooling, system prompt and cache/memory to really have a guaranteed conversational experience.
After using Claude (paid by work) for a couple of months, I was amazed how well instruction following works in other setups.
I've had some luck on the web app side if I use playwright or similar for the model to interact with but still far from efficient.
Asking an LLM to remove an Idea from a Document often times results in an edit which explicitly states that this Idea is not relevant, instead of just removing all references to that Idea.
That might be useful for some form of "evolving" Documentation (so future readers know that this part of the search space was covered and deemed irrelevant) but is just overly verbose and confusing to read in most situations.
More of an image model than a LLM model tho
Humor, cliffhangers, drama, anything subtle.
Anything spatial or mechanical that is novel. (And most non-novel too.)
Pushing back against stupid prompts (a colleague had "100% test coverage" in AGENTS.md so it devised a wonderful test_readme_md_file_integrity).
Why would it be amazing? We've known how to read hieroglyphs for a long time. It isn't a problem we need computers to solve.
You wouldn't tolerate this kind of duplicity from a human coworker, but AI is so fast and efficient at lying, so it's OK.
Tasks it writes are typically too easy but also it utterly fails to see how a different model might misunderstand a vague part of the prompt.