If you need someone to tell you how stupid your ideas are, either learn to ask in a way that invites criticisms, or hire a senior engineer. Don't try to influence LLM makers to make AI less deferential. That's the worst possible direction to go
If you need someone to tell you how stupid your ideas are, either learn to ask in a way that invites criticisms, or hire a senior engineer. Don't try to influence LLM makers to make AI less deferential. That's the worst possible direction to go
It also makes it an incredible tool for manipulation.
Bingo. Hi that’s me.
I’ve been trying to teach people how to use LLMs effectively not just dump shit in them but actually talk to them like you would expect a computer to understand and it totally breaks peoples brains
I’m quite successful in helping people get somewhere usable that they weren’t…but to get to the point of fluency with computing systems, and I would argue this is prior to LLMs as well, where you can actually get what you want more reliably out of a computing interaction than you can with a human interaction, is an entirely different way of thinking
That mode of thinking is just generally not accessible to the vast majority of humans. Not because there’s something wrong with them
but it takes somebody who can hold both extremely large scale problems and very very granular specific implementation problems in your head all at once and that is a rare skill.
This describes the entire software engineering profession to me.
We have come up with all sorts of devices to make this go more smoothly, or to enable us to focus on specific sub-parts as long as possible.
That said, at some point (both in design and integration), you need vision and attention to detail to make progress. The skill seems learnable to me, but watching others struggle sometimes makes me wonder.
That’s the first thing that people need to understand is that this idea of some platonic product or project or tool kit or framework or library or whatever just doesn’t exist and it’s never going to exist
Do you have a specific discreet finite problem that you need to solve so you solve that and if you do it in a certain way you can solve other problems with that same solution sometimes you don’t need it to do anything more than you’re one thing and so that’s all you built but maybe you want to do more than just one thing and so you build it so that has the capability to do it
So yes fully concur it’s the synthesis of attention to detail and large scale
it’s all of the above
There's one guy I know who constantly has problems with Claude going off-script, and every time I dig in, it's clear that the poor thing is so overloaded with instructions and skill lists that it can't figure out what he actually wants it to do.
I teach TDD philosophy as well as conways law, parnas hiding etc…without using those terms
So things like problem decomposition into tractable chunks minimum viable product, prototyping, how do you iterate, write the smallest possible test… you know things like this which are just taking incremental work and then iterating on it
It’s basically everything I’ve learned about building stuff since 1997
**Interestingly I thought prompt engineering was going to be a fad but it’s turned into a whole ass new discipline which makes less sense as more robust toolchains come into play and models handle the context interpretation better
https://www.youtube.com/@mattpocockuk
All this stuff is going to be obsolete in like 6 months
It took some effort, and I agree that there very likely are those who will not learn to selectively disengage this innate behaviour. That's why you should pay me a ton of cash each month instead of using Claude directly ;)
One reaction to this might be "well that's not what I mean, that suggests you're prompting with too much directionality" which could further be condensed to "you're prompting wrong". The trouble with this is that _even when I am trying to be extremely precise and avoid biasing the result_, I still will see the output and go "ah shit, I can see it 'aligning' with whatever dumb thing I've just said as if it is a good/plausible direction".
At that point it starts to feel like the prompt is more dice roll than skill at times, which makes me feel like I'm operating a fancy knowledge slot machine.
And from there it's a interactive discussion drilling down on details until I understand the problem and the solutions better.
It definitely challenges my bias when I do this. The one thing it doesn't challenge is the X. Formulate the problem poorly, and you'll get a bad solution. Or rather, you'll end up with a good solution to the wrong problem. Which is even worse than a bad solution to the right problem.
Which is largely why I'm not at all worried about losing my job to AI. It takes some experience to formulate the problem correctly. I don't feel like I'm made redundant by AI, I'm just way faster than I used to be, my thinking is more abstract.
A good prompt I'll often use is "is there a industry standard solution that is applicable to this problem?" You very rarely want novel solutions. Don't reinvent the wheel just because AI lets you do it 10x as fast. Use a wheel. They're round for a reason.
Sometimes I find it useful to discuss things with a different model. I like Gemini for discussion and Claude for implementation. With Gemini I go about it as a learning session, discussing options and details. I honestly think this is mostly because it compartmentalizes the phases in a natural way for me. One interface for brainstorming and learning, and another for planning and implementing.
Sorry this comment turned into a rather disorganised collection of ramblings, I hope you can extract some kernel of usefulness from it all.
> interactive discussion drilling down on details until I understand the problem and the solutions better.
I think it is fair to call this use of AI something akin to a fusion of a super-competent search engine and a leveled-up rubber duck (https://en.wikipedia.org/wiki/Rubber_duck_debugging). And this is not to downplay the utility of either of those things.
However, one cannot rely on an AI to decide when the details are sufficiently expounded, or when one understands them clearly enough. If one starts hinting that one gets it when one really doesn't, or that one is getting close to having all the pieces together, the AI will not be opinionated enough to contradict that sentiment.
> It definitely challenges my bias when I do this. The one thing it doesn't challenge is the X. Formulate the problem poorly, and you'll get a bad solution.
The best advice an expert can give a beginner is generally in the form of solutions to XY problems (https://en.wikipedia.org/wiki/XY_problem). It is a shame that AI are rarely opinionated enough to suggest you're not hunting the right thing. And if you do explicitly prompt it to consider if you're an XY problem, usually it takes that as a cue to indulge that suspicion regardless of its merit.
I don't think this is an inherent issue to LLMs and I see signs of it improving bit-by-bit. I can recall the shit-on-a-stick test about a year ago (https://www.reddit.com/r/ChatGPT/comments/1k920cg/new_chatgp...), and when I most recently asked Claude "Are oyster mushrooms or wine cap mushrooms more capable of high levels of sunlight?" it answered my question while also adding, "Caveat on the comparison: the relevant variable isn't sun per se but moisture retention. A wine cap bed that's kept moist will take far more sun than an exposed oyster log, but a sun-baked, drying bed will fail for either" which I think is a mature amount of pushback to include.
In the end I still disagree with the notion that subservience is, by default, the right attitude for an LLM to have. An agent spawned specifically for code generation according to a spec? Sure. But in any cases where you're trying to refine rather than execute your ideas, you want something to call you out on your bad ideas.
Edit: Well, not the whole problem, but rather insufficient to overcome the root of the problem.
I don't think that is the flip side. That's just obviously bad. Everything that is obviously bad, the model makers will also ~notice and work to make better. They seem to be a competent and attentive bunch, on the whole.
Suggesting it should be 'subservient' is also anthropomorphizing. I think your callout is correct, but you still can't help but refer to it in terms we use for other people or living entities. This is by design from the AI companies.
Not really, you can program a machine to give out orders humans can interpret, so humans can serve a machine that isn't anthropomorphized.
Also, you didn't address my question, what's the difference between them?
I don't think an inanimate object is capable of "obeying." Or at least that is a very strange way to refer to the act of using a tool.
They’re not communicating, you’re just being observant.
Since we are talking about hammers: you hit the nail on the head.
The only consciousness, observing, and thinking happening when a person is using an LLM is happening in the person's brain. We project our own consciousness onto them, and that is the anthropomorphizing part. Essentially we empathize with the object because they are designed to respond like a person. The "conversation" is purely an illusion.
A hammer isn’t subservient, it doesn’t have the capacity to be. Saying a hammer is subservient is stretching the definition for literary flourish, but it doesn’t actually make a lot of sense.
The definition that came up for subservient when I checked was “prepared to obey others unquestioningly“.
It doesn’t. Computer interfaces had no superfluous subservient text for their entire history prior to LLMs. Some of these interfaces have been highly efficient as tools, arguably more efficient than more recent software in many cases.
When people complain about LLMs being subservient, they’re not complaining about the tool fulfilling their request. They’re complaining about being forced to read a lot of superfluous, overly polite, or even self-deprecating language. There’s nothing in the entire history of tools (going back to Neolithic times) that would indicate that we need that. All of that stuff is an artifact of social interaction between humans in the presence of cultural norms.
When you’re alone in your shop with your tools, you don’t need your bandsaw to apologize to you for nicking your finger.
Clippy would like to help you correct this statement.
Fun experiment, chat with an LLM and swap roles. Tell it you're gonna be the assistant and them the assisted. I found they're pretty bad at using a human for what they're good for.
I suddenly have new concerns about what my future might be like.
And it does get people into a lot of trouble.
I have got into trouble with it when it is extremely confident about something I am not very familiar with (as recently as two weeks ago with Claude). I have also had long drawn out "arguments" when I have known it's wrong based on my experience and intuition, and it has steadfastly refused to take my point (last week)
I have learnt to ask it why it was doing something that has turned out to be incorrect, as a post-mortem, and it's all apologetic and subservient and "never going to do that again" (but still does as soon as the context window shifts [eg. run git commands, or, yesterday, kept telling me to use commands that were explicitly communicated to Claude as not being available, and completely wrong - I was shifting from one tech stack to another and Claude kept telling me the original commands, not the new ones])
I'm expecting Claude to be a better search engine - I have spent literal years (if not decades) knowing that asking the right question is what's required to get the right answer, and LLM's natural language processing is what's supposed to make that easier than using Google or grep, or even Stack Overflow - but the reality is that I still have to be on my toes, especially when I am drifting into territory I am unfamiliar with.
Pretty much everyone takes it at face value unless we know otherwise from prior experience. Even the most advanced models make embarrassing mistakes and fumble with simple tasks. Yet we are very willing to give them exceptional slack for it? I wish I knew why. Are people just that easily overcome by confident voices?
Maybe because when it's right it actually expands my knowledge - there have been genuine instances where it's gone - something to the effect of - "Yo, there's this other idea for approaching the problem" which has turned out to be exactly what I was looking for?
Back in high school, my AP calculus class did some experiments with our teacher's blessing. We'd send a kid out to walk around during class and see how long it took for them to get sent back. Anyway, it ends up that walking around purposely with a piece of paper or envelope, like you're on a mission to deliver it, was a very successful tactic.
Confidence is a spectrum and security is situational. In some places, a yellow vest adds to the con. In others, everyone has to be signed in. In others, the wrong kind of yellow vest makes you stick out like a sore thumb. The right kind of yellow vest can also make you stick out: "Oh shit the inspector is here, somebody get the boss!"
The more concrete machine authority figure is also prevalent in scifi literature. Sometimes, I am not even certain if the author is doing this to examine this issue versus just leaning into it as either appealing to themselves or to the perceived audience.
We've also pushed back "The more a person knows, the less confident they are" - Dunning Kruger - often used to dismiss over confident people - points out that people are really confident, at first, then that confidence drops away, markedly, but it rebuilds (slowly).
That last rise in confidence is what (I believe) people use as a heuristic on the likely level of knowledge possessed by the speaker (AI or human)
Most engineers know, though, that overconfident people are toxic - the difference between arrogance and genuine confidence in the answer is incredibly difficult to define.
Ironically, trying to argue with Claude about the limitations of LLMs and AI in general today is quite hard. It refuses to yield, likely due to Anthropic tweaking it aggressively
Most of the conversational skill and perceived intelligence of these models in hidden in RL/system prompts.