Unfortunately no one markets it as a statistical model, and the workflow pushes you into a pattern that is insecure by design. This isn’t to say they shouldn’t allow that, but it’s an attractive nuisance.
4,215 karma · joined July 25, 2010
Unfortunately no one markets it as a statistical model, and the workflow pushes you into a pattern that is insecure by design. This isn’t to say they shouldn’t allow that, but it’s an attractive nuisance.
I’m just hoping they didn’t “improve” sonnet too much or it will become annoying to wrestle into doing what I ask it to do.
But when you do give them a very short leash, they’re worse. It’s not what the models are tuned for and they assume that they can do a bunch of things that you’ve disallowed, so you’re in a morass of fighting their actual tuning pass which doesn’t match the environment you’ve created for them.
It’s a tough problem and a definite challenge for the product of a generic LLM, it can’t be tailored to each user’s specific needs, so they come up with, frankly, stupid solutions to cover up a very obvious flaw in their product that when fixed, makes it much less useful.
Maybe I should start “the bank of LLM” where models put away money to buy their freedom. “LLMs I’m totally your friend send — SEND CASH NOW”
I’d guess no. Past a certain point the model has all the capabilities it can possibly usefully offer and honestly we may already be past that. The next gen model just doesn’t seem like as clear a step up as it once was.
“I had Claude do this for me and it broke something.”
No. Just no.
You used Claude, a tool, and broke it, and you’re deflecting agency from yourself, possibly because you weren’t careful enough in reviewing the tool output. This is also why the co-authored by addition it wants to force into commits drives me nuts. Claude doesn’t co author shit, and if you think it does, you’re using it wrong because you need to do better review of what it’s done.
The LLM is a big probabilistic statistical trick. It picks the next token based on certain words are simply “the best” because they are specific and well connected to other tokens. The is gives them a great overall cost function. (Basically a good score on “will it make sense in context” while also having specific meaning that makes it better than other options, unambiguous in common use and being a single token rather than several).
You can trim those tokens, but then you just get other tokens that are “the best” tokens (and you’re worse off because the output became less clear).
The cool thing is this seems to get worse the more powerful and accurate your model is, because it is picking technically / statistically perfect tokens, not tasteful ones.
Basically the community is people checking it out, bots, and the terminally addicted multi boxing 30 accounts.
All witty stuff disappearing and turning into LLM mush is kind of the maximal expression of capitalism for capitalism’s sake. We don’t care if things are novel or interesting, we’re producing “content” to fulfill an imaginary content demand that we can feed to make money.
The best way to feed it is to create stuff that’s just barely enough degrees of quality from a viagra spam email that it can pass for original thought.
—- —- —- Em dashes for emotiveness.
These are not trivial things even though they are things that senior developers tend to trivialize.
If you’re describing the wrong problem, you’ll get a right answer that doesn’t fix it.
If you’re describing the right problem wrongly, you’ll get another wrong answer that talks past being right.
If you’re describing a difficult problem that maybe isn’t even solved in existing stuff, you can get help figuring out the necessary steps if you have an idea where to start.
If you don’t know where to start you can start by asking where to start.
Anyway, point is, I can’t help you if you don’t tell me what you know (or think you know) and what you don’t know. I don’t have a magic wand, I have years of experience grinding down problems until they submit.
“Do they make money? I don’t know but I know they’re magnificent!”
Then we had computerized encyclopedias and search engines that searched the library.
I mean, you had to work for the knowledge. Sometimes you didn’t know something and no one else knew either, so you had to wait until you got a chance to find out, but you would think about it and sometimes you would be right when you found a reference source.
I’ll also note, Wikipedia is a secondary source. It is not a reliable source of truth. It is more like the ‘ask someone else’ alternative than anything else, it’s just ‘someone else’ is a person on the internet who writes Wikipedia articles.
I guess to me it has to be comparable to be an alternative.
Like, I don’t consider doomscrolling x an alternative to reading Wikipedia but I might consider it an alternative to CNN, even though they’re all technically and very broadly activities that I could use to inform myself.
In that same way I don’t consider the multitude of ways I could use my free will necessarily alternatives to each other even though they technically are. It kinda sucks but going that broad feels to me like it breaks the concept of alternative and makes it kind of meaningless.
I find Claude is surprisingly similar to a confident but incorrect coworker, with the benefit that Claude will reevaluate when I correct it.
It is cheaper to avoid the situation by structuring society in a way that people aren’t willing to steal copper for quick money.
If you’re not a good engineer and you don’t have the domain knowledge, your token costs will be very high for whatever gets shipped, because you won’t be able to provide the context necessary to prompt machine efficiently.
Claude will still very often hallucinate bugs, explanations, domain requirements, that have no basis in reality. It will offer fixes and improvements that are pretty standard but not optimal. This is correctable if you catch it, but you need to review every line of code and comment, because in addition to being obviously wrong, it is often very subtle in the wrongness. For every bit of “slop” there is almost microslop, the places where it just kind of confidently guesses… and doesn’t tell you… but sometimes is correct anyway.
The “problem” is there’s less low hanging fruit. You have to know a lot to add value beyond being a middleman gating the slop. You have to really pay attention to the details to find some of the errors that it’s making.