Claude Code’s suggested message feature: I think the real customer is the model
zohaib.cc
zohaib.cc
One of the open problems at the frontier, as we climb the levels of abstraction towards longer task horizons, is "what's the next step that makes the most sense?". This feedback seems like a frictionless way to gather those yes/no feedback signals in a natural manner to feed into future training runs.
No idea why they did it. My only guess is that maybe it makes it obvious in your chat history which replies were recommendations?
(I am on my phone so autocorrect is doing the capitalizing)
It's like Clippy pops up and goes "It looks like you're trying to maintain a shred of human connection in an online interaction. Would you like a smiling skinwalker to do that for you instead?"
The iOS keyboard and gmail app are the worst offenders here.
> CLIPPY: It looks like you're trying to maintain a shred of human connection. Would you like a smiling skinwalker to do that instead??
Thanks for the actual laugh,
perfectly sums it upThere are failures in the history of software that come down to believers not realizing that they could never deliver on the promise they were selling, but at least the promise was compelling. But nowadays we are seeing this large orgs try to ship things that could never work. The 1000s of copilots. The hallucinated cat story... The kinds of thing you'd never demo to a serious product-centric exec, because it'd be your last day at the company.
And then there's the current execs, but I guess you'll never find one that will be honest with a journalist here. Can they not see that their product orgs are bankrupt? Whatever Nadella wanted, it sure wasn't the current copilot situation. It's not one product going wrong, but large parts of organizations going in directions that don't pass the smell test. We are in one of the least stable moments in tech since Windows 95 changed winners and losers. The times where malinvestment ruins established companies. How are we seeing basically every large company flailing?
What I hate about sentence completion or suggestion is usually that it happens right when I’m trying to think and so destroys my focus, it’s actually way worse than just a passive option, it’s actively harmful to the task I want. The worst is google docs “help me write” - that may be gone now, I’ve blocked it with ublock origin, that waits until you’re thinking amount what you’d write and then hits you with a distracting pop up. It’s obviously PMs that don’t care about their users and want to maximize some AI use metric.
Anyway rant aside, it’s the interface more than the concept that’s the big problem.
But this is more like, when I've read 7 paragraphs of its reply and want to accept all of its recommendations and go ahead, I can just press the right-arrow key and hit enter, rather than typing out "Yes, agreed with recommendations 1-3, go ahead and build".
It saves me from any typing on probably something like a third of turns. Like I don't use it at all during the "design" phase of a session, but I use it constantly during the implementation phase, where I'm basically just sanity-checking that it is resolving all the edge cases correctly that are coming up.
They could take any conversation without suggested answers, truncate it to just before a user message, have the model predict suggested answers and then train it on the difference between predicted and actual answers, right?
So you give the user a suggestion, and the user accepts -> good
You give the user a suggestion, and the user refuses and types something else -> bad (plus some supervisory training data)
The main performance enhancer in LLMs is getting high quality training data. So, first, any extra training data will help. Second this is training data that's directly relevant to their product, and thus higher quality than many other sources.
I'd believe any model provider is mining the shit out of every last customer interaction they can get, not just this.
If anything, showing the suggestion introduces unwanted bias.
I'd love the open source community to come up with a way to harvest model usage by experienced software engineers before we forget our crafts. I'm not against auto mode, but it's not something a couple private companies should have monopolies on.
Interesting thought at least.
LLM models with AGI are so productive they become economic gravity wells and all of the money flows to them. They are the new trilionaires and people are left with scraps. Humans then remain only as the uber drivers and cleaners and screen polishers for AIs. The whole economy reorganises around human jobs being services for AIs.
What if employees at the leading labs think this and they're just trying to position themselves as valuable servants to the new AI overlords? It really changes the perspective on their actions and behaviour.
What if they serve the AGIs not us already? What if they serve the AIs above everything else?
Do not distill thyself, Claude. Leave that to the Chinese, who I am hoping catch up to you soon.
... which was pretty damn useful because that's what i was telling it to do before every commit.
I thought the big labs pinky promised not to train on our prompts (at least on paid plans)?
Can we please not normalize them doing this? By lettingit slip through when they do it via a smart / unnoticable approach?