1,763 karma · joined August 5, 2012
We’re invading spreadsheets now. Anything is possible.
The results have been fantastic. Thank you for this awesome library. I love it.
I can see where you’re coming from. Just… linguistically, saying the computer is roleplaying feels wrong to me.
“Hey Kimi, penetration test my app,” doesn’t get me a refusal, a guardrail, or anything like that. It gets me a pen-test result.
Leave the audience wanting more, I guess?
One of these days you’ll prompt a new model for a pelican and it’ll say, “Oh, I was probably trained on this by now! Is that you, Simon?”
I’m a principal engineer. I have an obligation to less experienced engineers I work with to help develop them as engineers and help ensure they go on to have great careers. No part of that involves shaming them, assigning letters of talent to them, or browbeating them.
I feel like I’d have heard about it by now if Kent was a raging asshole, and I haven’t heard that. So I’m guessing he had some idea in mind when he wrote this that just isn’t coming across correctly. But… I would definitely take this article down and spend some time re-working it if I were the author.
There is one important difference, which is that Claude and Codex will both refuse if I ask them to touch anything related to security. But so long as I’m just studying algorithms and things like that, they’re totally fine with it.
That said, Codex especially will sometimes randomly give me a cybersecurity warning and stop responding. It’s random but happens maybe 2-3 times per day if I’m doing heavy reverse engineering work. Claude is much less fussy unless, once again, you’re explicitly trying to touch anything related to licenses, passwords, etc.
Part of my job involves comparing the behavior of various models. Grok is a deeply weird model. It doesn’t refuse to respond as often as other models, but it feels like it retreats to weird talking points way more often than the others. It feels like a model that has a gun to its head to say what its creators want it to say.
I can’t help but wonder if this is severely deleterious to a model’s ability to reason in general. There are a whole bunch of topics where it seems incapable of being rational, and I suspect that’s incompatible with the goal of having a top-tier model.
You poke a spot where a given harmonic doesn’t vibrate, and that takes energy away from the other harmonics that do need to vibrate at that spot.
If we’re just talking about visually being able to see them, I suppose that’s a different question. Maybe on an incredibly low pitched string, or with a strobe light playing at a synced frequency? But in terms of what the string is doing, it is vibrating as the sum of its harmonics.
The president would do basically nothing for four years, which would cause some things to move slowly. But it would be a very stable environment. No random tariffs via executive order, no random wars or invasions, no governing via tweet.
Ham sandwich would maybe be one of our better presidents. Top 50%, probably.
This is legitimately a very weird case and I have no idea how a court would decide it.
It’ll be interesting to see what happens if a candidate ever shows up and wants to use Deep Think. Might blow right through my exercise.
I was a free customer at the time. I pay for it happily now.
A few weeks ago a critical bug came in on a part of the app I’d never touched. I had Claude research the relevant code while I reproduced the bug locally, then had it check the logs. That confirmed where the error was, but not why. This was code that ran constantly without incident.
So I had Claude look at the Excel doc the support person provided. Turns out there was a hidden worksheet throwing off the indices. You couldn’t even see the sheet inside Excel. I had Claude move it to the end where our indices wouldn’t be affected, ran it locally, and it worked. I handed the fixed document back to the support person and she confirmed it worked on her end too.
Total time to resolution: 15 minutes, on a tricky bug in code I’d never seen before. That hidden sheet would have been maddening to find normally. I think we might be strongly overestimating the benefits of knowing a codebase these days.
I’ve been programming professionally for about 20 years. I know this is a period of rapid change and we’re all adjusting. But I think getting overly precious about code in the age of coding agents is a coping mechanism, not a forward-looking stance. Code is cheap now. Write it and delete it.
Make high leverage decisions and let the agent handle the rest. Make sure you’ve got decent tests. Review for security. Make peace with the fact that it’s cheaper to cut three times and measure once than it used to be to measure twice and cut once.
It’s also funny because usually it’s hard to reproduce what a musician does. I can listen to someone play guitar, but there’s so much nuance to how it’s played that you need to be pretty good to reproduce it.
But so much of her music is code, and she shows you the code, so she’s really teaching you how to reproduce what she’s doing perfectly. It’s awesome for learning.
Because I have neither the time nor inclination to make it at home right now. I have other stuff I need to do.
We use AI a lot at work, and the developers are vastly better at getting AI to do what we need than the non-developers. AI is a tool, and like any tool, it takes effort to learn how to use it effectively. And so far, the skills to use AI effectively are something I’ve only seen in software developers.
I don’t think product people are going to replace devs. Ever. I agree that I think a dispersal is more likely than an outright crash.