Also see: Oracle
2,248 karma · joined August 12, 2017
Also see: Oracle
A toy/simplified version lives here: https://github.com/chaboud/goulash
Consideration of mutation rate (a sort of temporal Shannon-ish coding/ordering) lives in there (with some RoPE-friendly structuring). Note: That was a vacation project, not the day job, but similar principles apply even with larger models.
I'm shocked! Shocked, I say!
Sometimes you just want something that works with all the things. That's pandas. But, like the bamboo eaters, it wont be long...
Goodhart's law strikes again.
https://en.wikipedia.org/wiki/Goodhart%27s_law
However, what is meaningful is whether something is able to create usefully adjacent output, like "let's make Minecraft, but with marching cubes, subdivision surfaces, and global illumination... and behaviorally accurate pandas..." (or something like that).
I have an 11 year old, and most of his game ideas are adjacent to other games he's played. He can make those now, or, at least, enough that he can see where it works and where it doesn't.
Compared to a few years ago, that's pretty cool.
By contrast, Claude Code's bias to make assumptions of reasonableness about underlying systems has proven to be immensely frustrating over the last month or two, both personally and at work. I've wasted days on "that was my mistake. I've been reporting numbers on the old architecture because I hadn't enabled the new one in the config" both at work and home. It's immensely frustrating.
But here we are. Wrestling with energetic idiots in model form, wrangled by over-specific harnesses that struggle to stay off of deranged side-quests.
What a time to be alive!
(Note: there's still a huge grammar penalty, so, ugh do think small.)
You are building to solve their problems or open up their capabilities, and if they knew how to do it, they would have already. So listen to your customers for the "why", but use your own judgment for the "what" and the "how" of it all.
And be honest about being differentially valuable. I once had a prospective consultancy customer ask me for a very specific and elaborate piece of software to be built. He'd been thinking about it for years. I dug in on what he was actually trying to achieve, and I realized that small modifications to an existing in-market product would be able to satisfy his actual needs. I connected him with that company, he ended up with exactly what he needed, and they ended up with a high-end expert user as a resource.
Of course, I ended up putting myself out of a lucrative consultancy job, but I don't regret the decision at all. I suspect that most folks on HN feel the same way. Find a way to make a difference, and be honest with your customers (and leads) about what you're bringing to the table. That kind of integrity pays dividends in the long run if you have something truly valuable.
However, I've been hearing "load-bearing" at least two orders of magnitude more often over the last few months, particularly after uncorking Claude Code for the team.
I don't think it's a dead give-away of AI usage, and I don't think AI usage is a problem. I just think we can introduce phrases into common use by having them be used by common tools. So let's train the models on obscure/archaic terms and see what happens. Heck, we can just prompt it...
If a model couldn't ever do that in the first place, it'll just get stuck.
I work in the "ZeroOne" space, working on concepts and prototypes for things that don't exist in market yet. Sometimes these models crank hard and immolate tokens while grounding themselves on expensive-to-ingest self-developed frameworks. If the results are well judged and the crank-turn latency is low, I'm okay with the cost as long as the model isn't wasting my time.
But when I want to do more boilerplate work, I turn down the model and thinking level and get more traditional about restraining action. For the really hard stuff, I reach for the models that will start a token bonfire in the back yard.
Sandy water at the beach that looks a bit brown? Suspension.
Mining run off that looks like soup and coats everything? Slurry.
What's the Half & Half between milk and cream but for tweaky techniques for fabrication? Slurspension?
Thanks for the openness. I got bit by this one and was, frankly, pretty surprised.
The funny thing about user-facing interaction mechanics is that everyone is part of some minority, and everyone comes with their own sense of what "natural" or "obvious" is. With something this impactful, communicated clarity of behavior will important. Your feature is also doing double-duty, serving as a last net against prompt-injection attacks by giving the user the final say.
(Also, BTW, folks outside of Anthropic are unlikely to be as tooled-up for long-running unsupervised Claude jaunts as you guys. The cost of wild success is wide adoption.)
One thing I'll suggest is that the mechanics of permissions and asking are presently pretty hacker/nerd friendly but simultaneously too-scary and not-scary-enough for non-coders.
Examples:
- Wild-cards on always approve is awesome, but, with prefixes like timeout and nohup, the "thing" that is getting done is buried and largely unexplained to the user.
- Auto is actually kind of a sweet spot (sometimes goes off into the weeds), but the designers and PM's I've been working with might as well YOLO. They have no idea if they're breaking things, but they gravitate between plan and auto mode.
- Fewer permission prompts is great, but it comes after a user has slogged through generation of a data-set to work against, like battle scars for paper cuts. It's the thermostat problem. The signal comes when the user is uncomfortable. And it's a way to learn me, but not me now.
I've had good fortune with Opus 4.8 and Fable just telling the system what phase of my life it's in. Things like "I'm going to go make dinner... Go profile the matrix or configurations and build the dataset for the next two hours while I'm away" have a pretty good hit rate. On the flip side "keep me in the loop and bring me your results before making structural changes" also articulates well with Fable. It will tread more carefully.
And these approaches are the ones we'd use with someone transitioning from SDE1 to SDE2. A little more autonomy, and the grounding in the bigger picture. Can we eventually translate to perfectly judging what the user wants in the moment based on incomplete signal?
No, but I'm glad you're trying. Keep the interaction model clear to your broad set of users, and we'll come along for the ride.
I guess that's the joke.
Heavies have been playing this "pretend internet money" game off and on for decades. In investment analysis, details matter.
Javascript systems have long had polyfills for varied browser feature comparability gaps.
Whether you agree with these, making probing detection via fingerprinting illegal would take away this lever. Making surreptitious tracking via fingerprinting illegal? Even for state actors?
Yeah, that's probably reasonable. If someone is going to wear a tracking collar in exchange for "free" services, a little disclosure makes sense.
Techniques that work for inexperienced engineers with high ability but limited judgment often work well with agentic coding systems.
- Give them clarity of purpose. Why are they doing what they're doing?
- Make explaining it back to you part of the job.
- Give them two-way doors. Make mistakes reversible.
- Put effort into thoughtful refactoring as an actual sub-task instead of just accepting piled on hacks.
- Make your operating rules crisp and make sure they store them in their memories.
- Be accountable for their work. It's not okay to crank out AI Slop and then say "Claude's fault".
We're all Software Development Managers now.
So, micromanage the LLMs if you want to, but you'll be missing out on chances to improve them for your purposes and, more importantly, to improve yourself as a manager.
In the early days of LLMs, we saw the classic hype-driven bi-modality of opinions. Folks were in the "fake news, fad" camp, or they were in the "omg, take over the world" camp.
Those of us closer to the space, with the awareness to know that there was some truth (and a lot of misjudgment) to go around, were in the middle of nowhere. When I co-wrote some driver code with Chat GPT, other engineers (and even one of our directors) told me to keep it quiet. At the same time I had directors and VPs asking me how we could accelerate adoption. For a while, I had access to a cheat code just because I had the audacity to not ask for permission. Folks were sure I would get in trouble for spending thousands per month in LLM operation, but a handful came along for the ride, burning tokens like firewood and learning along the way.
Tokenmaxxing is probably coming from at least a few things:
1. A course-correction for the practiced frugality that kept folks from jumping in and just learning at the ragged edge.
2. A willful and deliberate recognition that the best innovations in the later phases of a disruptive introduction often come from sparks of ideation in concentrations of activity. In other words, we don't know where good is, and we need to find it. (Charitable interpretation from the article)
3. Recognition that, even if they don't know why, leaders and product owners will get punished for not jumping in and, because of bullets 1 and 2, won't get punished for trying and missing. Even if they have no idea what they're doing, they're going to fake it until they make it (or slide into another job).
This last set is where the pain lives. An organization with healthy and increasing AI tool usage will see elevated token counts, but so too will one using LLMs to rewrite wikipedia articles without the letter "m" to keep token counts high. These are pathological behaviors brought on by conflated metrics.
We had discussions about this in the early LLM days, where my old team was looking to ship new capabilities for older products. There was a lengthy VP-level discussion about getting to "80% usage" of the new system vs the old. Because the new system was a superset of the old, I eventually said "we can do that immediately, but it's a cost goal, where we're just aiming to make our business more expensive to operate, rather than a value goal for our users". We didn't adopt the target, but folks were understandably frustrated that they didn't have a straightforward way to measure and report progress.
Tokenmaxxing is, inevitably, a conflated goal, but it's what we have right now. Take advantage of the moment, learn, build, and keep an eye on levers for efficiency.
Experienced folks who know how to describe and articulate through others have a huge opportunity here. I have ultra-quick interns in my laptop, waiting to apply aggressive and slightly presumptuous energy to any and every problem. I also know how to pull them back in and get them to focus that energy (because junior devs were the same).
New folks will sink or swim quickly, but they're less expensive and more plastic on average. They're raised in this. We'll see what that does to quality.
Deeply technical managers, designers, scientists, program managers, and product managers are now in possession of an incredible power, to be able to craft existence proofs to counteract the couched recalcitrance that engineering orgs have held over their judgment for decades. There's a certain intellectual integrity in this, even if nobody can actually read the code at the rate it's being produced.
I've been explaining it like this:
Programming was 1% judgment and 99% effort, where lots of folks could carve out productive careers carrying that effort and receiving that judgment.
Agentic coding has cut that 99% down by at least a couple of orders of magnitude for some work. Well-judged and well-described systems can manifest quickly where effort alone would fail. The 1% is still there, but, by ratio after optimization of the sweaty part, it's at least half of where the value is.
I had an example of this this morning, where Claude Code left to run overnight on an open problem had made an absolute hash of multi-source grounded clustering. I course-corrected it with a rule (I don't like magic number tuning on small datasets) and a specific approach (use clustering with separating anchors/seeds), and it had the system working in 15 minutes (confirmed after a couple of hours of processing). These are the same techniques that we would use with junior engineers.
Along the way, it drafted reports and ran experiments that taught me about some of the limits of SOTA listening/characterization systems that I otherwise would have had to spend time researching.
Just make teaching you an explicit goal of the system, and you'll be able to swivel from opacity to illumination.
- No partially loaded content. - No relayout while content loads.
Holding those as hard rules leads to delay or rejection. Instead, while I agree it's better to have everything up front, gracefully handling cases when we don't is important, and some degree of responsiveness, even with partially loaded content, often makes for a better experience for the user than a delay.
Just be up front about it and find ways to keep continuity of relationship and smoothness. Diffeomorphic mappings are your friend...
I closed my laptop and went to bed, but the moment sticks with me. The potential for para-social relationships is enormous. This is a wild time.
And that's a problem because Aggregability is NP-Hard: https://dl.acm.org/doi/abs/10.1145/1165555.1165556
So a tree is a way to take a high dimensionality graph and make it usefully lower dimensionality, but, given the aforementioned proof, that reduction is going to go from being a lossless compression to a heuristic. So any interesting problem (at least, any problem interesting to me) is only going to be aided (read: not solved exhaustively) by that hierarchy.
I'm okay with this. Being okay with this has been one of the most freeing things over the last 20 years of my career. Accept inaccuracy, and find usefulness in your data structures.
Were the man page musings written in response to the (alleged, but... uh... NSA) kleptographic backdoor in Dual_EC_DRBG? It requires multiple successive outputs to compromise and derive internal PRNG state, if memory serves.
In that one construction, /dev/random blocking on seeding would have a mild state-hiding advantage over /dev/urandom, I imagine... but, sheesh. Nobody use that generator.
Customer trust is a lot easier to lose than it is to gain. Moments of frustration are the perfect time to step up and prove to the customer that they can trust you to make things right.