7,927 karma · joined July 15, 2013
https://a.drien.com
https://github.com/drien
https://infosec.exchange/@adrien
That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.
I did also have a cool experience, deep in the backcountry at a high elevation camp some years ago: we went to sleep in a gale, the tent flapping all around, with a cloudy sky. I thought we’d get a big storm, but I woke up hot at 3am, the air totally still. I slid out of the tent and the sky was completely crystal clear. An incredible spread of stars. I watched for a while and noticed a reddish glow on the horizon. We were so far out and it wrapped around in all directions, so I knew it couldn’t possibly have been a town, but I was puzzled. I learned about airglow when we got out of the woods: https://en.wikipedia.org/wiki/Airglow
The AC thing interests me as a topic–purely my opinion, but I do think that living in a very narrow temperature range changes your relationship with the outside world, and there's all sorts of good stuff to be done outside that's much harder to get motivated to do if the delta with what you're used to is too big.
As an anecdote: I mostly disliked hot weather in the past, but years ago when I got hooked on running, I found that I disliked the treadmill more than the heat, so I ran outside through a muggy NYC summer. It was miserable to get started but got a bit easier every year, and over time I've noticed, from the digital thermometer on my desk, that the temperature where I have to stop working and turn the AC on has slowly crept upwards. My running club puts on a 5k summer race series–if you'd asked me 10 years ago whether I'd want to run up a hill in the swampy Brooklyn heat every other Wednesday, I would have laughed out loud at the absurdity of the idea–but today it is genuinely something I love, and has been a highlight of my year for several years now.
All that to say, your point about suffering for suffering's sake stands: I cannot sleep without AC above a certain temperature/humidity, and I don't see any reason to force myself to suffer through restless sweaty nights. On the flip-side, though... I have never had more visceral, euphoric appreciation for the comforts of modern life than after a cold and drizzly backpacking trip where we spent the entire time damp and never really, fully warmed up for multiple days. 20-something years later I have the most vivid memory of getting in a hot shower when we got home.
An accessible intro is Michael Easter's book The Comfort Crisis, which, while a little pop-sciencey, explores the idea that making our lives more and more comfortable (across various axes) creates a feedback cycle that shrinks our comfort zones, making some of the more rewarding (but potentially challenging/uncomfortable) things in life seem unthinkable.
For me this was a bit of a "once you've seen it you can't unsee it" kind of thing. Letting convenience be your primary metric for success often produces totally unsatisfying results, anywhere from how you build cities to the way you eat a meal.
The arc of the original World of Warcraft (broader social impacts aside) is an interesting microcosm to explore: players continually asked for convenience features that would simplify the most tedious aspects of the game and reduce social friction (organizing groups, traveling etc), but as they made things more convenient over the years, many people found that the game actually began to lose something important. The response culminated in a very successful re-release of the original, inconvenient, game after 15 years.
The harness is the whole thing here. AI generates text. Everything else is undertaken by harnesses and infrastructure humans provide, have control over, and therefore responsibility for.
Every action AI takes is fundamentally not independent, it requires an explicit choice to let the AI write code, have a physical machine to run it on, to have network access, etc. The concept that these things are "rogue" ignores the role humans play in giving them goals and tools to pursue those goals, and makes it seem like it’s a self-determined force, over which humans cannot exercise control at all.
First world problems, perhaps, but also a microcosm of our broader societal divergence, where a reasonably comfortable middle is increasingly being replaced by growing working and upper-income classes, with very different lifestyles.
I have seen coding agents on my own machine (in sandboxed VMs) start doing things while trying to accomplish what I've asked that I felt sort of exceeded my mandate (changing database passwords, poking at the egress proxy that's preventing them from accessing some domains). Not to the point of causing any real issues, but I don't have much trouble envisioning scenarios like this when using stronger instructions around pursuing the goal + a running in a misconfigured sandbox envrionment.
That said, there's a lot of potential upside for American AI labs if they're able to get people scared about AI, they can:
- To your point, claim the regulations slowed them down and paper over near/mid term financial concerns
- Get the government to create stupid regulations that don't actually slow them down at all, but do effectively lock out any future competition (and current global competition)
- Position themselves as the only organizations blessed by the government with the ability to make safe AI, therefore eventually allowing them to claim to be some flavor of "too big to fail" and worthy of a bailout, should the financials not work out.
- Effectively create a distraction that avoids further public conversation/accountability/regulation/liability re the more tangible sorts of problems their products cause right now.
There are arguments that these are factors that filter out the people who are not sufficiently motivated, but it's hard for me to imagine there aren't a lot of bright young people who might be interested in medicine, but see one of the various paths that exist today to making doctor-level money with only an undergraduate degree and in an environment that doesn't require a working schedule that actively harms your health.
Apprentices have their work directly checked by the person supervising them, who is also ultimately personally responsible for the results.
I am a programmer but also like to work with my hands—I’ve done some amount of "real" electrical, plumbing, carpentry, light construction, mechanical, and landscaping work. To me, this line of thinking overall falls into a category of "things computer programmers believe about other kinds of jobs," that I see fairly often on HN. It’s key to the whole argument here but kind of papered over with an unvalidated assumption it’d probably already work with today’s AI.
Anyone who has followed a DIY YouTube video more than a few times has encountered situations where having muscle memory and proprioceptive experience with actually manipulating the tools and objects you’re working with would significantly reduce the chance you mess it up. How tight is too tight? Am I going to shear this bolt head off? Is this saw blade getting dull and going to hurt me? Even the best guides and teachers don’t cover the combinatoric ways things can frequently go wrong–it’s accumulated knowledge in the heads and hands of people who don’t blog about it so it can be slurped into training data.
I think there’s an enormous chasm between "can AI provide reasonably correct guidance for an inexperienced person along a happy path to accomplish basic physical tasks" and "I no longer need to hire an electrician because AI guides me through every step."
User: Hello Uber support? Yes, the app shows my driver is in the ocean off the west coast of Africa? And his ETA is -2147483648 minutes?
Uber: Your app, your problem.
User: Hello phonebrain, can you fix my Uber app?
Phonebrain: You're absolutely right, I must have misunderstood the API docs. Squiggling... ... Shuffling ...
Phonebrain: Error, you've hit your usage limit. Please wait 12 hours, or switch to pay-as-you-go app development mode. I cannot estimate how much it will cost to fix your Uber app.
Stumbled on my old iPhone 5S in the back of a drawer recently and could not believe how good it felt in my hand, even having only used 12 and 13 Mini models since 2021. What an excellent form factor and construction. Would happily pay a big premium for a modern version of it. Dreaming of a sub-mini phone trend. Sell me a $1500 Zoolander phone—not even joking!
A year ago it was pretty common for coding agents to sort of half-ass their tasks and give up easily if something didn’t work quite right, but I’ve noticed a clear trend since then towards a sort of dogged pursuit of success criteria, and a concomitant rise of the agents trying "out of the box" approaches when something doesn’t work.
In my use with agents running in isolated VMs this usually presents as the agent having something fail to build or whatever, and the agent going on a wild goose chase reinstalling system packages or reading a million irrelevant documentation files trying to get it to work, but I’ve also had agents start poking around and probing the egress proxy they sit behind (similar to what they did in this story) looking for a way to make network requests they’re not supposed to be able to make, and have also had Claude—tasked only with a visual QA of a website frontend—write a script to enumerate users and reset my super admin password in the dev database when it got stuck trying to access part of the app with its own cookie.
It was totally annoying to me to not be able to operate the phone with one hand without feeling like I was about to drop it. I kept the 17 for most of the return window thinking I’d get used to it, but I just kept finding more situations where it bothered me. Battery life on the mini is not amazing, but a slim magsafe powerbank makes it largely a non-issue.
On the other hand, I operate an app that talks to (and logs prompts/meters costs) to many different LLM API providers, and I do not consider it painful at all. I have an AI agent to deal with any integration quirks, if needed. Mostly they provide OpenAI-compatible APIs anyhow. It's basically a no-brainer to go direct with the providers and save 5%, the great majority of the cost of an AI-powered app is no longer dev time implementing the integration, it's the tokens themselves.
I’ve been using this pretty extensively for a few months on Mac and Linux and have been super happy with it.
I evaluated Fly.io itself several years ago and found it unacceptably buggy as well for any real use case at the time, but I kept an eye on it and today I operate a few production workloads on it that are very stable and reliable—I figured the same would be true for Sprites and it seems to be taking shape that way.
The most recent one that's had me annoyed is the "Fullscreen" TUI feature, which is super unintuitive, implementing its own text highlighting and copy-on-select mechanics, overriding your terminal's native right click. Easy to disable but terrible defaults, IMO. It's not even really clear to me what problem it was actually supposed to solve.
> Same thing just happened on a Claude Mobile session in same Enterprise account. Common theme in both is Sonnet 5, first response after more than 5 minutes (cache miss).
Momentarily baffled, I realized that, despite appearances, the old frame was actually not square, in fact it was a parallelogram. I'd measured the height and width and assumed it was square. The previous (experienced) carpenter who'd built the doors I was replacing had clearly noticed this, and simply allowed for the misalignment in his design. He built perfectly square-appearing doors that mounted to the not-square frame. I had to go back and rework mine considerably for them to fit without looking ridiculous. They're still there and holding up well, but I also still think of this lesson on a regular basis in my day to day life now.