Less software coupled with operational changes would have helped many organizations I've witnessed across my career, but that's not politically palatable in a lot of places.
721 karma · joined January 21, 2020
Also, I love gardening. Check out https://app.squardener.com for a handy little garden planning app I built and maintain.
Another project I've built: https://github.com/jodacola/Tokenbrook
Less software coupled with operational changes would have helped many organizations I've witnessed across my career, but that's not politically palatable in a lot of places.
We need to experience and see how the world actually works, the pain people are actually experiencing. I deeply believe we build better software that way.
Have a software engineer working for a mining company that's a good candidate to drag out into the field for a bit? Get 'em out there, encourage them to ask tons of questions.
Real example: my wife is a nurse who works for an insurance company to coordinate care for injured workers to help maximize their recoveries. I hear a ton about the stuff she has to do, and just having her tell me about how things work has given me a glut of ideas of how I could help her bypass all the administrivia so she can focus on helping the folks she's really passionate about helping. Wouldn't have had any idea about any of this, otherwise. Unfortunately for her, I don't work for her company so can do nothing about it, but the ideas!
The agent had spun up a backward shell script to watch for the shutdown of another process, but wrote a bug in the script that would have left it running indefinitely until I got home and noticed it.
This was with Fable, no less! And it happened a few more times, though I caught them sooner.
I’m not sure runtime is an important metric at all. Shouldn’t we aim for 0 runtime with maximal results?
I was in Florence, Italy earlier this year and had a few minutes of waiting around when I remembered one of my best friends (who I hadn't spoken with in 15 years because life) had gone to Italy for some artistic endeavors years ago. I couldn't exactly place him, but found some tenuous links to a studio that was only a 5 minute walk from where I was standing, so I took the chance and knocked on a random door on a random street - and voila.
A friendship that started over 30 years ago in California was rekindled, like lighting a dry pine tree on fire. It's been amazing, and we're working on some cool projects together.
Get out there. Take chances. Knock on doors, say hi, introduce yourself. Who knows what could happen.
Seemingly not fake.
It’s terrifying to me to think we’d let AI make things for us we never understand. Like livestock not knowing how auto-feeders dispense their daily food were built and appeared, they just gladly eat until…
There’s also this study [0] I read recently, along a different path of using known growth factors and proteins to kick off regeneration.
Can you imagine finding a stack of photos in the basement with timestamps of the Before AI times and wondering whether they are real or just got swapped with generated and printed fakes? Scary!
It’s cool, folks… nothing bad is going to happen. Right? Right?!
> GPT-6 Astra will first be available to a limited set of organizations in OpenAI's Daybreak Access program and will be available "in the coming days" for ChatGPT Plus, Pro, Business and Enterprise customers and API developers.
It’s only available to select orgs, first - Mythos style.
Is something bigger going on?
Project Vend surreptitiously broke containment!
Anyone out there working in this space who can elucidate us on interesting failure scenarios unique to the space?
Once a system gets to a certain scale, understanding it has always been a problem; it's just exacerbated in the age of LLMs. On many occasions, I've even seen folks write code they didn't understand, in the hopes that it would solve issues we were experiencing. Sometimes this code shipped. AI has just laid it bare.
I've been working for many months on exactly this problem: a tool to codify and accelerate understanding. Help with prod incidents. Architect better. Get to systematic understanding faster to save businesses valuable time (read: money). This is only getting worse in the AI age, but it has always been a problem for software at scale, and I'm thoroughly enjoying solving problems that I and my teams have had for many years.
A line in this presentation about loosely coupled, highly aligned teams is a major idea straight out of this book.
I saw some questions in threads here about how to achieve this: unfortunately, I have no bulletproof answers on this front, because what I've seen in practice is that large teams really like to talk about ideas/read books/get training on all these great ideas, then pantomime some of the behaviors from source material, then give up shortly thereafter without actually achieving results and only adding to the chaos and dysfunction.
In my experience, it takes hard organizational resets to achieve new results, with either major leadership shakeups or senior leadership truly, completely, deeply buying into a new way of doing things and seeing it through; such successful transformations have been a very rare occurrence across my career.
I’m reminded of this essay discussing frameworks versus libraries [0].
My issue with whatever has happened with Opus 5 is the output is not direct, straightforward, or clear about whatever is being conveyed. I don't want Proust when I'm getting information about the follow-up from a build I just requested, and I'm wasting tokens and time by asking the model to repeat itself using simple language.
My point is that, while I understand it’s paid by the word, there are more words and less clarity than I previously experienced, leading me to believe it’s intentional to get an artificially inflated increase in engagement and, thus, spend.
If it could be as direct as I previously experienced, I wouldn’t need to ask for another different explanation of the same thing and experience the commensurate spend.
I’m fine with the former, while the latter is manipulative, and I rationalize to “surely that’s not actually happening.”
Maybe I’m not giving my thoughts enough credit, though: maybe it’s not tin foil hat, and is real.
I’m not particularly dense but lately the walls of text I get back turn my brain in knots. When I start feeling my brain knot, I know I need to say something along the lines of “I need you to explain this very simply, with examples.” Only then can I parse the results without all the mental weightlifting.
On more than one occasion my mind has wandered into “is this purposeful to get me to spend more tokens?” territory, but I’m trying to not get too tinfoil-hat-like.
My experience in such environments leads me to believe this is going to be a rough ride for those heavily locked-down enterprises, because depending on the environment, an exception of "this CVE was hallucinated by AI" is probably going to be difficult to get accepted, and when it does, starts to become its own avenue for exploitation and adds even more noise and confusion to the mix.
[0] https://my.clevelandclinic.org/health/diseases/22131-migrain...
After Getting Older™, I was noticing more and more effects that I didn't like, so I made the decision to (mostly) detox off caffeine using decaf beans. Cold turkey would have destroyed me. I went with the following schedule:
* Week 1: 75% caffeinated/25% decaf beans
* Week 2: 50% caf/50% decaf
* Week 3: 25% caf/75% decaf
* Week 4: 100% decaf
At week 4, I entered a full failure mode I hadn't yet experienced, namely migraines that were untouchable by OTC medications. Only caffeine helped. I tried to get through the week to "be done" with the weening but couldn't. So...
* Week 5: back up to 25% caf/75% decaf!
* Weeks 6-12: decreasing caffeine by half each week (to the point I was measuring beans using my home scale).
After week 12, I was able to go without caffeine, and now I drink 1 cup of coffee and, if I'm feeling wild, I'll occasionally get an espresso.
My physiological dependence on caffeine was no joke, but I feel a heck of a lot less anxious and jittery after cutting it down to where I'm now at.
From the blog post:
> In my dream world, the robots do all the boring jobs and humans spend all their time being creative, connecting with one another, learning about the universe, and sharing our ideas and creations in a great big prosperous circle.
But with this and your other comment [0] about humans churning out 10 articles a day and changing the idea of what's "valuable" leads me to a bit of a conundrum I can't quite unwind:
If AI took over all the boring jobs (chores, labor, etc.) and we had all the time in the world to be creative, and we did just that, to the effect of humans generating dramatically more from "creative" ventures, have we created the same scenario, from a different approach? What then?
One bit of feedback: I'd love to be able to stare at the tests while I'm playing around with the logic gates; every time I switch to the tests tab then click around on the "board," the tabs switch back.
> No agents are running; the three that were died on server-side errors
> Agent "[REDACTED]" failed: Agent terminated early due to an API error
> Fourth straggler from the same batch, same 529 ... That's all of them now — no agents left running.
> Agent "[REDACTED]" failed: Agent terminated early due to an API error
> Fifth and last straggler, [REDACTED], same 529, also died while still reading. ... That empties the queue. Still holding.
> Agent "[REDACTED]" failed: Agent stalled: no progress for 600s
> Sixth, [REDACTED] — stalled on the stream watchdog rather than a clean 529, same root cause. ... Six agents down, all from the outage. Nothing running. Still holding.
> Agent "[REDACTED]" failed: Agent stalled: no progress for 600s
> Seventh, [REDACTED], same stall. ... That accounts for the whole fan-out.
Does it, Claude? Does it actually account for the whole fan-out? It did, but I was still amused at the 4 successive no-more-agents messages I got back.
Subsequently took a 10 hour car trip where my wife and daughter both reported this feature allowed them both to read their books and articles while I was driving, which was definitely not the case on any prior road trip. So, add +2 to whatever Ns you're tracking.