2,029 karma · joined July 7, 2011
Obviously nation states will likely have significantly more resources than this, but this is not script kiddie levels of GPUs.
I put up with it when I'm actively using an LLM tool, because I just accept that this is how these tools write, but when reading a blog I expect to read words written in one of the vast varieties of normal human-written blog formats.
It reads to me an argument for self hosting, maybe a less capable model to deal with smaller compute resources, but when starting to use that model and inference software to build a benchmark you can easily rerun when you change things to observe the impact of the change.
Maybe that’s a similar level of diligence but they feel different to me.
Catching cheaters is still hard but it will keep honest people more honest.
Are the contracts actually take-or-pay-style? What are the terms? What are the amounts of compute and money involved for each future time period? What happens if the datacenter costs spiral upwards? What happens to the datacenter investment if the buyer goes bankrupt, goes public, or gets acquired (potentially by the datacenter owner)?
I personally think the datacenter build-outs are going to generally be a big swing and a miss. The problem 1-2 years ago was making models good enough to be useful for a variety of tasks. Now we have that. The next problem is making the models efficient enough to run a profitable business. Recent Chinese lab model releases (because they're already constrained on compute resources) and OpenAI price cuts on Luna seem to indicate that this transition to chasing efficiency has already started.
Comparable adult parent things would be up at 6am, prep kids for school day, in the office by 8am, 8 hours of work plus a lunch, departing the office around 5pm, transporting kids to/from some activities, dinner, helping kids with homework, somehow fit in regular shopping needs (food, essentials, etc), maybe some volunteering or hobby things, if lucky in bed by 10pm, and we're easily at 15 hours of "work" on a weekday.
Yes, it's a lot.
Still leaving manual approval for all edits. Combined with reading the full transcript of the exploration, I feel I stay in the loop pretty well in this first test.
Your summary approval idea is interesting and feels maybe like a mini plan mode. My biggest frustration with the existing manual approval system is when Claude is exploring it gets tedious to approve each command. Being able to approve a block of commands or a mini plan AND have auto mode audit them for safety would probably be something I would consider for the expiration phase of my Claude use.
The AI companies seem pretty bad at setting up tests. And really good at marketing those failures into spin at how amazing their products are.
If you don't mind Meta retaining your data, the "Contributor" pricing is deepseek-v4-flash-level of low, roughly 1/10th normal muse-spark API pricing currently. Attractive if you're OK with them retaining and using your data.
Texas is probably our national leader in terms of deploying solar and wind power, which feels ironic but is totally rational and logical.
Sadly, my investigations a few years back indicated that a suitably strong and robust 2-axis tracker that would withstand our local winds and snow loads is significantly more expensive than just buying more panels and facing some of them east/south/west to accomplish the same thing. Doing this on an existing shed also avoids some local zoning rules vs a dedicated ground mounted system.