Either way, really cool and impressive.
422 karma · joined October 16, 2013
Either way, really cool and impressive.
Matches my experience. Like 90% of the work is the spec.
Except instead of weeks researching its like a few hours talking with an agent.
interestingly my input is still pretty important
I am no longer:
- reading docs for hours and hours
- typing (barely at all)
- writing code
- manually doing tight debug loops
- using an IDE
to do this I had to give up reading or even controlling the code and focusing on behavior/design-level control (not superficial, still dictating overall technical architecture)
i have agents doing everything from writing the code, verifying the code, hardening, increasing test coverage, analyzing behavior, algorithmic perf improvements, managing/deploying to cloud resources, etc... (pretty much everything)
and I am accomplishing projects that would take months or years in a fraction of the time.
that's way more than 10x.
somehow, this is harder and more cognitively demanding than writing code
needle in a haystack is not good for this, yes it proves the model can attend to its context, but in its usual form, somewhat trivializes the query-key relationship.
something like long-form Q&A would be more ideal. Like reading a book and answering questions that require synthesizing information derived from either the whole thing or disparate portions of it. Like describing an entire character arc in a 1000 page novel with examples and evidential moments.
I wouldn't rely on it for large stuff like codex though. I haven't tried out deepseek/kimi, if we could run those locally it would be great.
orchestrator -> parallel subagents with investigation, authoring, verification, benchmarking subagents and integration / final verification handled by parent has improved my productivity too.
I feel like from here its agent swarms against a whole spec but haven't got there yet.
Still getting plenty of bugs in the more complex scenarios, but mostly (in some projects) i never have to look at the code and treat it like a black box
For many (including myself) who haven’t stretched those particular math muscles since diff eq class a decade or so ago, the paper is just an opaque wall of literal Greek.
In this post I describe my personal understanding of diffusion models in less-dense terms, focusing on intuitive understanding and personal mental models I use to understand diffusion.
Lots and lots of tests!
In another view, standard libraries do a pretty good job.
For day-to-day work it is far superior than endless virtualenvs clogging up my harddrive and hunting for brew package dependencies.
I open sourced a template in case other may find it useful.
{
name: "Large Input Slice",
input: []any{"A", "B", "C", "D", "E", "F"},
chunkSize: 3,
expectedChks: [][]any{{"A", "B", "C"}, {"D", "E", "F"}},
},
{
name: "Remaindered Large Input Slice",
input: []any{"W", "X", "Y", "Z", "1", "2"},
chunkSize: 4,
expectedChks: [][]any{{"W", "X", "Y", "Z"}, {"1", "2"}},
},Doing that for one dependency is bad enough, but for 100s it's a nightmare.
Personally, I prefer to just pull the pieces I need out of an open source library (unless it's very well maintained, or huge). It's like doing a code review at the same time, so you're aware of what's going on in your application.
No real point here, but I do find it curious that there is this tendency to build new UI frameworks all the time.
Personally, I still like html + js (and I have used React).
The largest reason is cost. My small deployment (3 nodes) is running around $100 / mo on AWS (That's my app, nginx, redis, and postgres).
It doesn't even need 3 nodes, I don't recall if 3 is the minimum, but realistically I only need one (for now). For larger projects this is probably a non-issue.
Second largest reason, is that really I have no idea what is going on on these nodes, and I probably never will. Magically my services run when I have the correct configurations. Not to say that's always a bad thing, but I've found it difficult to determine the default level of security as a result of this.
A third reason is the learning curve. This is less of an issue because I've invested the time to learn already. But like, the first time I tried to get traffic ingressed was painful.
As to what I'm moving to, I migrated one of my websites to a simple rsync + docker-compose setup and am pretty happy with it. In the past, I ran ansible on a 50 node cluster and it worked really well.
If you're looking for something more legit: Ansible
Source: I am also using Kubernetes in production, migrating off it soon.