HNHacker News
TopNewBestAskShowJobs

bthornbury

422 karma · joined October 16, 2013

https://github.com/brthor https://www.youtube.com/channel/UC4mmZUMvTNzzz2sS_tgbKVQ
submissionscomments
bthornbury··on Introducing System One Models and Jev
Is the tradeoff of the parallel output that we don't get arbitrary string generation? like output # of tokens is fixed ahead of time?

Either way, really cool and impressive.

bthornbury··on Engineers will do anything to avoid learning from history
"Suddenly, waterfall is the thing to do."

Matches my experience. Like 90% of the work is the spec.

Except instead of weeks researching its like a few hours talking with an agent.

bthornbury··on Franken.domains: Stitched-Together Domains, Because Every .com Is Taken
lots of good ones still
bthornbury··on 2x, not 10x: coding with LLMs in 2026
I discuss the testing approach and coverage with the model before and after, sometimes in a fresh thread that does a static analysis.

interestingly my input is still pretty important

bthornbury··on 2x, not 10x: coding with LLMs in 2026
codex pro plan currently
bthornbury··on 2x, not 10x: coding with LLMs in 2026
for me, almost all of the work is specs

I am no longer:

- reading docs for hours and hours

- typing (barely at all)

- writing code

- manually doing tight debug loops

- using an IDE

to do this I had to give up reading or even controlling the code and focusing on behavior/design-level control (not superficial, still dictating overall technical architecture)

i have agents doing everything from writing the code, verifying the code, hardening, increasing test coverage, analyzing behavior, algorithmic perf improvements, managing/deploying to cloud resources, etc... (pretty much everything)

and I am accomplishing projects that would take months or years in a fraction of the time.

that's way more than 10x.

somehow, this is harder and more cognitively demanding than writing code

bthornbury··on SubQ 1.1 Small
we need some better standard long-context benchmarks.

needle in a haystack is not good for this, yes it proves the model can attend to its context, but in its usual form, somewhat trivializes the query-key relationship.

something like long-form Q&A would be more ideal. Like reading a book and answering questions that require synthesizing information derived from either the whole thing or disparate portions of it. Like describing an entire character arc in a 1000 page novel with examples and evidential moments.

bthornbury··on Running local models is good now
the qwopus 27b model is good for grunt work style tasks, even across multiple files. Piping a bunch of things through, small factoring changes, stuff that just takes time to type out.

I wouldn't rely on it for large stuff like codex though. I haven't tried out deepseek/kimi, if we could run those locally it would be great.

bthornbury··on AI coding at home without going broke
promote yourself to PM only and use agents for authoring, verification, tests, checking the tests

orchestrator -> parallel subagents with investigation, authoring, verification, benchmarking subagents and integration / final verification handled by parent has improved my productivity too.

I feel like from here its agent swarms against a whole spec but haven't got there yet.

Still getting plenty of bugs in the more complex scenarios, but mostly (in some projects) i never have to look at the code and treat it like a black box

bthornbury··on An Intuitive Understanding of AI Diffusion Models
The classic papers describing diffusion are full of dense mathematical terms and equations.

For many (including myself) who haven’t stretched those particular math muscles since diff eq class a decade or so ago, the paper is just an opaque wall of literal Greek.

In this post I describe my personal understanding of diffusion models in less-dense terms, focusing on intuitive understanding and personal mental models I use to understand diffusion.

bthornbury··on How an inference provider can prove they're not serving a quantized model
Something like a perplexity/log-likelihood measurement across a large enough number of prompts/tokens might get you the same in a statistical sense though. I expect those comparison percentages at the top are something like that.
bthornbury··on How an inference provider can prove they're not serving a quantized model
AFAIK seed determinism can't really be relied upon between two machines, maybe not even between two different gpus.
bthornbury··on How an inference provider can prove they're not serving a quantized model
Is modelwrap running on arbitrary clients? I'm not following the whole post, but how are you able to maintain confidence in client-owned hardware/disks following the secure model the method seems to depdend on?
bthornbury··on Coding agents have replaced every framework I used
Why does there seem to be such a divide in opinions on AI in coding? Meanwhile those who "get it" have been improving their productivity for literally years now.
bthornbury··on My AI Adoption Journey
> got a load of ticking time bomb bugs

Lots and lots of tests!

bthornbury··on My AI Adoption Journey
Either really comprehensive tests (that you read) or read it. Usually i find you can skim most of it, but like in core sections like billing or something you gotta really review it. The models still make mistakes.
bthornbury··on My AI Adoption Journey
AI is getting to the game-changing point. We need more hand-written reflections on how individuals are managing to get productivity gains for real (not a vibe coded app) software engineering.
bthornbury··on IKEA for Software
I'm not too sure about this take. The larger code rewrite issue is constantly trying to be solved, which is somehow making the problem worse.

In another view, standard libraries do a pretty good job.

bthornbury··on We reverse-engineered Flash Attention 4
I'm pretty sure it's called "reading the code". That said, it is difficult enough in its own right.
bthornbury··on There are no new ideas in AI only new datasets
This generalization issue in RL in specific was detailed by OpenAI in 2018

https://arxiv.org/pdf/1804.03720

bthornbury··on [dead]
Recently, I've been using a local docker container to house the interpreter for all of my new python projects.

For day-to-day work it is far superior than endless virtualenvs clogging up my harddrive and hunting for brew package dependencies.

I open sourced a template in case other may find it useful.

bthornbury··on I have been using Mixtral everyday for coding and I think it has saved me days
Note that I had to remove two of the test cases to fit in the HN character limit:

                      {
   name:         "Large Input Slice",
   input:        []any{"A", "B", "C", "D", "E", "F"},
   chunkSize:    3,
   expectedChks: [][]any{{"A", "B", "C"}, {"D", "E", "F"}},
  },
  {
   name:         "Remaindered Large Input Slice",
   input:        []any{"W", "X", "Y", "Z", "1", "2"},
   chunkSize:    4,
   expectedChks: [][]any{{"W", "X", "Y", "Z"}, {"1", "2"}},
  },
bthornbury··on High taxes be damned, the rich keep moving to California
Might be able to afford an apartment
bthornbury··on Erlang/OTP by Example
As an Erlang fan, who hasn't used it in production services, I'm wondering if you can let us know some specific scaling issues you encountered.
bthornbury··on Evergreen: a React UI Framework built by Segment
Dependencies have cost. You have to monitor for updates, notify the maintainer(s) of any bugs, keep an eye out for security vulnerabilities, and sometimes (gasp) even step through them with a debugger.

Doing that for one dependency is bad enough, but for 100s it's a nightmare.

Personally, I prefer to just pull the pieces I need out of an open source library (unless it's very well maintained, or huge). It's like doing a code review at the same time, so you're aware of what's going on in your application.

bthornbury··on Evergreen: a React UI Framework built by Segment
My first thought was that this is a UI Framework (Evergreen) built on UI Framework (React) built on a UI Framework (js + html).

No real point here, but I do find it curious that there is this tendency to build new UI frameworks all the time.

Personally, I still like html + js (and I have used React).

bthornbury··on Kubernetes Is a Surprisingly Affordable Platform for Personal Projects
Kubernetes has a minimum node count. Moving to one node, saves cost, by a factor of 3. Not to mention all of the other resources it creates (load balancers).
bthornbury··on Kubernetes Is a Surprisingly Affordable Platform for Personal Projects
Sure.

The largest reason is cost. My small deployment (3 nodes) is running around $100 / mo on AWS (That's my app, nginx, redis, and postgres).

It doesn't even need 3 nodes, I don't recall if 3 is the minimum, but realistically I only need one (for now). For larger projects this is probably a non-issue.

Second largest reason, is that really I have no idea what is going on on these nodes, and I probably never will. Magically my services run when I have the correct configurations. Not to say that's always a bad thing, but I've found it difficult to determine the default level of security as a result of this.

A third reason is the learning curve. This is less of an issue because I've invested the time to learn already. But like, the first time I tried to get traffic ingressed was painful.

As to what I'm moving to, I migrated one of my websites to a simple rsync + docker-compose setup and am pretty happy with it. In the past, I ran ansible on a 50 node cluster and it worked really well.

bthornbury··on Kubernetes Is a Surprisingly Affordable Platform for Personal Projects
Sounds like you can get away with a bash script that calls rsync, and restarts the code on the server with ssh. (I deploy one of my websites this way, simple & reliable)

If you're looking for something more legit: Ansible

bthornbury··on Kubernetes Is a Surprisingly Affordable Platform for Personal Projects
The learning curve is sharp. The amount of things happening you don't have awareness of is also worrisome (to me).

Source: I am also using Kubernetes in production, migrating off it soon.

Page 1 of 7Next →