HNHacker News
TopNewBestAskShowJobs

weitendorf

1,307 karma · joined December 3, 2019

Fred Weitendorf

Founder at Accretional (accretional.com). Building an agent mesh based on open source

Sponsor for statue.dev and previously at Google working on Serverless Infrastructure for Cloud Run and Cloud Functions

fred @ company or https://www.linkedin.com/in/fred-weitendorf-40b505b6/

submissionscomments
weitendorf··on Half-Baked Product
Yes, that's the Internet Consensus and the reddit comment section every time a startup doing anything is mentioned.

You should never take a risk, business people are all evil and stupid, you should treat every employment or business opportunity as purely transactional because they'll do the same to you, there's nothing you can do about your job or employment, the only way to win is to cheat because everyone else is doing it, the key to happiness is educating other there's not really any cause and effect involved in the way things work unless you, personally, already know it. Just, you shouldn't do anything unless you understand everything about it, and if you don't it's not your fault.

> not flailing around is very difficult and unlikely

This is literally the defining trait of startups. What makes it stupid is that it's always more complicated than "engineer guy did everything he could but got screwed in the end" and that in real life, sometimes people do actually make money or establish businesses because of decisions they made, and conversely that there are real causes and effects behind things that don't go the way you want them to. Telling a story that doesn't contradict in anyway with consensus (so, directionally correct but always wrong) opinion has no point in the same way that there is no point telling a story where a knight rescues a princess by journeying through the kingdom making friends and overcoming challenges, then confronts the evil guy and kills him, the end. This is just that, but "the shady business guy and the screwed engineers"

weitendorf··on Half-Baked Product
I had the opposite reaction, this felt like a story that was literally purpose-built for pandering the hn audience without saying anything interesting.

Good fiction teaches you something you hadn't seen before, or challenges your perspective, or articulates a point of view or personality that you had never before considered. If it's just "some guy went to work and it sucked and he was right and everyone else was wrong and the Green People did classic Green People bullshit", and there's nothing else complicated or humanizing it, and no real-world lesson or stranger-than-fiction details to it, then what value does it have?

Like, what would happen if you asked a redditor with 10 years of experience reading about startups, but no real exposure to that culture/experience beyond the comment section, to write a story summarizing the consensus opinion on reddit of how startups typically work? Of course, because it's made up it's not wrong, but it exists entirely within the socially-contingent reality of the Internet Consensus.

In the real world there's politics, inter-personal relationships, personalities and personality flaws, and too much detail for "startup flails around" to be something you can reduce to "the startup flailed around". Of course it did, but why and how? A story that says "you know how it goes in all the other stories? yeah, that" or "there was a guy like you and he was good, and all the other guys were idiots and they were bad" has no point

weitendorf··on The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)
I built a very similar tool recently mentioned elsewhere in these comments. I think with the current state of LLMs, harnesses, and related tooling, being able to create or setup self-eval tooling is the biggest differentiator between merely using LLMs to write code vs realizing true 10x productivity wins.

I'm curious whether this is something LLMs are eventually going to be good enough at doing, or something the average developer knows when and how to do, or if this is going be something that's too specialized or difficult for most developers and maybe the next generation of developer tooling products. Now that we're several months into Claude Code crossing the threshold of legitimacy and adoption, I've been surprised at how few projects or developers are doing this yet.

To a certain extent now all you need to do is ask Claude Code for browser automation workflows and CUJ tests in your repo, and ye shall receive, but probably something that just uses base playwright. It would be even better if you could ask to install or use a self-eval tool that already did everything you needed it to do and also knew how to specify/setup automations. I'm assuming the level of agency or mental overhead of embarking on a browser automation side quest is beyond what most developers are used to in the course of their regular work, even though it's not really as hard as it sounds now. If so then self-eval tooling could be a very promising new product category to sell to enterprises.

BTW if you have a link to your project I'd be interested in checking it out! $5.69 for a UAT run sounds very high to me based on how many tokens it typically takes for agents to create automations or steer my similar project, but it could be that your test workloads are much more exhaustive or high-dimensionality than mine are. This is what a basic "go to amazon.com and search for a product, then take a screenshot" automation looks like for ours: https://github.com/accretional/chromerpc/blob/main/recipes/s.... And this is our interactive/dynamic remote steering mode: https://github.com/accretional/chromerpc/blob/main/chrome-pr.... I decided to implement against the Chrome Devtools Protocol (one layer under Playwright) and use grpc service reflection to allow agents to dynamically discover/describe the entire chrome devtools api surface. I just started working on a way to gather traces and monitor/manage the automation run internals because I think there's a ton of opportunity in this problem space for orchestration and RL

weitendorf··on The gauge broke: devs felt 20% faster with AI, measured 19% slower
I think most people who strongly identify with tools like vim do so out of a sense of identity-building to "be the kind of developer who is good at vim" / embody some kind of aesthetic or in-group signal moreso than an actual desire to be more effective at getting work done.

As long as you don't have some kind of stochastic or >5s impediment taking you out of a state of flow, most developers' productivity is going be vastly more influenced by their knowledge, understanding, and ability to focus on the problem they are working on than the marginal difference in time it requires to perform some navigation or editing task. Which is not to say that vim is bad or that you shouldn't use it, but that it's just a text editor and if you get triggered by someone not liking it or thinking it's more trouble than it's worth, it might be worth taking a step back and thinking about why it's something that triggers an emotional/defensive response, rather than the kind of reaction you'd have to someone liking strawberry more than vanilla.

weitendorf··on The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)
There are multiple developer subcultures (nominally for productivity but mostly hobbyists) pretty much exclusively motivated by installing and configuring complex, visually-dense, high-learning-curve tooling and editor setups driven by the same psychology.

Besides the cultural association between keyboard navigation and complex tools with being a 1337 h4xx0r, I think there is something to be said about the process of tinkering with and learning how one's own tools work, or more generally experimenting with new, "interesting" ways of working than the default choice (which around where AI was at the time of this study), and being more engaged and thus more knowledgable about one's own work or problem domain, even if the overhead ends up being a poor investment time-wise upfront. Personally, if it took me 20% longer to accomplish something but I understood it 95%, vs 75% if I had done it the "fast way", I would almost always take the 20% latency hit, with the expectation that more knowledge/exposure to different tools and techniques would have much better ROI over time than marginally faster delivery.

There's a certain kind of developer (much moreso the kind all-in on AI in early 2025) who thinks that AI is really smart and knowledgeable and assigns a high degree of confidence/deference to its responses, the same way you might to a venerated subject matter expert or wikipedia/stackoverflow/google search result. To this person involving the AI lends more credibility/confidence in their work and their own understanding of it (vs if they just uncritically copied code off stackoverflow). Better understanding this kind of user made me realize that the quality signals and mental models people build around productivity can vary immensely even within the same profession or team. Productivity-hacking is a lot more about vibes and identity-construction/tribal affiliations than most people would like to admit.

weitendorf··on The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)
I ran into this with my first attempt at building a static site generator "for agents" last year. https://statue.dev

I got very frustrated with LLMs and their inability to apply good taste or maintain consistent design languages, and put the project on ice. But I decided to double down on more tooling and learn as much about frontend as I could because I also realized that frontend itself - the problem domain, the engineering culture (or lack thereof), the historical baggage, the sheer size of the frontend api/language surface was part of the problem. And also there was/is a lack of good LLM and agent-oriented tooling that was a much deeper problem than I expected initially.

I originally thought I would just create skills/workflows and apis for generating sites from templates, but the problem is moreso that you need an entirely different kind of harness and development process for frontend, which doesn't really exist yet. Claude design is probably the most familiar gesture in that direction for most people but I think it's only scratching the surface. Our own "agentic playwright" is https://github.com/accretional/chromerpc/tree/main/chrome-pr... - IMO this kind of tool (both ours and Claude Design) is a major win for removing the largest, most frustrating frontend LLM painpoints (having a human doing QA and prodding the model to fix obviously-wrong outputs).

But the bigger problem is that the webdev tooling ecosystem is FUCKING AWFUL, and there are too many different ways to do something even using the actual base browser apis, let alone all the random ass low-quality tools and cargoculting that seeps into the models' way of working and thinking. That's not to say that tools like React are bad, necessarily, but that there is so much pre-LLM slop and churn and low quality/inconsistent work in the frontend ecosystem that you really need to be MUCH more knowledgable about the way browsers and the web actually work than the median frontend developer (especially the ones participating in the endless hype flavor of the months, generating all the noise that defines the engineering culture) to effectively use them. Or even better, if you know enough you can also NOT reach for them because you're able to just implement it via raw html/css/browser primitives instead of through 2000 node packages.

To be clear, I'm not saying frontend development is slop, but that it has a very high skill ceiling and requires a lot of very particular/thorny knowledge to be good at. I think the reason AI frontend looks so much like slop is that it hasn't been RLed against the actual web-standards in a way that lets it learn how to actually build good sites, it just has the median frontend engineer archetype from its pretraining and then some kind of RLVR to get it to produce workable, not-fucked-up code (the 3x3 grid, the slop hero, the unnecessary blinking green buttons, etc.). And also, for LLMs, maybe engaging with the webdev tool ecosystem beyond the core infrastructure layer and base apis/languages is more trouble than it's worth, because they often optimize for "I want a particular kind of UX and don't know how to implement it directly, but I do know how to find a package and call it, then prod it into working".

LLMs need something more like a browser-harness, a meta-design system, per-design-language component management tooling, and a non-slop build system. They also generally need much better support/more sophisticated UX for hierarchical iframes, CSP, etc. which is a space that is not very well-explored despite its potential, because most frontend devs find it too hard or complicated.

People are already starting to build these and I think we'll get there in the next year. The hardest piece of the puzzle is figuring out how to structure RL training envs to learn frontend directly against web standards, because web standards are very complex and high-surface; but this is also the most promising because it's how you get Mythos-like superhuman performance. We have a project to build some of the base domain modeling/search tooling needed for frontend RL, eg https://github.com/accretional/proto-css, but it's early days. You should definitely try agentic browser tooling if you haven't yet because it makes a huge difference in getting existing LLMs to be more effective at frontend, and automating most of the debugging. It's what allows us to eg fully automate creating gifs of models interacting with our site in the context of a user journey when we run tests: https://github.com/accretional/proto-css/tree/main/chrome-te...

weitendorf··on What can you confidently guarantee about your software?
I have been working on metaparsing and some related tools for spec-based-development, fuzzing, and property-based-testing off and on for the past few months. I hope to make it easier for application developers to get the benefits of formal verification without having to spend inordinate amounts of time or resources writing unfamiliar, difficult formal-verification software or tests. You're right that it's not useful for most developers to concern themselves with this now, but I think that's mostly just because nothing makes it easy enough yet.

> Even property-based testing is mostly unhelpful for e-commerce apps like these.

Only in practice because of lack of application-developer oriented tooling. Why can you not have a callback assertion or database trigger that alerts or opens a ticket when a refund is not processed within a specified amount of time? Why can't you map out your payment processors' APIs and, if amenable, define a more transactional interface on top of their CRUD API to automatically rollback incomplete or partial state transitions? (Obviously, not all payment processors will expose an interface allowing you to do this in every case; but often times they do provide the fundamental pieces in-practice, just not in a way that is easy to discover or implement). Why can't you define cross-dependency static analysis guaranteeing that eg the core checkout flow continues to work before even running any integration tests? With the right tooling you could, it's just unreasonable to expect every developer working on ecommerce to get a PhD in automated theorem proving to build their own tools for that.

> tricky logic with just a bit of I/O

I actually think that IO is the #1 most important and high ROI surface to apply formal verification (and fuzzing/other related tools which are downstream of formal methods). How many ecommerce apps are secure against DOS or runtime panics from an adversarial client sending them junk data? How many write their own janky/insecure parsers against input formats or external APIs and introduce security vulnerabilities there? I would hazard a guess that most 5+ dev webapps are guilty of basic IO sins that formal verification tools could identify in short time and also prevent from being introduced if they were added to the core development process.

The approach I'm taking is to use codegen to automatically convert OpenAPI specs and grpc service definitions (and in some cases actually just methods and http handlers themselves), and wire-format data, and application-level data structures into formalized representations that I can use downstream with autofuzzing/trace/black-box integration testing/etc tools. There are certain things that only make sense to do with AI, that are very hip and envogue, eg black-box or API-surface-informed adversarial integration testing. But you can actually much more cheaply just use AI to set up deterministic tests like this once; the problem up until now is the lack of good end-user tooling and support for those tools, which are hard to write and test especially in real-world applications that don't conform to the specific industries/academic niches that formal verification are already adopted within. That's getting solved too.

weitendorf··on What can you confidently guarantee about your software?
Nothing stops you from writing a formal spec for an implementation or the technical parts of a problem that the product / upstream requirements don't explicitly demand. In fact I would say this is a large part of what software engineering fundamentally is; the reason specs are written in a certain way is because they model a problem domain and encode fundamental (or incidental, which could still be very well-reasoned, eg due to backwards compatibility, limitations of dependencies, or practical constraints) properties that are unlikely to change. Once you've actually written software that other systems/components are depending on, you've in practice created a spec. And it's a false dichotomy to think that just because there is no formalized spec for something before it's implemented at all, that this means that once the domain is better understood or stabilizied, you cannot use formal verification on the de-facto spec that has been arrived at/de-jure rules and structure that are prescribed but not yet formalized.

Even when you are writing something for the first time, like a parser for a certain kind of data format, you often can greatly accelerate your production-safety/security maturity by formalizing your spec enough to be fuzzed and to use existing parsing tooling that avoids the footguns most developers will introduce writing their own parsers from scratch. This is IMO the most common situation in which formal verification becomes very useful; even if the data format is purely an implementaton detail, any time parsing untrusted/user-provided input is involved, it would be much more foolish to try to just throw together some random string processing functions than to use something like flex/bison, an existing serialization format/protocol, or a metaparser (which may themselves be formally verified).

weitendorf··on What can you confidently guarantee about your software?
I endorse this, and IME find there are major benefits to “relaxing” even moreso through reductions or decompositions into well-known models rather than defining your own.

A lot of problems in practice are just too complex, too theoretically difficult, or some combination of not-theoretically-interesting-enough for an academic or specialized worker to be attracted/assigned/exposed to it, or just too difficult/expensive/time consuming/unimportant to be worth directly formally modeling. Or, they’re big problem spaces and you want to start where you get the most bang-for-your-buck. A (partial) reduction lets you still realize practical benefits for problems that may not perfectly or entirely be amenable to simple/small models, or which themselves may be “jagged”/gnarly problem domains due to real-world constraints like backwards compatibility.

Reduction to a formal target arise naturally from defining type conversions (likely to a codomain that is either actually a strict superset of the real reduction image, ie overly generalized/non-surjective, or bijective for some partial subset/step). Mathematically or computationally, it may be less than ideal to use an overly broad/flexible/general abstraction to model something more structured. But practically, it lets you use actual software written for the target reduction and apply any results/properties of the more familiar or better studied model.

The less-than-ideal reduction may be easily verifiable/simple and so obviously safe that, in-context, it’s clearly worth choosing the technically-suboptimal/less-elegant/partial approach to free ride off existing formal verification, without needing to wade deep into theoretical territory, than to chase perfection.

For example, if you can model a data format as a Context Free Grammar, or a subset of all context free grammars, a developer could spell out that grammar or grammars and then use existing metaparsing tools to safely process input data rather than write their own parser. And they wouldn't even need the format to be a complete CFG to leverage CFG tooling! Suppose portions are context-sensitive - for example, decoding the payload requires applying some length-prefix extracted from the header - just encode and parse the header format as a CFG with an opaque payload blob suffix with a length equal to eg the max payload/packet/fragment size. Perhaps some part of the header also specifies the payload’s structure in addition to the length: you could have one CFG for reading the header, a simple wrapping CFG for a fixed-length frame containing a header followed by junk (FRAME_HEADER_CFG :=, a CFG for each payload format, and one wrapping/router function that handles the very simple non-cfg step of using the first-pass parsing output to determine the length/format of the second pass.

FRAME_HEADER_CFG := HEADER_CFG {BYTE}

...

Parse(byte[] in) {

  AST frame = ParseCFG(FRAME_HEADER_CFG, in[:FRAME_SIZE]);

  Header h, byte[] p = frame[0], frame[1];

  assert(len(p) >= h.payload_size);

  return ParseCFG(PAYLOAD_CFGS[h.payload_type], in[FRAME_SIZE-len(p):h.payload_size]);
}

Then even though your data format is not technically a CFG, the vast majority of your parsing occurs within a CFG, with only a small part that does (in pseudocode). And you did not have to do or run any formal verifications, or use special model checkers, only encode the CFG-compatible subsets of your data format into something like EBNF, then use a compatible formally-verified metaparser, and take a leap of faith that the small amount of non-formally-verified/not-nice-to-model code is safe. That's much more realistic for a developer to do than to implement a formal-verification of a non-cfg format de novo.

weitendorf··on DeepSeek Introduces Vision
I thought this way until I tried it, and the main difference is that when I'm managing tons of agents at once or just reviewing some plan / approving next steps, or need to give quick feedback/ask a simple followup, the voice interface makes me much faster and more likely to continue because it's lower friction (and in many cases that's good, though not all) and can be hands-free.

Actually, my thoughts on this matter changed so much that it inspired me to get much more into voice controls because I realized how this same problem was basically why some people sucked at remote work or weren't able to properly use tools like claude code, because it was essentially the same problem but worse (typing / messaging feeling too high-friction or raising the barrier for participation). I have a way to let Claude call me now to tell me stuff when I have a bunch of instances out doing stuff and then leave to go home.

I'm trying to get that better integrated in my devloop because I think it makes managing >4 agents simultaneously much more feasible and natural for some people (I used to play Starcraft a lot so I'm used to the multitasking, but it still takes sustained willpower to be constantly "driving" or monitoring things, or to field questions), especially ones who have never served as TLs or people managers before. IMO it's a big performance roadblock for a lot of developers to be treat directing multiple agents simultaneously as some kind of high-stakes/high-cost thing. The kind of developer who would not say anything in a team meeting unless prompted or who thinks everything is stupid by default (because they are afraid of making decisions / being wrong even if only briefly) is both very common and reluctant to work this way, but also really probably needs it to be as productive as more skilled developers.

weitendorf··on Local Qwen isn't a worse Opus, it's a different tool
Sure, my company has been working on a broad swathe of infrastructure projects and developer tools, which requires prompting models to seek out other tools/apis/docs/examples but in a way where we can't just dump all the context on the model up front. We also need the models to oftentimes look up technical documentation and specs, and sometimes build custom parsers for specific documentation websites that only make the data available embedded across 200+ pages of html.

First, I almost always try to seed every new project or context/domain with canonical technical specifications or examples I found elsewhere. When I set up this project recently, I linked to a bunch of the official Apple docs for sysctl, and told it to use a specific technique for calling assembly code from Go, that from experience it almost never realizes it can do or knows about (and similarly for sysctl, I knew it kinda sorta knew about it, but not in its entirely): https://github.com/accretional/sysctl/commit/da52438233e5b33...

The other thing I did was tell it to enumerate all the test cases ahead of time rather than to just directly implement them; again this is something where you have to explicitly tell it to go digging for information where it has blind spots and get it to set up properly grounded self-eval in a way that it can test against. I usually tell it to take notes as it works or commit notes to itself that will persist over sessions: https://github.com/accretional/sysctl/blob/main/FINDINGS_2.m...

Once we get back to working on this project we'll just have it implement / validate the rest of the sysctl feature support against the full inventory we had it uncover: https://github.com/accretional/sysctl/blob/main/cmd/darwin-n...

Another thing we do is have it specify an API that it can produce against; then in other projects we have them consume the API via reflection (and our special sauce we've been working on is the ability to discover and integrate against these automatically across thousands of APIs from many providers, which we've got working and can share if you're interested in using it as an early customer): https://github.com/accretional/sysctl/blob/main/proto/sysctl... This isn't the greatest example because it doesn't actually fully specify the sysctl keys yet. But I did have it create a knowledge base trying to cover the 1000+ keys as best as it could, to reference as it continued: https://github.com/accretional/sysctl/tree/main/macos-sysctl...

We have a better example in eg https://github.com/accretional/proto-sqlite/tree/main/lang where we were able to encode the entire sqlite grammar into a grpc interface so that you could eg find the exact structure (and sanitize) of a select statement: https://github.com/accretional/proto-sqlite/blob/main/lang/p... This way integration and discovery becomes a matter of telling it "use reflection against this endpoint to discover the sql interface, then implement against it" and we can model formats/input validation as formal grammars via EBNF (all magic words) vs just adhoc

We also tell it to set up and use a browser automation toolkit/testing and always run it at the end of testing workflows (often in a way that auto-opens screenshots on our local machines + commits them to git) via tools like https://github.com/accretional/chromerpc#headlessbrowser-aut... so that whenever we produce UIs it can evaluate its own output and iterate without direct human intervention. This is another case where the knowledge-discovery problem becomes a problem so we tell the models to use reflection to discover the browser automation apis. That ends up giving us things like this where it records user journeys through sites and creates visualizations without us having to debug them or do them ourselves: https://github.com/accretional/proto-css/tree/main/chrome-te...

weitendorf··on Local Qwen isn't a worse Opus, it's a different tool
While I agree that the harness is part of it, I think it's also a lack of epistemic understanding or awareness for what it means to actually solve a problem vs just get something kinda working; maybe if Claude Code or other harnesses made web search more likely or had a better way to make technical documentation and specs available to models, it would be better solvable there.

I often tell it to stop asking me and just keep going until it accomplishes X task; unfortunately it tends to assume I want something that only just barely works, in the sense that it means it's time to stop once its there, which is I don't think a harness by itself could easily address (ultimately the model itself needs to determine the stopping points unless I literally specify by hand hidden evaluation criteria).

That's why think it's at least partially a training issue where the model gets rewarded for "solving" the problem within a certain amount of context/time without access to grounded knowledge (eg looking up the actual spec for a format) nor adversarially/rigorously evaluated against a reviewer capable of finding all the edge cases/shortcuts preventing something from being a properly generalized solution. I don't want it to ask me for guidance when it's working on a well-specified problem, I want it to either find the right parser and use it, or to completely implement one against the spec, rather than write some half-assed string inserter that eg only works on the specific select statements my examples use right now. My understanding is that the Mythos/Fable models were better trained for this but from my brief foray into using Fable for work I wasn't that impressed. For me they need to get better at agentic search and self-eval still

weitendorf··on Local Qwen isn't a worse Opus, it's a different tool
The parsing thing, or the willingness to instantly drop into janky unsanitized string manipulations, or to constantly push back against work on infra projects because some random package on GitHub has 200 stars so it’s totally the safer approach, is driving me insane.

On one hand I’m glad Anthropic is only just now starting to get into infrastructure because it means there’s opportunity there, but it’d be great for their models to be more knowledgeable or able to seek out that knowledge on their own, or for the UX of Claude code to be more amenable to launching 5 in parallel and picking the best one, so I don’t have to spend time arguing with a robot. I think there’s a much better balance to strike between just charging ahead towards the goal at all costs vs being lazy and pushing everything back up to the user. Basically they write too much code that’s too contingent/brittle outside its exact current context and don’t do a good job distilling out the essence of the problem “cleanly”. Almost all of them are like this right now, it’s partially a problem with long-range planning but I think a real bias from over optimization for certain RLVR outcomes vs others.

weitendorf··on Local Qwen isn't a worse Opus, it's a different tool
One thing I used to test quite a lot was rerunning the exact same prompt on the same input, or semantically equivalent (in my mind) but differently framed or worded input, and seeing how much they diverged. In particular I’ve done this quite a lot between Sonnet vs Opus and across Qwen models.

I recommend everybody do this because you don’t need any special data except what you are already using, and the results will be very eye opening: there is WAY more randomness or instability involved than you would otherwise assume. A lot of what you might think is a better prompt technique, or a particularly good or bad outcome, could just as well be random chance or just different behaviors across model version or sizes. And your results can be massively biased by small differences in input. We’ve been calling some of these “magic words” at work, specific technical terms or references/techniques that you need only mention to get vast improvements in outcome.

There’s a skill to it. With agentic loops if you get the model into a self-eval structure where it’s hard to cheat or take shortcuts, and it’s in the right structure or domain that models its training, you’re golden. But it’s hard to find the sweet spots (pro tip, have Opus 4.8 convert PyTorch models into ONNX or quants or get them running on different hardware, I swear it was like I activated some kind of savant-like skillset; meanwhile I can’t for the life of me get it to properly write/test EBNF formalizations of common languages and formats without cheating).

The worst part is that it changes so much so frequently that it’s almost useless to really go digging for this kind of knowledge unless you’re actually the one training the models. I wish this kind of “stability” in output was more emphasized in their training so they’d be predictable. I assume it’s hard to do without overfitting or breaking the explore-exploit loop but also, I would spend so much more on LLMs for batch workloads if they could do them more reliably…

weitendorf··on The founder's playbook: Building an AI-native startup
I think you either need to split your time in between 100% product and 100% marketing or have a founding team that is able to do both - not just “marketing” but the ability to build and leverage a personal + corporate brand and get specific influential people to work with you on social media and for PR.

The “build in public” thing sort of works if your customer base is other aspirational founder-developers but also has saturated and become less useful. I think the stakes for what constitute a credible and trustworthy product are higher now that AI has lowered the bar for new entrants to get “something” working; in a way it’s almost discrediting now to market something not-quite-ready because randoms on reddit can just as easily do that, and it signals something very different now than it used to.

There’s also a bias towards noticing the noisy advertisers a lot and not how many other startups (some of which are quite large) mostly market through different channels or specialized buyers, and so assuming they don’t exist. HN doesn’t realize how many deeper tech companies there actually are because of that; generally IMO the bleeding edge SV technology companies are about 1-2y ahead of what you see as a product consumer/commentariat.

weitendorf··on TorchCodec 0.14: HDR Video Decoding for CPU and CUDA, and Fast Wav Decoder
> TorchCodec now has a dedicated WavDecoder for decoding WAV files. It bypasses FFmpeg entirely and reads WAV data directly, resulting in significantly faster decoding.

I'm working in this area recently and very keen to use this given the claimed performance benefits, but I tried all your links and didn't see any actual performance numbers. Do you have any to share?

IMO a fair performance benchmark for those not tied to the full pytorch stack would have ffmpeg and the wav already loaded into memory before execution. Given that torchcodec relies on the user-supplied ffmpeg installation I suspect that may not be the case for ffmpeg already, at least not by default.

I understand why meta wouldn't want to do this (then you are inevitably distributing exploitable security vulnerabilities in pytorch, because ffmpeg will probably always have them) but I've been statically linking fmpeg and keeping the binary in-memory while still using separate processes for different batches of audio, with I/O through UDS between the parent and ffmpeg; then the parent does VAD on the pcm on CPU before any further inference. My implementation for static linking is similar to the pattern in https://github.com/amenzhinsky/go-memexec#static-binary - would be interesting to see if this is possible in the pytorch/python ecosystem, or maybe it's already been done.

weitendorf··on Did Anthropic ask for this?
Work on distributed/federated learning. The main doomsday scenario is exactly what Anthropic is enabling.

RSI within a top-down, concentrated power structure creates an unstable equilibrium and will instantly devolve into an epic power struggle. If you control the company/computers/military with the one most powerful thing, you control everything else. It will probably just be seized but then you instantly become a target to all the other dispossessed/scared/opportunistic people who think they would do a better job than you, or just want it. It’s a stupid power fantasy to think that people with guns and will just let you reign from above as some kind of benevolent researcher-king with absolute, unaccountable control over the economy. Even worse if there’s mass joblessness and nothing else keeping most people busy.

If we can build a horizontal, federated (not the performative kind like Bluesky, has to actually be performance competitive and distributed without jank) intelligence on a mix of commodity and specialized hardware - which BTW is exactly what even Anthropic would have to work with too, a bunch of datacenter gpus of different generations + traditional compute + edge compute/network devices, maybe some ASICs, then finally consumer devices and webgpu - then there is much less risk that AI will be used to concentrate power or amplify bad actors (without their actions being immediately reacted against by the overwhelmingly larger set of good or neutral-with-something-to-lose actors).

The main barrier to federated learning is figuring out how to economically structure this, it has to have a self-funding mechanism that is hopefully more grounded in actual value than something like crypto where it’s purely forward-looking demand (/speculation) or artificial/enforced scarcity. Also it obviously has to be secure, but the risks are different vs companies like Anthropic that are trying to guard their IP - in this case it’s mostly just protection from bad actors trying to pollute training data in a way that would only be noticed after it’s expensive to fork away from, plus just generally using it as a malware distribution or data collection mechanism.

weitendorf··on Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
Unrelated but I’ve been putting off learning about post-abliteration technique and want to use it for an upcoming open source “retraining” project I have on my backlog. I’m not interested in the refusal layers though, more like deep fine tuning but in a way that might let me prune out or consolidate layers, if that makes sense? Do you have any pointers or links to the current SOTA in this area?

I guess I’m looking for a kind of bulk/sticky dropout (which was in fashion way back when I studied DNN in school).

weitendorf··on Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
Yes, this is the problem. They are business interests of Anthropic and have nothing to do with “safety”
weitendorf··on I design with Claude more than Figma now
You should note that Claude Design is most likely a DPO->PPO->Actor-Critic bootstrap play: https://arxiv.org/abs/2305.18290 / https://en.wikipedia.org/wiki/Proximal_policy_optimization / https://spinningup.openai.com/en/latest/algorithms/sac.html

It's much harder to RL out design taste because it's not self-grounding, and human labelers have no real skin in the game, so this (having a human with a vested outcome in the process directing a model's work) is the best way to get LLMs better at design/"taste"/aesthetic judgment themselves. We were working on the same thing 7 months ago and then I realized that winning over designers to do this would be a huge uphill battle setting up an inevitable fall from grace later on.

What makes me most suspicious of Claude Design is that when you disconnect and reconnect later, it loses context and nags you that the product doesn't work like that. Bullshit. It's at best an anti-abuse/implementation detail (to keep you from launching 10 at once and coming back to them later) or product shortcoming that just so happens to be optimized for keeping you from continuing your design in better tools than theirs for the inevitable followups.

It's great for one shots and it makes sense when you're trying to build a vertical product development stack like Anthropic but I'm disappointed it feels more like a tool optimized for keeping you in their product than for what you're working on. If a company other than Anthropic had shipped this - it's not that hard to build a visual self-eval loop, just use Chrome Devtools Protocol to run headless chrome and take screenshots -> feed into a judge LLM for feedback -> continue - I don't think it would really have seen much adoption.

That said, AI trained on Actor-Critic with a tight human feedback loop definitely seems like the right approach to solving the problem, just not something I want to spend my time training for someone else unless I can do so with higher "entropy" ie high parallelism/optionality

weitendorf··on Three of our worst VC stories
It’s not cheaper to run Claude in your own GPUs rather than the $200/mo for certain workloads. For a large portion of what I work on, the bottleneck is my time, not tokens. You certainly could throw more tokens at it but if you need it to work a certain way for certain reasons, and your plan/goals are beyond the scope of what the top-capability models can do, then throwing them at the problem just bogs you down in extra cruft or reviews/iteration that you could more effectively do being the primary driver of the work.
weitendorf··on Three of our worst VC stories
This is what I’ve been doing. I’m not even against external funding, I just see it as instrumental to the ultimate goal of building a sustainable business. Venture capital is basically a super high interest loan so it’s only something you want to take when and where it can be effectively deployed.

Most other founders/business owners and investors I know don’t see that as a controversial statement (that it makes no sense to see access to capital through a scarcity mindset, when it is factually accessible) but because most customers or potential employers aren’t one of those, it’s been a problem because this isn’t what you’re “supposed to do” and so they read into it from a social/legitimacy angle.

Regardless, it’s quite rewarding IMO and I highly suggest it to people with the means to pursue that path. I don’t see why people get so worked up on the whole VC thing, at the end of the day it’s just lending. Having dabbled in angel investing on the end you get a lot of people lining up for what they clearly perceive as unsecured loans or a social signal and on the other so many people get caught up in the dynamic of dangling said unsecured loans (when I first started it felt like some of the investors reaching out to me were just doing it to boss me around or something, like sir you can clearly see I started this two months ago and you dm’d me on LinkedIn to chat, I don’t need your money).

IMO most good founders/investors are credibility-maxxing but because of the social dynamics and moral hazards inherent with spending OPM you get weird other behavior

weitendorf··on Claude Opus 4.8
There are basically two tiers of "Chinese models" in this context, the "edge" sized ones with ~30B parameters or less, and the big ~1T models that can basically only run in the datacenter.

I don't think it's as simple as saying China's hosting is subsidized, they have generally cheaper electricity and labor costs than in the US and don't have access to the top tier models, and a large internal market where the big models are the best thing they can run with what they have. So obviously they max out on their top models (which are trained with their hardware market in mind, not ours) and get the economy of scale from that, and can run generally the same hardware for less money than in the US because

The edge models are very cheap to run and can do so on inexpensive hardware. They are like 95% cheaper to run than Haiku, so the math is in their favor for certain batch workloads. Most people just run the models for themselves when they do that without making it available on openrouter or whatever, because you can just provision a gpu node and use it as needed, and it's not that expensive to run this family of models.

Is your problem that you want to call Chinese models hosted in the US because you're worried about the data handling?

weitendorf··on Claude Opus 4.8
It's through my startup, so both I guess. Generally I find my bottleneck to be attention and focus, and the opportunity cost of not going back to work at my prior employers absolutely dwarfs the amount of money I spend on tools, so it's not hard for me to justify spending $200/mo on something I use every day that makes me more productive and generally removes bullshit from my life.

At my prior job there was still what felt like a strong enough correlation between my actual performance and my pay that I don't think I would have had a hard time justifying the expense there either; now I absolutely don't. With the current state of the models, it's baffling to me to hear about professional software developers planning their work around their $20/mo subscription's quotas.

Obviously it's more complicated than more tokens = more productive, but I see them less like SaaS and more like gasoline, where if I run out or need more to do what I'm doing, as long as I'm not being wasteful, I just buy more. Why would I waste a day walking 30 miles by foot when I can just pay $5 for gasoline and drive?

weitendorf··on Claude Opus 4.8
I pretty strongly feel the opposite way. Granted I have not used deepseek enough to “know” their model idiosyncrasies as well as Anthropic, so there is a partial skill issue. But I just find it really hard to justify using a less powerful model while I work.

The most I’ve ever spent in a month extra on API tokens for my own work is $200, and I pay for the $200/mo Claude. I use these models quite a lot, though not idly (I usually just walk around and do other stuff until I know how im going to approach the next set of problems). So it costs me about $3000/year to get as much as I want of the best model available. Already that seems low enough to not be worth stressing out too much about optimizing it, because it feels like an indisputable good value, and trying to save money with a less powerful model would be optimizing for a $1000-$2000 saving at the expense of a large portion of my work taking longer or being more frustrating and iterative.

That’s not a flex or anything, I get that in other countries $3000/yr is a lot of money for a software developer and also a lot of people would perhaps rationally be better off doing X% worse at work or spending Y% more time on tasks to save $Z, if their productivity improvements didn’t translate to more salary. Otherwise if your performance has more upside I really do think that the smartest models are better with the current pricing scheme. Deepseek and the other Chinese models spend a LOT of time thinking, and tend to be much more jagged (benchmaxxed) in performance. How can dealing with that over an entire year be worth $2k?

The only situation I can think of where sacrificing my own time/performance to save on inference is batch compute (of course, $1k vs $100k is different from $30 vs $3k) or work where the tier 2 models have crossed the “good enough” threshold. But I think Opus is not even close to that threshold generally yet. As it gets smarter I, and I think most others probably, just try to do harder things faster and hit the next wall.

weitendorf··on Why the smart home bubble popped
Working on it, they all already expose a tool interface, you just have to know where to look for it and how to use it!
weitendorf··on Magnifica Humanitas
A formative moment for me was reading Richard Stallman's writing on the GNU website and seeing him quote [0] Rabbi Hillel [1]:

"If I am not for myself, who will be for me? If I am only for myself, what am I? And if not now, when?"

This inspired me to seek out more about Rabbinic Judaism and its theology more deeply, and I found the language and analogies concerning the idea of "repairing the world" (which you referenced, but which I think at first glance aren't necessarily something most people would identify as a specific core doctrinal theme) particularly inspiring [2]. To me it's frankly beautiful and something I recommend anybody interested in metaphysics or ethics/morality looking into; it also ties into the Kabbalah. IMO this aspect of Jewish theology deserves to be more widely known because it's something all of us can learn from.

[0] https://www.gnu.org/gnu/thegnuproject.html

[1] https://en.wikipedia.org/wiki/Hillel_the_Elder

[2] https://en.wikipedia.org/wiki/Tikkun_olam

weitendorf··on Memory has grown to nearly two-thirds of AI chip component costs
It would require much more than a couple of queries per day, I want to basically do bulk ingestion and search/evaluation/integration across tens of thousands of videos and software projects (if it were cheap enough and smart enough). It would basically be setting up and operating a pretty large data ingestion and coding agent pipeline, which I would want to itself be mostly automated.

It’s ok if you don’t want to do the same kind of thing but I find it weird how dismissive so many people get about wanting to use LLMs for large projects, or how anybody who says they’re using them for these kinds of things (I’m doing similar for other stuff) gets challenged on what they’re doing it for.

weitendorf··on Memory has grown to nearly two-thirds of AI chip component costs
I think there is a reasonable basis for taking a gamble that small models capable of fitting on a 32GB card will continue to advance over the next 5 years and eventually approach Gemini Flash 3.5 / Sonnet 4.6 levels of capabilities, which I would consider to be past the threshold of “probably worth the cost and hassle of running 24/7” if the upfront cost of the hardware was palatable.

My use case would primarily be in search, integration, and indexing other software projects with my own, as well as transcription/indexing of interesting video and audio content (eg Dwarkesh interviews) that I don’t have time to watch but want to easily search and apply to my projects, and search/indexing for useful information from things like Linux kernel and security mailing lists. Basically there is a lot of stuff that, if the cost were low enough, I would point a reasonably intelligent AI at to distill out useful information and apply it to my projects, or just cherry pick the interesting things out and surface them to me so I don’t have to wade through all the mundane stuff and man-made slop getting in the way.

weitendorf··on Memory has grown to nearly two-thirds of AI chip component costs
In the long run cloud gaming is inevitable, it’s just more economically efficient for the cost of the hardware required to render graphics to be amortized across consumers and not sit idle when being unused by collocating them with game assets in POPs.

Once enough gaming compute runs at the edge it also allows for more technically advanced games than would currently be economically feasible (but aren’t made mostly for lack of a market/adoption of cloud gaming and the resulting lack of technical know-how). So I think it will stick and probably end up winning over the holdouts, once the cost of rendering the games they want to play with consumer hardware becomes too large to stomach.

← PreviousPage 4 of 15Next →