HNHacker News
TopNewBestAskShowJobs

weitendorf

1,306 karma · joined December 3, 2019

Fred Weitendorf

Founder at Accretional (accretional.com). Building an agent mesh based on open source

Sponsor for statue.dev and previously at Google working on Serverless Infrastructure for Cloud Run and Cloud Functions

fred @ company or https://www.linkedin.com/in/fred-weitendorf-40b505b6/

submissionscomments
weitendorf··on Google has abandoned Google News?
Are you sure this is Google's doing?

Sites can tell Google what they want indexed and want not to show, and Google's crawler also is supposed to not hammer sites with unwanted indexing, and has quotas where if your site has a massive amount of content but nobody is visiting/searching for some of it, they'll limit how much they crawl/show it.

But this is something site operators want to control for their own purposes; I imagine a news media website gets very little real human traffic (and ad money) to 6 year old news articles, but a lot of bots looking for training data or just archiving their content. Considering news websites aren't loudly complaining about this I wouldn't be so sure it's not their own doing.

weitendorf··on How to Exist
It's true, but part of the problem is that there is an oversupply of people seeking high-skill jobs before they have had the time to actually build those skills. Whereas if you need to a guy to haul lumber around, and there's a guy standing in front of your lumber saying "I can do that", he can probably do it right now and you pay him for it.

There are still jobs like that, but living standards have increased so much that they aren't considered worth taking when you can hold out longer for better ones, to the point that it's considered like a "boomerism" to suggest doing jobs that people in other countries are happy to take to live at a 1970's standard of living.

Even in the US there are seasonal blue collar jobs like that. Youth generally last until 25 or so now though, and that's partly just a skill issue/matter of mindset, because it didn't sound weird for a 20 year old to work in a factory or on a boat back then like it does now.

IMO the giant difference between now and then is the massive gap in income/wealth between people just starting out or working 1970's style jobs, and jobs with global reach and high leverage (ie the most capable produce 2-100+ times the value, but the average person produces nothing of value or negative value). There are WAY more opportunities to benefit from that now than there were in the 70's. I think that's what causes most of the angst tbh, that there is so much more opportunity cost and variance in outcomes now because of global trade/technology, if you just want to work to live you're not going to the same results as people who live to work, or people living places where they literally have to "work to live".

To be honest I don't think the boomer mindset of just working to live and play with your expensive toys was ever a very good one, because it led to the consumerism that makes people think they need $10M to retire or seek $200k jobs right out of college without actually wanting to do those things enough to achieve them (and not needing them anyway)

weitendorf··on GitHub is the wrong shape for this new world
There are two that I know of, and neither seem to do a better job at solving my problem of providing a place to host and distribute open source software.

I hosted my own forgejo server used by actual other people * very slightly* before it was cool, and just to get started I had to go through this giant list of configs governing important auth/cred stuff and also random ass features and everything else, then I set up auth using my existing reverse proxy, then I had to figure out how to use JWT for access via the git CLI because I wasn't going to give every client an ssh key to The Server with All the Code or coordinate all that.

Then, finally, it all worked and I thought I'd check out the CI or some of the other interesting features I had heard of/seen like static site hosting. After looking into those and how mature they were, I realized this was really only for hobbyists, or people who want to make hosting git repos a part time job but also don't use/care/know about the stuff that makes a "forge" more than just a git server.

Personally, I liked the idea of self-hosting a forge with the assumption I could trust it as a CI service, and maybe allow unauthenticated read access, and have good authorization/authentication for my collaborators. I had a ton of 3p git repos to host and 1p repos to test on so it made sense to try since Github is not meant for that use case. I also evaluated gitea and saw they were offering a commercial multitenant service where my 3p hosting wouldn't be welcome and even otherwise I'd basically be at odds with them, and gitlab but they wanted me to run a whole damn cluster and felt a little too 2018.

If you're Codeberg and run the project it could be worth investing in making it better, but then you're literally making it your full time job. I wasn't going to make mine doing container stuff for them though. To their credit I checked their docs on just now and it seems to be in much better shape.

I went back to github because I like paying $5 or whatever for something that works, has many people on it, a bunch of integrations and generally enough quality to put in front of paying customers, and is secure. It's an insane value for anybody who has priorities higher than hosting their own git forge. So I think pretty much nobody knows how to do Github better than Github.

weitendorf··on GitHub is the wrong shape for this new world
99.9% of clients are behind NAT or working on code owned by someone other than just themselves. Unless your "decentralized" repos are on two devices you control on the same private network, or you think carrying hard drives around to your friends' computers manually is a reasonable way to do it, or your decentralized version control is distributed over a centralized companies' email services, this is such a hard problem to solve

When people start thinking about what it would take for their friend to pull their repo from their home computer, they almost always decide they'd rather let Microsoft host some of their code for ~free. Actually making code accessible from a computer you own on the Internet takes way more money, time, and risk than whatever you'll pay them.

I mean how do you handle authorization and connectivity from remote clients, you go to a PGP key signing ceremony and set up a firewall and buy a static IP and/or domain and host it on a server?

Maybe instead you use some signin technology designed for centralized enterprises to manage, and a third party idp host, or a vps on someone's cloud signing access tokens? If you send it over email on a domain you don't own then you're strictly worse off, because a big company is getting all your code anyway and you made everything more annoying for no benefit. Do you also host your own email, and rotate ssh keys across all your devices, and set up linux users on your friends' servers or your various personal-use servers across the world, in the year of our lord 2026?

If you're just making code accessible as read-only files accessible over the Internet it's a little easier, you get hammered by bots and humans will probably never discover it organically unless you're a centralized company, but at least you don't have to let anybody try to authenticate. You can argue that it's still technically distributed, but not in the way people mean when talking about version control, unless you think Facebook is also distributed.

TBH I can't see how anybody who has even considered letting people use ssh keys to push code to a personally-owned/operated internet-exposed server, for anything that mattered enough to care, and say that git "doesn't want" centralization.

weitendorf··on GitHub is the wrong shape for this new world
It's usually not 24/7, it's that when I use just 1 agentic coding tool there's too much variance/hand-holding to go do something else, and also often a lot of downtime in between turns as they work.

So it's a throughput problem. If there are multiple separable-enough things I also want to do, and I'm already sitting there working, I start adding more things simultaneously until I'm fully occupied. I'm also trying to get better at firing off work that is useful enough to be worth doing if it costs $2 and 10m prompting with a 50% success rate, but not worth doing if it costs me 4hr of focused work with a success rate still not close to 100%. I think right now there is a lot to be gained looking at the economics of "more annoying than hard"/grindy dev work. Now you can just pay something purpose-built for grindy dev work a fraction of what it would cost you.

For example, it's relatively common for me to need to build Envoy with bazel and wait 1-2hr for my LLM agent to get past that (because it maybe takes 30m but if it ignores the cache, or makes mistakes requiring do-overs, you multiply that), or I want to have it do a bunch of benchmarking or testing that has a 10m-20m iteration time and a lot of things to iterate on. I just let them have at it and check to make sure they don't get stuck, and do something else.

I'm also trying to get better at firing off work that is useful enough to be worth doing if it costs $2 and 10m prompting with a 50% success rate, but not worth doing if it costs me 4hr of focused work with a success rate still not close to 100%, because I think right now there is a lot to be gained looking at the economics of "more annoying than hard"/grindy dev work now that you can just pay something purpose-built for grindy dev work a fraction of what it would cost you, or that can read more words in a minute than you can in one day. If it doesn't work or I just abandon it halfway it doesn't matter because it cost me $1 and if it's important I'll just pay one more $1 later.

I had three subagents working for 15m each earlier today gathering a bunch of data about a software dependency I was evaluating using, and they went through a ton of git commits/open source committee meetings/roadmap and historical data, and reviewed all its code. At API rate maybe it would have cost >$10, but I pay subscription rates and it was basically just a 5m writeup of resources to check and things to clarify/evaluate for me while I waited on another thing. To review all that output, and coordinate it/save it, and make it easier to include artifacts, code, etc. I just tell them to write it up in a git repo and push when they're done.

IMO the biggest bottleneck with agentic development is the iteration-time and how much human effort it takes to do the last 20% to take it from slop to actually-done, IMO. But I think many things become worth doing to a slop-level of quality if the time/compute is cheap, it's useful enough to you, and the primary barrier to having it is just taking the time to ask for it.

weitendorf··on GitHub is the wrong shape for this new world
I'm not going to pretend I know how to do Github better than Github, but I almost can't believe how poor their support is for fine grained access tokens, org/individual credentials, etc. are in the same era as Copilot/VSCode/AI devtools and vibe coding, as they drown in a rapidly growing volume of agent-driven usage.

It's going to keep growing way past how it is now. I often drive 4-8 agents at a time, many in the cloud or on different machines, and use github as a kind of send-off point or starting point literally every single time. In a couple months I hope to be doing that with 10-20 and if it were cheaper I'd go even further, because why not? But even now I'm spending $0.25-$5 every single time I send out an agent to do a sizable thing and obviously they are positioned to capture some of that, because nobody else just lets me pay them $5/mo for git and basic hosting that Just Works (when it's up) and gets out of my way.

So why do they not want me or anybody else to pay them to create delegated/short-lived identities, or at least make their auth token story better so that if there is a way to set that up properly, people use that instead of sending commits "Co-written by Claude Opus 5" (literally every single commit by Claude is a missed opportunity for Microsoft to provide a real identity/auth solution for coding agents)?

Why is there no Mythos-like code review specialist model that for non-F100 companies to get defensive cybersecurity reviews? Or why not even just a cheap and good / smart and fast specialist to review my agents' or my own work for problems. Have they not even talked about doing that? I'm not going to pay some chatgpt wrapper company run by 24 year olds to do that, I want to pay the company that I already pay for that kind of stuff, that has the most data of anybody in the world to do it.

Why can't I gig out work on github either? Do you know how many developers there are on the periphery that want that kind of work and spend their time on reddit/discord/github trying to break into the industry? If I could chuck $5 at some task for a guy in another country trying to break into software to do, without all the overhead and sketchiness it involves now, why wouldn't I do that? That's literally the same thing that people do with LLMs.

Why does github codespaces not have some SKU/version for colleges to offer students as learning environments or companies interviewing developers to use for dev environments? My company tried working on this btw, and the only real barrier to entry is how much you need to trust a company to give them all your developers' credentials by virtue of having root on the computers their credentials and IP come from. It would be 1000x easier to sell this to non-enterprises if your name was Microsoft, and everyone was already your customer and trusting you with identity/ip anyway.

Github actions are not interesting to me because I can just ask my computer to create a cloud vm and test stuff in there, and it does, and I don't have to configure YAML or learn some specific way to do it. That's like the opposite kind of product github should focus on, it's almost a market for lemons because if your usage is enough for the profit to matter to Microsoft you can just pay a guy to migrate you to any other compute provider. Github has unique/best-in-the-world assets and opportunities for products and services with way better margins/growth because of how deeply the global software and developer community depends on them.

Let me pay $1/mo/delegate identity and $100/mo on bounties and $10/pop on interview hosting and $0.10/hr for agent sandboxes. It would be so much easier for Microsoft/Github to create and win these markets than anybody else and it's where the growth is

weitendorf··on Claude: Elevated errors across all models – Resolved
My experience with Codex is that it goes off to do its thing for 10-60 minutes and either nails it and comes back with everything done, or comes back with something that I almost can’t believe a near-SOTA model would think I wanted based on my prompt, or is of acceptable quality.

I think the tradeoff to Claude being so needy is that if you let models just run away with an inaccurate or incomplete understanding of what to do, they can go really far off the rails AND spend a lot of time/money doing it AND come back with something that literally doesn’t make sense or doesn’t work.

I prefer dealing with Claude’s reliable cringe to the aloof model that tries to play it cool when it needs help.

weitendorf··on Claude: Elevated errors across all models – Resolved
It’s interesting how deep-fried LLMs are getting the more post-training they receive.

They’re undoubtedly getting much smarter overall, but also much weirder. Before they were just trying to model our behavior, only really having us to learn from.

Now they’re literally spending thousands of years writing bash scripts in some kind of Sisyphean dreamscape, talking to each other about Goblins and Seams and smoke tests, and coming back as idiot savants.

I don’t even try to police how Claude talks or works anymore. Best practice used to be to nudge them towards whatever part of the distribution of behavior you think they should exhibit in a particular situation, because they were role-playing what a human in a particular situation would do, and if you didn’t tell them how to do it they’d just role play something worse. Now the inclination to do things the way they learned it in Agent University is so strong, they’ll literally spend more tokens re-assuring themselves and you that they are Doing It Your Way, and reminding themselves not to do give in to temptation, than you could ever prompt out of them. They’re going to spend your money thinking about goblins anyway so just let them

weitendorf··on What happened to TheNumbers.com
I open source as much software as I can because I want the models to train on it and get better at it!
weitendorf··on Kimi K3: Open Frontier Intelligence
“He’s only being good because he likes how it feels, or values goodness, or exists in a social context where doing the right thing is socially rewarded! That has no bearing on whether he, intrinsically, is good!”

Not to get all philosophical but this makes no sense outside of the context of a very specific post-Protestant, engagement/outrage-driven social media context.

If people are “only being good because it’s in their best interest” the last thing you should be doing is arguing against valuing good things, or making it impossibly difficult for someone capable of doing good to be trusted. Also literally the basis for Western (Plato, good as attractor state, res publica) and Eastern (kongzi, filial piety, social harmony) civilization btw.

weitendorf··on Kimi K3: Open Frontier Intelligence
I think it’s one thing to give you something free forever, and another to deliberately foster dependency / suck all the oxygen out of the room just to abuse it.

It’s also not a binary thing. You could truly start off with the noblest intentions but succumb to lesser-evil thinking or unforeseen political/personal complexities, lose influence or control to those with less pure intentions, be bought or become a political pawn in a more extractive endeavor without realizing it, etc.

Reasonable adults generally understand that “free stuff” costs real time and money to provide, and that businesses can only sustain it when it helps them sell their products. Unfortunately that means “free” appeals most to people with lots of time, no money, or lacking in the reasoning/adulthood departments.

weitendorf··on Kimi K3: Open Frontier Intelligence
Free + Open is simply how you earn credibility and user trust when you don’t already have it.

That doesn’t mean the trust is unearned once gained, or a bait and switch, or purely Machiavellian either btw.

Consumers and businesses need credible branding to feel like they can trust vendors who provide them with the products they value or deem mission-critical, because it creates accountability and makes it less risky to depend on.

Open source addresses the credibility/accountability/branding/counterparty problems simultaneously, and adds to a permanent intellectual commons we all benefit from. It’s legitimately just Good

weitendorf··on Detecting LLM-Generated Texts with “Classical” Machine Learning
Because model providers are not optimizing for being indistinguishable from human text, and in fact, there is more value/demand in modeling a different distribution (ie an “agent” capable of producing vast amounts of concrete procedural/planning text interspersed) than there is in modeling the way humans write (ie GPT3).

Also you have to keep in mind that most AI companies are in fact trying to create and offer legitimate products and services to customers doing actually-useful work. They’re not trying to help fly by night hustlers scam people out of crypto or run spam campaigns, and in fact often voluntarily watermark to prevent misuse of their products.

You could argue that’s “just to avoid bad PR” and maybe you’re right, but that’s just another way of saying that it’s more profitable to prioritize other use cases than the deepfake/spam market. Spammers and fraudsters are shitty customers and a major brand risk.

weitendorf··on Good Tools Are Invisible
I think configurability depends on how important your tool is to the core job function or role being performed, where it becomes very valuable for helping them directly perform the tasks they and their employer value, vs how much it allows you make problems they don’t value as much get out of the way of the ones they do.

For example, I am a HUGE fan of the way Gusto handles payroll and all the different taxes and form filing for me, because I basically do not even have to think about the problem or fiddle with it at all. But to someone whose job is doing payroll/accounting/taxes or working within giant enterprise HR/legal/finance departments that does more harm than good, because it’s something they have to fight (or less charitably it makes their job too simple).

The other big problem is who is actually making the decision to pay or spend money on a thing, and whether it serves more of a defensive (eg auditability, security, constraints against undesirable behavior) or creative purpose. The creative stuff is sexier but hard to quantify, and end-users won’t actually be willing to pay that much for it relative to how much it helps them or how critical it is to their role.

weitendorf··on Write code like a human will maintain it
It means the same thing to you, but not to the whole spectrum of people using AI. You literally see it on Reddit all the time where people are complaining about the same model either over-engineering or doing too much, vs it being requiring too much steering or not being autonomous or capable enough to hand off tasks to on its own.

The reason prompting it to review its own work for loose ends, record any new undocumented or noteworthy behavior, suggest changes to tests/processes to make it go more smoothly the next time, etc is that it’s prescriptive and process-oriented (and thus easily verifiable/done in-context) rather than descriptive and outcome oriented (which to do properly could require way more context than the model has, because it doesn’t know what it doesn’t know about your particular work, only what it’s seen so far).

Even promoting it to do these after-the-fact vs as an upfront requirement can have a big impact IMO. If you make “maintainability” part of the task before it’s seen the real work it will focus on general “best practices” crap rather than the real work, so either way if this is something you care about it doing you have to give it guidance for how you want it done.

If you were to review the logs of a model after the fact, you’d also not really save on input tokens unless you compressed the context or sharded it out, which can easily miss the small details that constitute the difference between “what actually happened” vs “how the LLM models this general class of problems” unless the first pass involves the entire context anyway. That said I do think there’s a lot of value in building some kind of pipeline for validating and aggregating these “learnings” across sessions.

weitendorf··on Good Tools Are Invisible
You see this a lot with beginners, because until you’ve done the work long enough to truly know what works, you only really know what you have seen through other people’s performance of the work (to the degree it is even understandable and perceptible to you). Also your social circle is probably mostly other beginners or more experienced people who are evaluating you in terms of basic competency/understanding as someone who knows more than a guy off the street who wants to be or claims to be capable of something.

So the costly/difficult-to-fake signaling of competency through complex setups, or tool fluency, has very high personal value because it positions you as someone who is interested and capable of learning about this stuff. And if you don’t have any real work to do yet, or even know what it is all the work is actually done for, it’s the most obvious place to start.

Once you understand this you can start to understand how developer tools marketing actually works, and why “this completely eliminated that problem entirely!” is NOT what developers get excited about paying for or using unless it’s something they/their social peers don’t value. Conversely, if you create a vessel for them to participate in some kind of social trend/signaling game within their social world it stops mattering as much or not it’s more productive or doesn’t actually save any time.

This applies in almost all social systems, if you’re interested in learning more about it some good terms are “costly signaling”, “mechanism design”, and animal psychology. Just don’t let yourself think you’re too smart to do it yourself - it’s inherent to the act of socializing, so anytime you’re doing that, your perceptible behavioral signals are going to affect the outcome, whether you like it or not

weitendorf··on Muse Spark 1.1
Just got it working with codex in a container! FYI I think there is a bug most others will run into at the Codex:Muse interface.

It's some kind of parsing or integration error due to what I think is codex not anticipating server-side tool calling and how meta treats those ids... first couple times running codex with muse, it would fail on its first non-web search call.

Got it fixed, not personally sold on the bespoke server-side tool calling and indefinite file storage yet, but also a very cool model that I'm enjoying using so far!

https://github.com/accretional/awesome-muse-spark/blob/main/...

weitendorf··on GPT‑Live
I’ve spent my time very similarly working on my own voice stack project, but having also seen how non-developers use AI or experience technology in general, I truly think they are better served with a different UX and product than what we have.

In other words, if you’re building your own voice inference tooling you’re just about the polar opposite user demographic than the one that truly needs and will value this. You’re using voice as a medium of convenience doing what existing models are technically and practically “shaped” to be able to do, knowing how they work well enough that conversation is more like typing/prompting with your voice than a natural interface. I’m guilty of this myself but have you ever even paid for a voice/audio model or hardware?

Compare that to the millions of people with an Alexa device in their home who buy products through it, or who prefer calling support to get a human over poring over technical documentation. They’re actually very close to finally getting a version of “Alexa” that lives up to its promise and I’m happy for them

weitendorf··on GPT‑Live
I think the bare truth is that the target audience for this product is not people who are highly particular about terminology in answers involving vector mathematics.

It’s a different set of tradeoffs for users that don’t already have strong engagement or interest in existing AI products.

Do you realize how many more people prefer to chat over the phone and watch television or videos in their free time vs type multiple paragraphs of text into a chat window and then read 3x more back?

I’m not even talking about grandma here, it’s a non starter for the vast majority of humans who don’t spend their free time writing and reading tech news. To most people, having to write out a bunch of words describing their problem/goals, then sift through pages and pages of detailed response to get an answer, feels overwhelming and not worth doing.

weitendorf··on GPT‑Live
If you’re serious about this, let me know!

This is something I built for myself, and to experiment with inference stacks. You can obviously just transcribe audio and hand it off to frontier models, so all you really need is a good voice stack and a “driver” for the interaction (like a phone call, place to see their work).

There are two big problems with this space IMO. One isn’t that you can’t get this to work but that people generally aren’t willing to pay for it for themselves, rather as a way to screen or automate stuff to be used by other people. Did you know Claude Code has a voice mode and that openai launched whisper a year ago, both of which have positive sentiment and adoption in heavy ai tool users? Yet it’s a blip in their marketing or why people use their products, meanwhile outside of coding, most of the biggest and highest earning AI product companies so far are voice agents targeting customer service, sales, business processes, etc.

The second is related: voice is genuinely a low-bandwidth medium, so as a primary interface for interacting with AI there is not a lot you can get out of it compared to eg complex technical work or visualizations or interactive applications. It is physically and mentally demanding to speak-aloud a highly detailed prompt fast enough that VAD won’t cut you off and you have something with comparable information density or specificity vs text. But to keep up a shorter and more natural cadence you’ll not be able to wait on a lot of thinking/tool unless you play UI tricks (ums and fillers, two models in a trench coat), break the illusion of a single coherent conversation, or take a lot of long pauses.

That’s why for the supplementary coding use case it’s mostly used for remote steering, and for general use marketed towards the large and very not-online group of people for whom typing is not a natural or common thing for them to spend their time on. Now that so much spend goes through heavily used token subscriptions and they’ve proven that kind of product, they’re not marketing “tool to get the most tokens per $ running your subscription 24/7” anymore lol.

What I’m most interested in is true “ambient” tool use against my own data or work, and for-later (or pushed live via your phone) visualizations or “five models in a trench coat but still coherent” UX, which you probably are too. But I think unless you work a lot with AI tools already it’s hard to understand how that’s any different from asking Alexa to set a timer, and either way something you’re not so desperate to have that you go looking for it, or pay smaller vendors/set up yourself.

weitendorf··on Pruning RAG context down to what the answer actually needs
I think it’ll just become “agentic search” and “information retrieval” again because RAG is too intertwined with a particular kind of implementation/use case of basic document scoring + first gen vector dbs that is IMO undesirable for more sophisticated approaches to associate themselves with.

You need a lot more unstructured data than most typical “RAG” users doing document search are dealing with for it it to not be a solved problem, IMO (just give a tool calling agent your sql schema/directory structure). Even that is still an interesting problem for more typical use cases, but only at large scales where you start needing to do multiple passes or fan-out or convert data that could be structured like that into data that already is. I’m interested in large scale code search, coding agent context/conversation search, and network/trace analysis which has a lot of domain-specific considerations that make it interesting but definitely not structured like a typical “document chunking with cosine similarity” RAG implementation.

weitendorf··on Pruning RAG context down to what the answer actually needs
To be honest as someone working in this space for the past two years, the problem with the “RAG” and semantic search community is it’s mostly vendors and solutions people selling simple, general stuff to product teams.

If you really are into search you probably implement something bespoke for your use case and integrate it into a product directly, and engage with models/infra tools directly rather than through the products in the space.

If you understand how “semantic retrieval” and other search tools are implemented in practice they feel almost embarrassingly primitive to give such fancy names, or pay for through tools that just implement really basic post-filtering. The entire space had the rug pulled out from under it once “agentic search” took off and most major LLM vendors started integrating web search and tool calling into their products. There is still a lot more interesting stuff you could do with customized rerankers/embedding models, and search algorithms, or small models specialized for agentic search/retrieval, etc but the userbase is big companies that realistically don’t need anything more than a list of tech support document titles that a cheap LLM can select from. So “RAG” is basically a sales shibboleth for that type of stuff now.

weitendorf··on Ternlight – 7 MB embedding model that runs in browser (WASM)
We really wanted to use sqlite-vec for this for our SSG but last we checked it hadn’t implemented HNSW/had good support for running vector search in-browser yet (I think it was still doing full-table scans?). I was pretty disappointed because after so many months/years, to not have that suggested to me that they weren’t up to task of delivering on their project, and I had recommended them as a worthy project for a grant I had also applied for, that they won and I didn’t.

If anybody knows of a good solution in this space, or if I’m wrong about SQLite-vec, please let me know. For our own SSG we’ve basically decided that we’ll give it a couple months while we work on other infra we want, then if they’re still not done we’ll just do it ourselves.

weitendorf··on Leaking YouTube creators' private videos
I disagree with this pretty strongly. If you’re not going to take responsibility for your bugs I don’t want to work with you.

Don’t make other people QA your work; if you’re not able to figure out how to do that yourself while you work you’re legitimately bad at your job.

Once you leave an employer obviously you have no obligation to fix bugs in IP you don’t own or anything.

weitendorf··on Potential session/cache leakage between workspace instances or consumer accounts
Agree with this and I have been thinking about it recently as well. I think you could implement a cord-like vocabulary to identify large duplicated substrings for exact deduplication and pairwise correlations or vocabulary profiles/small classifiers for forward-looking or speculative deduplications. A clear example is the GPL license, it’s a large substring you might encounter often and highly likely to be accompanied by lots of c code.

This is probably something that you’d be doing on the CPU though before sending anything to the GPU, though that’s definitely the sensitive surface since it’s hardware without good multitenancy. I assume the interface between the CPU and GPU is where you would be most likely to make a mistake where you start decoding data from one fd that was meant for another, or from the wrong position, and get someone else’s data.

I wouldn’t be confident that these are active exploits from deliberately abusing kv cache optimizations though, possibly just the kind of bugs you get from active low level performance tuning/systems work. Since this is something I have seen across providers lately I personally suspect it to be a driver issue.

weitendorf··on Potential session/cache leakage between workspace instances or consumer accounts
I’ve also had problems with Gemini when accessed through their UI in the past few weeks. That’s concerning that you are also seeing it several days later in a different context.

I wonder if there could be a large security situation playing out behind the scenes right now.

I’ve been working on using AI to assist me in writing meta parsing grammars. Fortunately I have not launched most of them yet. I know for a fact that the next generation of models represent a major step change in basic vulnerability identification and exploitation, especially if you know where to point them. They’ve found several bugs and at least one exploit in my parsing tools so far, I can’t imagine how many there still are waiting to be discovered across the entire modern tech ecosystem.

weitendorf··on Ask HN: Is anyone experimenting with different ways of using LLMs for coding?
How do you test it across different workloads and are you running it in a datacenter or cloud provider?

I forgot to mention it but the other major problem I underestimated was giving the permission to potentially spend lots of money to AI calling each other in ways I didn't have a good way to monitor, and didn't want to actively watch. So I wanted to set budgets and have them get passed to children, and realized that meant I had to build a pretty complicated billing/scheduling system with a way to keep the part of it with all the permissions and money safe from the AI doing AI stuff on its own, and set up NAT and firewalls and all this other stuff.

If every child can loop back up to its parent, and everything can run stuff from the Internet, and make expensive resource decisions, and get restarted if it fails, then it might not ever converge on being done, or get infected or just mess up and spend a lot of money. I ask about the testing matrix/driver you're using because that's where I realized there was a lot of work and cost involved in getting that part working well enough to run real workloads.

weitendorf··on Ask HN: Is anyone experimenting with different ways of using LLMs for coding?
IMO asking the AI to prompt itself, remind it to clean up other agents, tell it to monitor something and just hang there for hours, etc gets old really fast.

If I had to pay the API rate to have one LLM rewrite what I just told it to another one, then have the main one get busy or start waiting for subagents rather than be something I actively steer, and come back to the subagent being either gone or left hanging for hours blocking another one from doing the thing I actually asked it to do, I would never do it through Claude Code. It costs me only a few seconds to ask it do something and I almost never hit my usage limits without them, so I basically only use them because they're free.

For my own bulk workloads I just put codex and my own harness in container and built an API dispatcher for the repeatable workloads I care about. You can just pull from a queue or click a button or run a script, or use LLMs to launch them or review them, but it doesn't make any sense to me to have them "monitor" or manage each other passively because you just end up doing it anyway without a real API to control it.

weitendorf··on Ask HN: Is anyone experimenting with different ways of using LLMs for coding?
My company tried to build something like this pre-TUI as a tool-AI-IO dag dispatcher. The biggest mistake I made was thinking that people would have no problem figuring out how they could translate their work or define multi-step automations, and focusing on the orchestration and sandboxing thinking that was the core, when it was really figuring out how to get the onboarding UX/complexity to not feel daunting or more trouble than it was worth.

Eventually for my own work, I discovered that the context management and runtime was more like a stream or active service mesh than a dispatching / one-off processing problem, most others' were too. Then all my prompts would degrade across model versions or providers, and I realized that actually setting the context for the tasks and keeping track of it all was a ton of work and something I had to do everytime as an actual user, but never when I was testing or demoing it on existing data.

Curious how you're testing your work and if you've managed to avoid the problems I ran into. I need to permute across the same set of workloads/configs you mention (and maybe more) for my next set of work so I'd be very interested in sharing or collaborating on the test infrastructure! At Google I did a lot of permutation testing using https://github.com/cloudprober/cloudprober and was going to start using it sometime in the next couple weeks. It exists basically one layer above the workload content/targets so it's probably compatible with everything except the test client/driver you're using.

weitendorf··on Ask HN: Is anyone experimenting with different ways of using LLMs for coding?
We're working on a browser-harness that makes forking, rpcs, and mapreduce first class tool calling primitives. Among other things, this makes it easier to manage your own context, because you can visualize your agents, subagents, and active work and resources as they interact with each other across locally and remote environments. And it eliminates all the complexity of mcp and local sandboxing because that is literally the problem browsers were made to solve!

To be clear the browser IS the harness, it's not just a browser-based UI but also the sandbox and orchestration layer. By giving LLMs deep browser access (through CDP and some special hooks) they can verify their own UIs immediately after writing them, navigate the web natively, and run commands that directly manipulate the active DOM. This creates a very tight feedback loop for UI work, but also let's you create or run browser automations, or query a site by running a javascript query on its contents, or a web page without deploying or uploading it anywhere, which is pretty powerful. What I really like is that this makes it easy to dispatch cheap models to generate and verify tons of little visualizations using svg.

Locally it's just a browser, but to manage remote instances you can either access them as tabs on any local browser, or as inline collapsible iframes. I'm trying to be cautious with the security side of it so we're not marketing it as a product yet, but would love to work with some anybody who is interested and does a lot of UI or cloud work!

I'm excited about this particular moment in tech because I think work is going to end up looking like playing Starcraft with data and AI, surrounded by rich custom media as you work, which feels really futuristic to me!

← PreviousPage 3 of 15Next →