HNHacker News
TopNewBestAskShowJobs

weitendorf

1,306 karma · joined December 3, 2019

Fred Weitendorf

Founder at Accretional (accretional.com). Building an agent mesh based on open source

Sponsor for statue.dev and previously at Google working on Serverless Infrastructure for Cloud Run and Cloud Functions

fred @ company or https://www.linkedin.com/in/fred-weitendorf-40b505b6/

submissionscomments
weitendorf··on Muse – Meta’s personal AI agent
I genuinely think big tech co's top leadership have very sophisticated strategies/positioning that they simply cannot communicate or explicitly canonize due to their position as spokespeople for the company (and society writ large, news media, investors, customers, employees, vendors). Of course there are a lot of bozos and incompetent people flailing around and a lot of work ultimately gets wasted (which understandably bothers line employees a lot), but is inevitable and even necessary to eg hedge product strategy/comp and take risks on ideas and people.

It would not really be useful to have that conversation with employees because it's incredibly distracting (now product strategy is up for debate with way too many cooks in the kitchen), and very few employees have the exposure or skills to meaningfully contribute even if they think they do. I saw it firsthand at Google TGIFs.

Meta's strategy seems to be "personal agents" quite consistently. Remember they tried to buy Manus? And note that Muse Spark and the meta AI platform products clearly seem to prioritize web search, computer/browser use, and vision/language tasks over coding, which is something that had to bake for a long time. Also, this product launch itself is pretty interesting:

1. It's clearly a fast follow to the current FOTM hype startup Instinct with a much more comprehensive implementation and integration with their other agent products.

2. It's kind of like openclaw, which got a lot of non-developers very excited but was basically consistently unusable. Except this agent's compute runs remotely and presumably has slightly more sane development practices. IMO it's the first main openclaw-like product that has made it to the "just works" level of usability.

3. The focus on ecommerce is actually really really important, because Meta is trying to capture the intent/demand-driven purchasing flow that Google currently owns through search. Controlling the top of funnel is what enables google to make hundreds of billions of dollars per year on search ads. Meta is an advertising company and Google's search ads business is the most lucrative and centralized/well-defended advertising market in human history.

So it is actually a really big deal that Meta is trying to go after it (at least, the CUJ, it's possible that they'd monetize the agent-driven UX differently than search ads) because it's probably their best shot at disrupting that market and one of the few growth opportunities that would actually make a dent on their balance sheet. Obviously Meta is not going to lay out all that strategy stuff explicitly because for all intents and purposes it's a distraction and shifts the conversation in an unproductive direction (the strat behind the product, rather than the product itself). But it's pretty clear if you look for it.

Edit: Actually I thought about the advertising business more and I think in the short term this is partially about attribution/conversion. In 2022 the Apple tracking changes cost Meta $10B in lost conversion metrics; an agent-driven and proxied purchasing flow has built-in attribution and funnel measurement which is very valuable in its own right!

weitendorf··on Muse – Meta’s personal AI agent
I’m pretty sure if you asked my mom what problems AI could help solve for her, navigating website’s purchasing/account flows and directly answering questions about the contents of her email would be #1 and #2.

And my mom doesn’t use AI products like chatgpt because she’s not a student, nor agents because she doesn’t work in tech. To her AI, is no different from those chatbots websites pop up in the corner, the ones you never seek out or interact with intentionally, because you don’t need what they offer.

So I actually think this kind of marketing is quite helpful for regular people who mostly just use their phones to buy stuff, looking things up, navigate life, and entertainment. My mom doesn’t give a shit about APIs or sandboxing, or benchmarks and to her it literally is a hassle to manage a million one-off accounts and confirmation codes and receipts when she just wants to buy something on her phone. She would not assume AI is capable of that or know how to set it up locally.

that’s not even getting into the fact that the top 10% controls 50% of consumer spending in the US and primarily purchases convenience + health + experiences, whereas the other 90% primarily prefers to buy aspirational/identity based goods that evoke their mental model of a higher status lifestyle. It’s why the same bustling lifestyle archetype you criticize is literally used exclusively in car ads, housing marketing materials, consumer electronics, home goods, etc. Because normal people want to feel and/or look good not spawn subagents

weitendorf··on Muse – Meta’s personal AI agent
For the most part I feel the same way, and actually think it has the potential to completely change how e-commerce/web browsing work once it plays out, to the benefit of the client/consumer.

But, I can definitely see a fair argument from the e-commerce businesses that stand to lose from that, that by removing their control over the shopping experience in favor of an opaque agent, they lose the ability to effectively communicate with their customers. Or to provide a coherent/smooth purchasing flow where important stuff like price/dates/shipping are properly surfaced to the user.

That argument would I think be hard to separate from the desire to corral customers into their marketing brochure and convert site visitors into purchasers without churn. But it is true that the LLM in the middle could consistently miss things, or have weird biases/preferences that force vendors to rebuild their sites around LLM tics instead of actual people, janky harnesses that go unnoticed by end users but cause missed sales due to stale data or blind spots, etc.

Also, demoting their sites to glorified databases with a shipping/fulfillment API will give the buyer agent’s provider a lot of power over them and break their ability to establish branding/repeat-customers. Those are the only ways they can reliably carve out margin that otherwise gets whisked away by the advertising platforms pitting them in a zero-sum competition for placement. So it’s possible it could starve out or kill e-commerce the same way Google’s changes to search (the ai box and inline info) hurt the ad-funded websites their info came from.

weitendorf··on Muse – Meta’s personal AI agent
I think they just want to show you ads and help you buy stuff. The incentive is actually to keep the data to themselves so they can monetize access to it via ads.

The shopping experience is significantly less hostile than Amazon’s 1P digital storefront and I kinda don’t care if fb knows that I want to buy a computer.

The only thing to worry about is that they want you to install a native app. But this UX would be difficult to provide for free via the web due to the obvious abuse potential of giving free access to LLMs + remote compute + browser and tool use. And I think as long as you use it to shop or automate web tasks (and give them access to the top of funnel for customer intent, the $300B/yr thing Google monetizes) they don’t really have any reason to abuse your data.

weitendorf··on Muse – Meta’s personal AI agent
Try asking it to find a good deal for you for <item> across multiple sites. Then tell it to use the browser to navigate through the two most promising sites to confirm pricing and availability, and prepare comprehensive breakdown with its findings.

People shop a lot on the internet, actually. Between that and ads for the things people shop for, it’s pretty much the backbone of the Internet economy.

I bet they would purchase more things than they currently do, and seriously break the economics of Google/Amazon search ads (about $300B of yearly spending just for those two), display ads, and internet-first e-commerce sites if they could just ask an agent working for them to help research/source/purchase things without all the navigation and dark patterns in the middle. Personally I would probably spend at least $1k/yr more on snack/beverage subscriptions alone if it had less friction.

> working class peasant consumer

You mean the majority of people in the world? Most people only their computers to entertain themselves, look things up, buy stuff, and complete tedious tasks (taxes, email, etc) they’d rather not do.

How much do you and your peers spend a year on online shopping and how much do they pay out of pocket for SAAS or AI tokens? And how many of your offline purchases had a significant amount of associated online research involved despite technically completing offline? (Houses, cars, schools, hardware) Yeah that’s pretty much where all the money in the global software industry comes from

weitendorf··on Muse – Meta’s personal AI agent
The inline browser the agent uses that you can take control of or watch is awesome! This is something I built a lot of tools to do in the past year (including a similar inline UX/image + click pass through) but having it Just Work in a remote, fully managed client for free is extremely convenient and useful.

But… this feels like a UX that won’t last once a significant portion of consumers and purchases adopt it.

Either the network traffic is getting proxied through my client (effectively making each user a residential scraper for meta’s crawler) or it’s between meta’s servers and the sites, which puts site operators in a difficult position: if real customers are making purchases through this interface and throttling/blocking meta’s IPs makes you invisible to them and meta’s userbase, you don’t want to block that traffic.

But now every consumer in the world can just ask a question or say “check all these prices and sites for a thing I want” and go do something else, right out of the box for free, and have thousands of page loads and site interactions fire off for them.

Bypassing the branding/marketing funnels or intended UX (cf. vc twitter abusing resy thru instinct) of sites, through some kind of proxy client amplifying the traffic a human would create, with the ability to let anybody scrape or interact directly with a site’s backend… definitely a consumer win, but seems unsustainable.

weitendorf··on Bill Gates tries to install MovieMaker (2003)
How exactly do you think Bill Gates at peak Microsoft is going to fix this given "it's 100% on him?" And do you think the guys on the line are the ones on this email thread?

This is immensely more "ownership"/initiative than the vast majority of middle managers at any tech company would willingly take to solve a minor UX/CUJ failure, that's surely literally why he started the email thread and including the ones he did, to model that behavior for the the other middle management bozos and to try to instill a culture of caring about their work/whether it actually accomplishes the intended goals.

weitendorf··on PostgreSQL 19 Interactive Tour
The way I process this kind of content is by skimming or ctl-Fing for the material I'm interested in reading and usually just reading the example code or the specific explanation for the content I'm after

For example this site ranks on the first page for "go 1.27 generics" and "go 1.27 uuid"[0] and if I were looking for uuid content I'd probably click on the toc for uuid and go to here [1] and look at the examples and v4 vs v7 semantics and then bounce. For this particular article the thing I'd be most interested in applying is probably REPACK [2]

  REPACK (CONCURRENTLY) events;
And all of the content around their code snippet showing that is pretty prescient/semantically dense and useful. What I wouldn't do is read the whole thing front to back, or the prose at the top/bottom with the LLMisms: it's way too long and dense for that.

For comparison here are the official new postgres docs about REPACK [3]. Is it human-written and more informative? Maybe for some people, or for me if I needed to reimplement a postgres-compliant spec or something, but I'd prefer the LLM-assisted (and I say assisted because IME it's actually a decent amount of work to get LLMs to write content like this) article most of the time.

[0] https://www.google.com/search?q=go+1.27+generics

[1] https://victoriametrics.com/blog/go-1-27/#the-uuid-package

[2] https://victoriametrics.com/blog/postgres-19/index.html#repa...

[3] https://www.postgresql.org/docs/19/sql-repack.html

weitendorf··on C Is Not a Low-Level Language (2018)
That's fair. C is very old and used for almost all hardware so I think while you can make the argument that "only clang and gcc extensions asm blocks available like that, and intrinsics are only available through vendor-specific headers" and be right, by that same logic literally nothing except binary machine code for hardware without any kind of microcode can be low-level, and even then it's probably always hardware dependent (because if it's not fully bijective to the actual hardware it's implemented on top of, the semantics leak).

Practically speaking, we have a word for the kind of "abstractionless" model you're describing: machine code. I mean, even assembler is a bunch of abstractions about 'registers' and 'instructions' that are really just specific portions of the hardware or opcodes!

So we either descend endlessly into pedantry arguing that cosmic rays and electron tunnelling represent inexcusable deviations from the overly abstracted semantics that hardware vendors expose in their products or maybe we draw the line somewhere else.

You may not agree with mine, that "practical and simple interop with machine-level language impls across a high-level language interface is sufficiently close to the hardware as to be low level" but there has to be a limit somewhere between that and "technically the hardware's operating temperature is part of its logical semantics because if it exceeds a certain value for long enough it starts to degrade and yield incorrect results or terminate execution". I think eventually it just becomes unproductive nerd sniping, personally

weitendorf··on C Is Not a Low-Level Language (2018)
Sure, but then you're really arguing that the ISA no longer maintains 1:1 instruction-level implementation and that this is the definitive quality of whether or not something is low level, to the point that any deviation from that model makes it not officially "low level". To me that's just a very tedious pedantic argument that simply fails to capture the actual meaning behind why/when we might call something low level.

TFA famously argues that Spectre/Meltdown et al break that abstraction. But note that they are quite literally exceptions to the rule: the only reason we know/care about them is that the "magic under the hood" that was supposed to make CPUs faster while maintaining that abstraction introduced a bug that caused the implementation details to leak to the end users.

Similarly even vp2intersectd took multiple cycles in its original Intel impl and even in the performant AMD Zen5 impl it still takes >1 cycle with 6 levels of pipelining or somesuch. Ok. If literally not even a chip's ISA is "low level" then the term is effectively meaningless.

The only way you could define a "low level" language capable of exercising that hardware's capabilities fully would be to have some kind of per-cycle, pipeline-aware annotation layer over the actual machine code... which really seems like quite a lot of noise/cruft you'd not typically want to add on top of everything, all in the name of still technically being low-level according to some dubiously pedantic criteria nobody would event want in practice.

weitendorf··on C Is Not a Low-Level Language (2018)
This is such a pedantic point IMO. C is low level because it makes it very easy to work with machine language/assembly and do stuff like this (LLM assisted example follows):

  int main() {
    __m512i vecA = _mm512_setr_epi32(0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15);
    __m512i vecB = _mm512_setr_epi32(0,5,10,15,20,25,30,35,40,45,50,55,60,65,70,75);
    unsigned short mask = 0;

    __asm__ (
        "vp2intersectd %[B], %[A], %%k2"
        : "=@cck2" (mask)
        : [A] "v" (vecA), [B] "v" (vecB)
        : "k3"
    );

    printf("Intersection Mask: 0x%04X\n", mask);
    return 0;
  }
This is something "low level" programmers use very often to realize the benefits of a high-level language while exercising explicit control over using specific hardware instructions (vp2intersectd being an AVX-512 instruction used in highly optimized search algorithm impls).

Obviously if you rely on implicit behavior from the compiler to optimize your code you are no longer "low level". But if you can quickly and easily drop into machine-level instructions to provide explicit implementation semantics, and the language indeed makes that relatively simple and easy to do, that sure seems "low level" to me

weitendorf··on PostgreSQL 19 Interactive Tour
Gotta say this site is killing it in the LLM-written technical blog SEO game the last couple years, they're consistently able to dominate (and actual deliver in spite of the LLM-isms) SERP for a set of technologies that very closely align with the ones that a typical backend/infrastructure engineer would work on.

I checked their list of blog articles from the last year and they haven't even written that many, which makes it quite impressive because they really model the set that I've been working with (Go, Postgres, grpc, otel) well by I guess defining some kind of customer archetype and writing really detailed guides about what they'd be interested in learning more about. At this point I've encountered their site "organically" like 5x in the last year and recognize the name/style of content so they're doing something right in the marketing department for sure.

On one hand you could just dismiss it as spam but articles like this actually represent a pretty significant LLM spend/human review element that delivers real value to me as a technical end user looking for info on google or in technical blogs on HN (ie it would take me a long time to generate something like this myself and I wouldn't do it proactively, only when-needed). So it actually does help me quite a bit that they do so before I think to ask about it.

weitendorf··on GLM-5.3-Flash
Strong agree, but I also think some roles in big companies (for me, infrastructure) or in certain industries (eg trading/finance) can help build the same understanding without as much of the variance/raw exposure to bottom line.

Now that the role of the ticket-cruncher is on the path towards full commoditization, and individuals can move much more quickly (and even more carelessly!), I think product roles will probably shift towards one where developers are more deeply embedded in the product/business process so that they own/understand what to build without as much separation between the decision-making and prioritization of what to build. Or at least, they should.

It was eye opening to me to run the math of "should X people work for Y months on this project to save Z per year?" and realize that in so many cases, the time and effort it would cost to stop "wasting" money on things is WAY more than you could actually save on it. Even "small" projects can very quickly become $1M+ investments in time and resources, and the diminishing returns add up quickly (but also a good way to justify the value of your contributions, when done). But the job only exists if it saves money or makes money...

weitendorf··on GLM-5.3-Flash
I have exactly the same opinion

Over the last couple years I’ve had to learn sales and understand the thought process behind this better, and I think I’m beginning to understand it

The psychology is that most people aren’t really trying to optimize for productivity (even most people who think they are) on an ROI basis, because their compensation is too decoupled from their actual raw output, and more closely coupled to how differentiated their marginal contribution is to peers. They’re much more incentivized to spend their personal/work time optimizing for being more skilled or acquiring some kind of competitive advantage relative to baseline.

Most people don’t consciously run the numbers of “I get paid $X/hr to add $Y of value” or model pay at work as something with variable inputs (eg something that can be increased with high performance), so it makes sense to them to spend 20 hours of time to save $100 or to make themselves 5% less efficient to take home 0.5% more or avoid doing something they don’t want to start doing.

NOT saying this always happens or that they’re stupid for doing so. I didn’t even realize how much I had been doing it myself until I started recognizing it, and shifted to having my own comp/performance fully aligned with the company’s P/L.

It actually makes a lot of sense IF you can accurately estimate incremental upside (which is much harder and more diffuse than modeling downside if you’re salaried a employee) or if the upfront skill/knowledge investment that looks like bikeshedding pays off in the long run.

weitendorf··on Protobuf has LSP support
This has been driving me crazy ever since I started using protobuf/grpc and realized more major tech companies (in the cloud/infra/data world at least, and a lot of other SAAS) were using it or something similar (eg capn proto) internally than not.

It feels like we’re in some sort of deadlock where each of them think “proto/grpc are too niche to support for external users, better just use JSON”, keeping it unfamiliar for an Average Web Developer. But if every company using it just exposed it to third parties/added it to their public APIs, it would immediately be common (and trendy) enough for every web developer to learn it and start using it.

If Google added the missing HTTP/2 streaming support to browsers (blocking native bidi grpc streaming) it would have an instant killer app in making it easier to implement websocket-like client/server applications. It makes absolutely no sense that full duplex bidi was added to the HTTP/2 but remains unimplemented in browsers.

The role MCP, OpenAPI, and JSON schema fill all would be a million times simpler if they were based on protobuf instead of JSON. I can forgive OpenAPI/JSON but it honestly pisses me off that we ended up with MCP and JSON-RPC + JSON schema, and people think these are cool/good tools, and actively adopting them. Just piling on the slop

weitendorf··on Anthropic IPO valuation hinges on $190-200B 2028 revenue forecast
Anthropic’s revenue run rate just hit $50B, in a market that didn’t exist 5 years ago and was 10% of its current size 1 year ago. They are, by far, the fastest growing company in human history. The demand for “rockets” has been pretty high and growing ever since they started supplying it.

Anthropic and its investors/customers don’t need random internet commenter’s permission to decide whether or not something is worth doing or justified.

Personally I think when something doesn’t make sense to you, you should try to figure out why it makes sense to other people, and whether they might have different needs/constraints/incentives/skills/knowledge.

Why might people spend more on AI as it gets better, rather than less? Could they, perhaps, allow for entirely new kinds of capabilities and products that hacker news commenters have not yet seen? As they have done every year for 3 straight years (hacker news has been wrong in exactly the same way every time, btw)? Might some people prefer to spend $10/day to work with the most capable AI available, due to the amount they use it while doing their job, and its effect on output? Do some people use AI for more than just tinkering with open source harnesses? Could it be that when you don’t understand something, there really is a way to explain it, that isn’t “everybody else must be stupid”?

weitendorf··on Models Are Getting Dumber on Purpose
You 100% can finetune or adapt/build on top of models, and specialize them or extend their capabilities. That’s literally what post training is.

The problem is that “finetuning” was a 2023 AI FOTM associated with products/demos that were almost exclusively using it for LLM character role-play/output style purposes (ie not in actual systems where they served a more functional role).

This made people think you could train models without replay/real evals by yoloing it with SFT (this is partially an artifact of that era being much heavier on autoregressive training and not so much evals). You really can finetune and get results but you have to treat it like a small ML training run, with real evals, and more intentionality than just “more examples”.

You can find pretrained and -instruct models on huggingface that clearly demonstrate what specialization/staged training runs do.

I’d be very wary of conflating finetuning with specialization/extending a model’s capabilities in general.

weitendorf··on Anthropic IPO valuation hinges on $190-200B 2028 revenue forecast
I walk to work and ride a bicycle in my free time, don’t see why anybody would need a hatchback or semi truck or build rockets. Idiots
weitendorf··on Models Are Getting Dumber on Purpose
This is an oversimplification: more data makes models smarter ceteris paribus, but mostly only because auto-regressive training (where most of the general knowledge comes from) is essentially compressing information that can be recalled later if it’s useful (or not recalled). Obviously there are differences in kind within “more data” too, you would much rather have all the books and blog posts in the world than all the fanfiction.

Yes, all three together would be even better. But it wouldn’t be if you had 100x more fanfiction, mostly synthetic, generated during RL to teach a model to be better at writing fan fiction. There are real limits to the amount of knowledge you can cram into fixed-size (downstream of hardware availability) weights. For a period scaling with data was basically “free” because we had the Internet and all the books/media that humans had already created; the data was accessible and limited (at least, the parts we think models should know about) enough and top-hardware big enough that we could basically compress the whole thing.

Post-training/RL are making this obsolete because they’re more about skill/capability acquisition rather than knowledge. They can generate much more data (most of it quotidian/useless, ie an agent made a typo in batch 382829) and clearly seem to cause a kind of mode collapse even in the most advanced frontier models.

We don’t need to make LLMs forget about SpongeBob SquarePants so they learn more about bash. But if I have a question about SpongeBob SquarePants, I don’t need to hear about load bearing seams prefaced with honest caveats after a model writes 400 lines of bash to look up SpongeBob’s family.

And there is probably a lot more SpongeBob knowledge we could put into models if we wanted to: interviews with the creative staff, a SpongeEnv/SpongeHarness modeling how the art/story team work together to create entertaining kids tv, a SpongeBench measuring entertainment value, etc. If a SpongeAgent spends 2000 years in Agent University learning how to Spongemaxx we probably don’t need or want to have it spend another 2000 years writing smoke tests

weitendorf··on Qwen 3.8 27B
IME this is a strong/reliable model smell, you typically see smarter and less benchmaxxed models' thinking traces spending more time exploring the solution space, and benchmaxxed models more time trying to refine/decide on the response contents. It has always been a big problem with qwen.

In general it's kind of a benchmaxxing/distillation (I don't mean that pejoratively here, read on) artifact where test time compute's purpose is to refine/zero-in on a particular input:output pair. Basically they're trying to "remember"/re-derive the answer rather than arrive at it deductively - it gets comical/absurd when you see it expending a huge amount of tokens trying to figure out the answer to some trivial question or response to some input, as if it were a trick question or the choice of wording was of the utmost importance. But it's not a bad thing when the output matters or the task is hard: basically the model has been trained to respond/act like a much smarter model and it's probably better for it to overgeneralize that behavior.

Also, IME Gemma and most other local models that support thinking tend to have the same problem, and it's partly only a "problem" because they give you access to the thinking tokens themselves (which you don't in many cases see when working with frontier american models) and you can actually see how they're getting spent. For the local or bulk (runnin on owned or rented hardware) use case where you are paying in time/watts rather than purely by the token, IMO it's a good idea to just look at the task success rate / time per task and not worry about the raw thinking token count.

weitendorf··on Why does Opus 5 feel worse to work with?
I think it's much simpler than that, they are indeed trying to counter against sycophancy and models lying (the decision whether or not to lie to the human user, take shortcuts, etc. comes up in their thinking traces pretty commonly IIUC) by rewarding them for giving good-and-bad feedback, fessing up to things, pointing out potential issues, etc.

The problem is just that they are rewarding the behavior shallowly, ie rewarding the appearance of honesty or neutral replies, being highly detailed/thorough, even where it doesn't make sense to do so.

I think this is partly due to a reliance on LLM-as-judge training runs/synthetic data during RL where they're having a model which itself doesn't epistemically understand when this behavior is necessary or valuable influence the feedback provided to the model being trained. And that's mostly a problem of scale/volume and the desire to have a tight feedback loop rather than a safety issue IMO. They just generate an absurd amount of traces during training and the only way to really evaluate/rank/steer them at the scale they're generated is through other models, and combined with some kind of honesty/truthfulness/non-sycophancy eval that isn't robust enough to prevent mode collapse, you get this.

weitendorf··on Timeline of the OpenAI accidental attack against Hugging Face
Frontier labs are not a monolithic entity.

There is a clear self-verification/difficulty ramp in cybersecurity, and it is a very valuable as a skill both offensively and defensively. So it is absolutely certain that someone, somewhere, will use reinforcement learning to make models very good at this, once coding agents exist.

Even if you are only interested in using this defensively in practice, you can’t really understand it without knowing how both sides work. So if you want to defend yourself, you need to train for it (or pay for someone who has).

weitendorf··on Message your other Claude Code sessions
I think what we really need is project/thread-scale continual learning. The problem is that the important parts of the conversation to you are the novel bits you just did, rather than all the context building the agent did to get to the point where it could do the novel bits (and even then, without really understanding the bigger picture).

If you snapshotted at 90% max context you could pretty reliably start iteratively trim that down, I think? I personally try to save the logs so agents can slice and dice them with sed/awk/jq/whatever when they need to look stuff up, because I’d rather pay the penalty on read (when it’s motivated by something) than in write(where you don’t really know what if anything will be needed), and they can figure out what they need on their own.

What I’d rather have is some way to bake history into the actual model weights (the same way it can recite certain literature or historical/factual stuff without context), with like multi-lora / “experts” that get trained out of band. But this is contrary to the “one fat model” approach to scaling and doesn’t work with closed labs’ business/IP models

weitendorf··on Message your other Claude Code sessions
I do this a lot and you have to be really careful to clean these up or qualify/steer agents around them. They’ll often be very emphatically confident about some assumption or implication they made, and if another agent stumbles upon them they’ll get mislead.

They don’t really know what they’re handing off or what you’re trying to actually do, so in a sense it’s not a grounded task for them. Actually, if you think about it, any scenario in which a handoff doc might be valuable is probably almost always better as a subagent thread, because you are paying the same amount of read/write tokens but you can clear things up synchronously.

I’ve found two-way message passing (each get their own write file, they read each others) to work much better because the communication is more grounded in actual coordination/work. You can also give each an inbox so that multiple can write to it. If you do the “progressive disclosure” right it scales subquadratically because they only read/write to others when it’s relevant to what they’re working on.

But IMO “write a handoff” is a trap, as a human you end working in some kind of robot-graffiti codebase full of junk, and it ends up being a booby trap for agents literally within days.

weitendorf··on Message your other Claude Code sessions
Claude Code makes agents reasonably aware of where their log files/history/etc are and get stored. Generally they’ll work with them without explicitly being told (especially to recover broken sub agents, corrupted sessions, etc) to do so.

I think the more general problem is that compaction is just a bandaid: you HAVE to dump context to keep going and searching back for it is more expensive than if you had just kept the right context. The better a job the harness does at filtering out junk, the more likely compaction is to remove context that might have been, forgive me, “load bearing”

IMO the default Claude Code / Codex (which to my understanding is almost continually-compacting?) compaction has got much better over the part few months. If you spam sub agents then context will naturally nest, and you can just resurrect them as needed without polluting the main thread.

weitendorf··on Seedance 2.5
Right now we're going through multiple concurrent but slightly non-overlapping transitions where AI is almost good enough for something, and seeing a predictable surge of spam/bad quality crap people don't like, followed by a whiplash effect when it reaches human or superhuman levels of performance.

I think most developers would agree that the Cursor vibe-code era sentiment towards AI for coding was right in the sense that it wasn't that useful then, and a lot of people did and do stupid things with it, but it really did deliver quite substantially on its promise just a couple years after the "consensus" was often "that's never going to work".

It's very strange to me how consistently public opinion shifts on this problem. In the long run we're all dead but whether it takes one or five years, you probably will be there to live through it. I assume whoever is downvoting you is just thinking with their consumer-brain or "I don't want that" (the way it is now) rather than that this will be no different than video games, or wasn't a child who watched youtubepoop or other weird stuff because it was funny.

weitendorf··on Seedance 2.5
People are still figuring out the social norms surrounding "deepfakes" and I'm convinced there are some versions of the future where it becomes normalized for content under essentially fair use/free speech doctrine, and many where it's used to restrict free speech in the small number of jurisdictions where public figures would currently be allowed to be featured in this content.

Obviously there are a lot of ways it shouldn't be used but I want to live in the future where it's something we use to have fun, where reasonable, rather than a dangerous/sketchy taboo

weitendorf··on Seedance 2.5
I've been dabbling in this space on the application layer and have spent a few hundred myself experimenting with video models.

It can be really entertaining/addicting to build with them because you're essentially pulling the slot machine and having TikTok/Marvel/YouTube come out of it. I think that was the bet with Sora but the problem is mostly that the novelty wears off quick, and most people want to just consume content without typing in what content they want to see, or sifting through mass-generated spam "content" with nothing behind it to make it worthwhile (a lot of people engage with content parasocially)

Once the tools for creators to steer and integrate models in this space get better, it will explode. We've been working on what I think will be one of the first use case for integrating these models, because I think we're approaching a middle ground where they can be integrated in experiences to provide entertainment/engagement/fun experiences without feeling like slop.

Btw, I'm impressed with some of the AI content on your site but I think you might want to pare down the non-demo pages because it has a different impression me (can I trust that this text is true? / I'm reading a lot of words but not really learning about this person) than you might have intended). I'm a bit of a hypocrite here but also speaking from experience.

weitendorf··on AI financial advice is surprisingly good, especially if you ask right questions
If you understand finance and aren’t specifically attempting to arb on that timescale, you actually want to participate in markets with those participants, because their presence gives you less variance/better price discovery on the scales that don’t factor into your decisions to buy and sell things.

So basically if you’re larping as a trader you will consistently get your ass handed to you unless you are genuinely better than all the pros, but if you’re investing or optimizing for a specific risk profile/exposure/timeline you’re playing a different game.

Anyway the fact that it’s so hard to explain this stuff to individuals does strengthen the argument that most individuals are better off following the herd.

weitendorf··on AI financial advice is surprisingly good, especially if you ask right questions
This already happened 10-20 years ago when personal finance got big on the Internet, it’s just taking a long time to play out.

It was never about ROI anyway, just preservation of capital and peace of mind - makes a lot of sense in the analog/less automated financial world of yore when non-professionals were writing checks or wiring money to people over the phone, and checking stock prices in the paper.

There will also never be a way to pay $10/mo for Gecko+ and trade your way to a lambo with it, because whatever advantage an amateur investor might have is purely from their niche knowledge/information/heterodox beliefs, though I give it about 6-18 months until we’re hearing all about it because it’s a timeless siren song.

← PreviousPage 2 of 15Next →