HNHacker News
TopNewBestAskShowJobs

weitendorf

1,260 karma · joined December 3, 2019

Fred Weitendorf

Founder at Accretional (accretional.com). Building an agent mesh based on open source

Sponsor for statue.dev and previously at Google working on Serverless Infrastructure for Cloud Run and Cloud Functions

fred @ company or https://www.linkedin.com/in/fred-weitendorf-40b505b6/

submissionscomments
weitendorf··on Show HN: Pac-Bench – How well can models one-shot a Pac-Man game?
Gonna be rude and say I don't think this is an interesting or useful benchmark tbh.

Clearly some new RLAAS/dataset/env is being used for this now (it doesn't even seem that complicated, you have one LLM judge whether gameplay is recognizable as the original game or not and another trying to implement a logically/semantically identical version of the game). It's why the performance improvement on this workload has been so dramatic.

Everything is going to go from 0->1 on this benchmark in short order because of that.

weitendorf··on I don't want to read what you didn't write
For better or for worse I’ve always preferred this, and IME a large portion of people (especially younger or who don’t write/communicate much for recreation or their work) will not read it or assume it was AI.

IMO it is a costly/goodhart-resistant way to “show your work” and help other people understand or challenge your mental model. (IE a justification for something you believe to be true). Overly polished writing is performative, it’s hard to take seriously once you’ve read The Economist/LW enough to see how poorly “well written” correlates with truth

To a certain extent are all wrong or ignorant about almost everything because our knowledge/time are very limited. But it is really important to understand what other people think in order to coordinate with them/align human goals and understanding.

It’s good that the average person taking the lowest-friction path to using an LLM in bad faith is easy to identify now. The more obvious and disliked it becomes, the more they’ll be hit with the stick to actually know things and not bother people. It’s so much worse to be “bad and stupid and not care” than “possibly cringe or wrong”

weitendorf··on Show HN: Radius – A Meetup.com Alternative
That matters for discovery, but lots of event organizers and hosts just need a consistent, simple UX for event signups, sharing details to attendees, and managing updates/invites/changes.

If they have a mostly consistent or invite-only set of attendees, the network effects don’t matter. Imagine a Boy Scout troop, Alcoholics Anonymous, pub trivia, blood drives, local alumni association, churches, user groups.

Many businesses have this problem but don’t use Facebook or discord because of the login wall and unprofessional vibes. They would pay for a real tool that makes consistent event/group hosting easy. IMO you’re getting interest/engagement in spite of the critical feedback because people who regularly hosts events need that

weitendorf··on Show HN: Radius – A Meetup.com Alternative
Do not implement any more documentation/pricing/marketing until you've really nailed down the UX. I'm not sure what your background is but you can just write down a list of things like "person looking for events, visiting the site from social/search" and think through the steps/impression they'd take from your landing page to completing their task. It helps to follow actually users/dogfood your app/have some basic analytics to see where traffic drops off or churns to understand this if you have trouble understanding it.

For example, do you think it makes more sense to aggregate groups by subject matter, or by location? Why would I want to browse events on the other side of the world/guess which category a group got bucketed into? The group/events UIs are actually good but the groups directory in the way is not, and your landing page makes the activity feed/posting something prominent for some reason (why would I care about the activity feed or want to post an update to some random site?) instead of the thing that they would actually want to use (groups/events), that actually makes an activity feed/status update make sense.

Sincerely trying to help because I've made the same kinds of mistakes before and it pains me so so much. The main reason is that I find it very hard to trust a product that seems to prioritize aggressive conversion, or strange prioritization of actual user needs vs larping as a saas, and I think it's hard to recognize that in your own work because you're thinking "I'm building a saas" and your users are seeing something that cares more about Being SaaS or prioritize random stuff more than even the basic thing it's supposed to help them do.

The impression you get isn't "more is nice to have even if it arrives in the wrong order", it's actual frustration/distrust seeing all this random stuff and what you want it do or be perceived as, prioritized over the actual thing it's supposed to do. Make it easy to find a group or trust that this is an actual thing people can use, I promise you that you don't need an Events API with embeddable widgets until you can do that.

weitendorf··on Show HN: Radius – A Meetup.com Alternative
Please create a better UI and make it easier to view events / try posting one before signing in. Something card-based where I can quickly type in my location and see what events are nearby.

This isn't just a complaint about the UI being "vibe-coded-coded" or incomplete, but the product itself not making the happy path very useful or confusing.

For example why is the landing page trying to sell this like a b2b saas app or hit me with a bait and switch login wall, or why is the screenshot not just the actual directory/a search bar, why do I have to know to go to groups? (Also, why are you showing me activity feeds or funneling me into that when I'm on your website primarily to look at groups and events? Why would I go to your site to look at that) Then on the group page it would be very convenient if there were just an edit/post button that I can click to draft or actually make a group (and then maybe quickly group by / filter by location or region). Same thing with adding an event, you could have a + button with a modal to draft it and only hit me with the login wall when I try to post.

The event/group UI actually looks decent but I really think you need to prioritize the core activation flow and design. Right now it makes activation, understanding what the product is, and doing the first thing I would want to do (look at groups or events) not obvious. I'm somewhat of a hypocrite on this but also speaking from experience - there's no point in spending time marketing a product with own-goals like that.

Why are your docs so much better designed/clear than the first thing I see or want to do! You don't need docs for this, it's like three database tables you can search and post to! Why the globe?? I'm not even sure why I would want or need to directly browse your users but why is there so much more effort here than in the groups UI (the thing I would be using this for)?

Since this is a pretty conventional crud app the only reason people would really have to use it is UX (or an audience you don't have yet). Trying to be helpful here because I think you should keep working on it and it's something I'd want to use if there were more attention to detail, which you could prob fix in less than a day with a coding agent.

weitendorf··on Durable execution without history replay
It's a problem I've spent a lot of time on, just not through the products marketed like temporal. I commented on the article because it's about program replay with checkpointing which is similar to what I worked with/mentioned.

I didn't see the article mention idempotency anywhere, and you didn't in the original response to the guy who said it just sounded like a buzzword; atomic snapshotting with deterministic execution is literally how you make a program continuation an idempotent function!

And the article is about solving the problem at the language runtime level so it doesn't even do that. So it would be reasonable to assume it is a buzzword if it does not have the essential property you mentioned and I was referring to. I was not even intending to disagree with you but just add what would make it less of a buzzword in this kind of case.

weitendorf··on Durable execution without history replay
Well, you have to either capture/eliminate/persist side effects and control the environment tightly, or it's limited in what it can do.

In a distributed or concurrent system, for full granularity, that can require specialized timing or virtualization techniques up to ensuring fully atomic snapshots and deterministic execution environments (and whether or not that properly models the SUT in real environments, or introduces bias/breaks the reproducibility in a way you care about)

Otherwise if you're only running against fixed checkpoints you have something closer to traces that maybe you could re-run or test against, in some cases, if you put in the work to set it up. In distributed systems that can be a lot of work so it's a bit vague if left unspecified. Because it's not enough to merely replay something if things can drift or don't accurately model the real system

weitendorf··on Durable execution without history replay
Good model. Anybody interested in actually training models or designing agentic systems should be doing this.

My company started around working on this problem because it's the basis for how you train programming models/reliably deploy LLMs to do specific tasks. It allowed me to build a much better mental model for LLMs because I saw how weirdly fickle/inconsistent/picky they could actually be outside of a "chat" where it feels like they have a coherent persona or consistent knowledge/capability.

Initially I thought of it as a search over prompts for capability at completing specific tasks, but now I think the speed/reliability and operations (eg can I switch models without degrading perforamnce?) benefits are even bigger benefits for most users.

A little "secret" since labs are making it harder to even use their models in this way and it's important that it be more widely understood: distribution-aware replay/re-sampling is a key technique in post-training LLMs. But it's also something that allows you to automatically identify the best model for some subset of your tasks, which can save you a lot of money.

weitendorf··on Don't call yourself an artisanal programmer
Cars use AI for steering control (et al) and technology very similar to RLVR (hold the RL), eg property-based testing and formal verification, to prove the soundness of their embedded systems. Most of us in San Francisco trust Waymo with our lives more than human uber/Lyft drivers

As long as you can verify/test and take accountability for the thing you put your name on there’s no reason not to treat it as a process or search problem rather than one you assemble yourself by hand. The only problem is that it’s ironically much harder and more engineering than most “software engineers” are willing or able to do.

I spent several years working on permutation testing/experimentation and creating e2e verification of infrastructure because at scale, or when reliability/correctness are critical, you cannot rely on a single person’s mental model, or for the world to not drift around a system as it works now. That kind of system is what allows you to use LLMs or engineers who don’t know everything about it to improve or change it. It’s more science than art, which is often (but not always) what you want

weitendorf··on Don't call yourself an artisanal programmer
But when nobody will die from your decision (and to be clear I think security is extremely important, but moreso for banking/healthcare than a private wow server), then “shoddy” becomes a matter of reputation/taste vs value/marketability.

The demand curve is different because it’s low stakes, like throwing a bad party or oversalting food. And part of the problem in software to begin with is too much LARPing about scale/engineering for things that don’t need it, as well as lack of accountability or care for things that do.

You can still be an “engineer” working on a game, it’s just more about making the game fun than making it safe. Or, you create a process for making and test hundreds of experimental bridges, and refine/invest additional time in understanding and verifying the safety of the best one.

weitendorf··on Don't call yourself an artisanal programmer
Strong agree. I think the fundamental challenge of working in fields that increasingly become AI-enabled will be the ability to understand and direct large or intricate systems without prior knowledge/the advantage of having built the model as implemented. That’s already how it works in complex domains or large businesses.

It does require a different kind of ego/abilities than before. My (negative) framing of the whiplash effect is that it’s a reckoning of “process fetishism”/a bad kind of careerism in the tech hiring market (because for the labor market to work, candidates need to be evaluable and sortable by businesses, and many people build an identity/optimize for legibility around “best practices” or very particular “technologies” which might get them a job).

Ultimately, you need to know and learn/be responsible for stuff, and be able to help people with your labor, not be “a type of person” that isn’t effective at the task of helping. But at the same time knowing things and being able to take accountability/help people remains critical, especially because that’s what people will want to pay for even as “time spent typing it in” decreases.

Personally, I think it will be a good thing because software and “tech” will become a more strongly domain-driven/enabling medium for real-world or specialized things. IE it is the end to “software for its own sake” or “willingness to type it in and play with Jira/jenkins/frameworks” and the beginning of something that is more applicable or knowledge-building rather than “being the X for Y at Z”. Harder but more fun :)

weitendorf··on Desert Ant Labs: local, fast models that run on device
They're not selling it to you/it doesn't come out of your salary?

They're selling it to your employer. You're literally not even the customer for this product. There's no need to get angry that they're charging enterprise/vendor rates for a product.

And it's free for up to 100k devices. If you are running anything on 100k devices you should be able to afford $0.50/device and. If you aren't then why do you care? It's still a very customer-permissive and flexible business model so I cannot understand why this is upsetting or offensive

weitendorf··on Muse – Meta’s personal AI agent
I genuinely think big tech co's top leadership have very sophisticated strategies/positioning that they simply cannot communicate or explicitly canonize due to their position as spokespeople for the company (and society writ large, news media, investors, customers, employees, vendors). Of course there are a lot of bozos and incompetent people flailing around and a lot of work ultimately gets wasted (which understandably bothers line employees a lot), but is inevitable and even necessary to eg hedge product strategy/comp and take risks on ideas and people.

It would not really be useful to have that conversation with employees because it's incredibly distracting (now product strategy is up for debate with way too many cooks in the kitchen), and very few employees have the exposure or skills to meaningfully contribute even if they think they do. I saw it firsthand at Google TGIFs.

Meta's strategy seems to be "personal agents" quite consistently. Remember they tried to buy Manus? And note that Muse Spark and the meta AI platform products clearly seem to prioritize web search, computer/browser use, and vision/language tasks over coding, which is something that had to bake for a long time. Also, this product launch itself is pretty interesting:

1. It's clearly a fast follow to the current FOTM hype startup Instinct with a much more comprehensive implementation and integration with their other agent products.

2. It's kind of like openclaw, which got a lot of non-developers very excited but was basically consistently unusable. Except this agent's compute runs remotely and presumably has slightly more sane development practices. IMO it's the first main openclaw-like product that has made it to the "just works" level of usability.

3. The focus on ecommerce is actually really really important, because Meta is trying to capture the intent/demand-driven purchasing flow that Google currently owns through search. Controlling the top of funnel is what enables google to make hundreds of billions of dollars per year on search ads. Meta is an advertising company and Google's search ads business is the most lucrative and centralized/well-defended advertising market in human history.

So it is actually a really big deal that Meta is trying to go after it (at least, the CUJ, it's possible that they'd monetize the agent-driven UX differently than search ads) because it's probably their best shot at disrupting that market and one of the few growth opportunities that would actually make a dent on their balance sheet. Obviously Meta is not going to lay out all that strategy stuff explicitly because for all intents and purposes it's a distraction and shifts the conversation in an unproductive direction (the strat behind the product, rather than the product itself). But it's pretty clear if you look for it.

Edit: Actually I thought about the advertising business more and I think in the short term this is partially about attribution/conversion. In 2022 the Apple tracking changes cost Meta $10B in lost conversion metrics; an agent-driven and proxied purchasing flow has built-in attribution and funnel measurement which is very valuable in its own right!

weitendorf··on Muse – Meta’s personal AI agent
I’m pretty sure if you asked my mom what problems AI could help solve for her, navigating website’s purchasing/account flows and directly answering questions about the contents of her email would be #1 and #2.

And my mom doesn’t use AI products like chatgpt because she’s not a student, nor agents because she doesn’t work in tech. To her AI, is no different from those chatbots websites pop up in the corner, the ones you never seek out or interact with intentionally, because you don’t need what they offer.

So I actually think this kind of marketing is quite helpful for regular people who mostly just use their phones to buy stuff, looking things up, navigate life, and entertainment. My mom doesn’t give a shit about APIs or sandboxing, or benchmarks and to her it literally is a hassle to manage a million one-off accounts and confirmation codes and receipts when she just wants to buy something on her phone. She would not assume AI is capable of that or know how to set it up locally.

that’s not even getting into the fact that the top 10% controls 50% of consumer spending in the US and primarily purchases convenience + health + experiences, whereas the other 90% primarily prefers to buy aspirational/identity based goods that evoke their mental model of a higher status lifestyle. It’s why the same bustling lifestyle archetype you criticize is literally used exclusively in car ads, housing marketing materials, consumer electronics, home goods, etc. Because normal people want to feel and/or look good not spawn subagents

weitendorf··on Muse – Meta’s personal AI agent
For the most part I feel the same way, and actually think it has the potential to completely change how e-commerce/web browsing work once it plays out, to the benefit of the client/consumer.

But, I can definitely see a fair argument from the e-commerce businesses that stand to lose from that, that by removing their control over the shopping experience in favor of an opaque agent, they lose the ability to effectively communicate with their customers. Or to provide a coherent/smooth purchasing flow where important stuff like price/dates/shipping are properly surfaced to the user.

That argument would I think be hard to separate from the desire to corral customers into their marketing brochure and convert site visitors into purchasers without churn. But it is true that the LLM in the middle could consistently miss things, or have weird biases/preferences that force vendors to rebuild their sites around LLM tics instead of actual people, janky harnesses that go unnoticed by end users but cause missed sales due to stale data or blind spots, etc.

Also, demoting their sites to glorified databases with a shipping/fulfillment API will give the buyer agent’s provider a lot of power over them and break their ability to establish branding/repeat-customers. Those are the only ways they can reliably carve out margin that otherwise gets whisked away by the advertising platforms pitting them in a zero-sum competition for placement. So it’s possible it could starve out or kill e-commerce the same way Google’s changes to search (the ai box and inline info) hurt the ad-funded websites their info came from.

weitendorf··on Muse – Meta’s personal AI agent
I think they just want to show you ads and help you buy stuff. The incentive is actually to keep the data to themselves so they can monetize access to it via ads.

The shopping experience is significantly less hostile than Amazon’s 1P digital storefront and I kinda don’t care if fb knows that I want to buy a computer.

The only thing to worry about is that they want you to install a native app. But this UX would be difficult to provide for free via the web due to the obvious abuse potential of giving free access to LLMs + remote compute + browser and tool use. And I think as long as you use it to shop or automate web tasks (and give them access to the top of funnel for customer intent, the $300B/yr thing Google monetizes) they don’t really have any reason to abuse your data.

weitendorf··on Muse – Meta’s personal AI agent
Try asking it to find a good deal for you for <item> across multiple sites. Then tell it to use the browser to navigate through the two most promising sites to confirm pricing and availability, and prepare comprehensive breakdown with its findings.

People shop a lot on the internet, actually. Between that and ads for the things people shop for, it’s pretty much the backbone of the Internet economy.

I bet they would purchase more things than they currently do, and seriously break the economics of Google/Amazon search ads (about $300B of yearly spending just for those two), display ads, and internet-first e-commerce sites if they could just ask an agent working for them to help research/source/purchase things without all the navigation and dark patterns in the middle. Personally I would probably spend at least $1k/yr more on snack/beverage subscriptions alone if it had less friction.

> working class peasant consumer

You mean the majority of people in the world? Most people only their computers to entertain themselves, look things up, buy stuff, and complete tedious tasks (taxes, email, etc) they’d rather not do.

How much do you and your peers spend a year on online shopping and how much do they pay out of pocket for SAAS or AI tokens? And how many of your offline purchases had a significant amount of associated online research involved despite technically completing offline? (Houses, cars, schools, hardware) Yeah that’s pretty much where all the money in the global software industry comes from

weitendorf··on Muse – Meta’s personal AI agent
The inline browser the agent uses that you can take control of or watch is awesome! This is something I built a lot of tools to do in the past year (including a similar inline UX/image + click pass through) but having it Just Work in a remote, fully managed client for free is extremely convenient and useful.

But… this feels like a UX that won’t last once a significant portion of consumers and purchases adopt it.

Either the network traffic is getting proxied through my client (effectively making each user a residential scraper for meta’s crawler) or it’s between meta’s servers and the sites, which puts site operators in a difficult position: if real customers are making purchases through this interface and throttling/blocking meta’s IPs makes you invisible to them and meta’s userbase, you don’t want to block that traffic.

But now every consumer in the world can just ask a question or say “check all these prices and sites for a thing I want” and go do something else, right out of the box for free, and have thousands of page loads and site interactions fire off for them.

Bypassing the branding/marketing funnels or intended UX (cf. vc twitter abusing resy thru instinct) of sites, through some kind of proxy client amplifying the traffic a human would create, with the ability to let anybody scrape or interact directly with a site’s backend… definitely a consumer win, but seems unsustainable.

weitendorf··on Bill Gates tries to install MovieMaker (2003)
How exactly do you think Bill Gates at peak Microsoft is going to fix this given "it's 100% on him?" And do you think the guys on the line are the ones on this email thread?

This is immensely more "ownership"/initiative than the vast majority of middle managers at any tech company would willingly take to solve a minor UX/CUJ failure, that's surely literally why he started the email thread and including the ones he did, to model that behavior for the the other middle management bozos and to try to instill a culture of caring about their work/whether it actually accomplishes the intended goals.

weitendorf··on PostgreSQL 19 Interactive Tour
The way I process this kind of content is by skimming or ctl-Fing for the material I'm interested in reading and usually just reading the example code or the specific explanation for the content I'm after

For example this site ranks on the first page for "go 1.27 generics" and "go 1.27 uuid"[0] and if I were looking for uuid content I'd probably click on the toc for uuid and go to here [1] and look at the examples and v4 vs v7 semantics and then bounce. For this particular article the thing I'd be most interested in applying is probably REPACK [2]

  REPACK (CONCURRENTLY) events;
And all of the content around their code snippet showing that is pretty prescient/semantically dense and useful. What I wouldn't do is read the whole thing front to back, or the prose at the top/bottom with the LLMisms: it's way too long and dense for that.

For comparison here are the official new postgres docs about REPACK [3]. Is it human-written and more informative? Maybe for some people, or for me if I needed to reimplement a postgres-compliant spec or something, but I'd prefer the LLM-assisted (and I say assisted because IME it's actually a decent amount of work to get LLMs to write content like this) article most of the time.

[0] https://www.google.com/search?q=go+1.27+generics

[1] https://victoriametrics.com/blog/go-1-27/#the-uuid-package

[2] https://victoriametrics.com/blog/postgres-19/index.html#repa...

[3] https://www.postgresql.org/docs/19/sql-repack.html

weitendorf··on C Is Not a Low-Level Language (2018)
That's fair. C is very old and used for almost all hardware so I think while you can make the argument that "only clang and gcc extensions asm blocks available like that, and intrinsics are only available through vendor-specific headers" and be right, by that same logic literally nothing except binary machine code for hardware without any kind of microcode can be low-level, and even then it's probably always hardware dependent (because if it's not fully bijective to the actual hardware it's implemented on top of, the semantics leak).

Practically speaking, we have a word for the kind of "abstractionless" model you're describing: machine code. I mean, even assembler is a bunch of abstractions about 'registers' and 'instructions' that are really just specific portions of the hardware or opcodes!

So we either descend endlessly into pedantry arguing that cosmic rays and electron tunnelling represent inexcusable deviations from the overly abstracted semantics that hardware vendors expose in their products or maybe we draw the line somewhere else.

You may not agree with mine, that "practical and simple interop with machine-level language impls across a high-level language interface is sufficiently close to the hardware as to be low level" but there has to be a limit somewhere between that and "technically the hardware's operating temperature is part of its logical semantics because if it exceeds a certain value for long enough it starts to degrade and yield incorrect results or terminate execution". I think eventually it just becomes unproductive nerd sniping, personally

weitendorf··on C Is Not a Low-Level Language (2018)
Sure, but then you're really arguing that the ISA no longer maintains 1:1 instruction-level implementation and that this is the definitive quality of whether or not something is low level, to the point that any deviation from that model makes it not officially "low level". To me that's just a very tedious pedantic argument that simply fails to capture the actual meaning behind why/when we might call something low level.

TFA famously argues that Spectre/Meltdown et al break that abstraction. But note that they are quite literally exceptions to the rule: the only reason we know/care about them is that the "magic under the hood" that was supposed to make CPUs faster while maintaining that abstraction introduced a bug that caused the implementation details to leak to the end users.

Similarly even vp2intersectd took multiple cycles in its original Intel impl and even in the performant AMD Zen5 impl it still takes >1 cycle with 6 levels of pipelining or somesuch. Ok. If literally not even a chip's ISA is "low level" then the term is effectively meaningless.

The only way you could define a "low level" language capable of exercising that hardware's capabilities fully would be to have some kind of per-cycle, pipeline-aware annotation layer over the actual machine code... which really seems like quite a lot of noise/cruft you'd not typically want to add on top of everything, all in the name of still technically being low-level according to some dubiously pedantic criteria nobody would event want in practice.

weitendorf··on C Is Not a Low-Level Language (2018)
This is such a pedantic point IMO. C is low level because it makes it very easy to work with machine language/assembly and do stuff like this (LLM assisted example follows):

  int main() {
    __m512i vecA = _mm512_setr_epi32(0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15);
    __m512i vecB = _mm512_setr_epi32(0,5,10,15,20,25,30,35,40,45,50,55,60,65,70,75);
    unsigned short mask = 0;

    __asm__ (
        "vp2intersectd %[B], %[A], %%k2"
        : "=@cck2" (mask)
        : [A] "v" (vecA), [B] "v" (vecB)
        : "k3"
    );

    printf("Intersection Mask: 0x%04X\n", mask);
    return 0;
  }
This is something "low level" programmers use very often to realize the benefits of a high-level language while exercising explicit control over using specific hardware instructions (vp2intersectd being an AVX-512 instruction used in highly optimized search algorithm impls).

Obviously if you rely on implicit behavior from the compiler to optimize your code you are no longer "low level". But if you can quickly and easily drop into machine-level instructions to provide explicit implementation semantics, and the language indeed makes that relatively simple and easy to do, that sure seems "low level" to me

weitendorf··on PostgreSQL 19 Interactive Tour
Gotta say this site is killing it in the LLM-written technical blog SEO game the last couple years, they're consistently able to dominate (and actual deliver in spite of the LLM-isms) SERP for a set of technologies that very closely align with the ones that a typical backend/infrastructure engineer would work on.

I checked their list of blog articles from the last year and they haven't even written that many, which makes it quite impressive because they really model the set that I've been working with (Go, Postgres, grpc, otel) well by I guess defining some kind of customer archetype and writing really detailed guides about what they'd be interested in learning more about. At this point I've encountered their site "organically" like 5x in the last year and recognize the name/style of content so they're doing something right in the marketing department for sure.

On one hand you could just dismiss it as spam but articles like this actually represent a pretty significant LLM spend/human review element that delivers real value to me as a technical end user looking for info on google or in technical blogs on HN (ie it would take me a long time to generate something like this myself and I wouldn't do it proactively, only when-needed). So it actually does help me quite a bit that they do so before I think to ask about it.

weitendorf··on GLM-5.3-Flash
Strong agree, but I also think some roles in big companies (for me, infrastructure) or in certain industries (eg trading/finance) can help build the same understanding without as much of the variance/raw exposure to bottom line.

Now that the role of the ticket-cruncher is on the path towards full commoditization, and individuals can move much more quickly (and even more carelessly!), I think product roles will probably shift towards one where developers are more deeply embedded in the product/business process so that they own/understand what to build without as much separation between the decision-making and prioritization of what to build. Or at least, they should.

It was eye opening to me to run the math of "should X people work for Y months on this project to save Z per year?" and realize that in so many cases, the time and effort it would cost to stop "wasting" money on things is WAY more than you could actually save on it. Even "small" projects can very quickly become $1M+ investments in time and resources, and the diminishing returns add up quickly (but also a good way to justify the value of your contributions, when done). But the job only exists if it saves money or makes money...

weitendorf··on GLM-5.3-Flash
I have exactly the same opinion

Over the last couple years I’ve had to learn sales and understand the thought process behind this better, and I think I’m beginning to understand it

The psychology is that most people aren’t really trying to optimize for productivity (even most people who think they are) on an ROI basis, because their compensation is too decoupled from their actual raw output, and more closely coupled to how differentiated their marginal contribution is to peers. They’re much more incentivized to spend their personal/work time optimizing for being more skilled or acquiring some kind of competitive advantage relative to baseline.

Most people don’t consciously run the numbers of “I get paid $X/hr to add $Y of value” or model pay at work as something with variable inputs (eg something that can be increased with high performance), so it makes sense to them to spend 20 hours of time to save $100 or to make themselves 5% less efficient to take home 0.5% more or avoid doing something they don’t want to start doing.

NOT saying this always happens or that they’re stupid for doing so. I didn’t even realize how much I had been doing it myself until I started recognizing it, and shifted to having my own comp/performance fully aligned with the company’s P/L.

It actually makes a lot of sense IF you can accurately estimate incremental upside (which is much harder and more diffuse than modeling downside if you’re salaried a employee) or if the upfront skill/knowledge investment that looks like bikeshedding pays off in the long run.

weitendorf··on Protobuf has LSP support
This has been driving me crazy ever since I started using protobuf/grpc and realized more major tech companies (in the cloud/infra/data world at least, and a lot of other SAAS) were using it or something similar (eg capn proto) internally than not.

It feels like we’re in some sort of deadlock where each of them think “proto/grpc are too niche to support for external users, better just use JSON”, keeping it unfamiliar for an Average Web Developer. But if every company using it just exposed it to third parties/added it to their public APIs, it would immediately be common (and trendy) enough for every web developer to learn it and start using it.

If Google added the missing HTTP/2 streaming support to browsers (blocking native bidi grpc streaming) it would have an instant killer app in making it easier to implement websocket-like client/server applications. It makes absolutely no sense that full duplex bidi was added to the HTTP/2 but remains unimplemented in browsers.

The role MCP, OpenAPI, and JSON schema fill all would be a million times simpler if they were based on protobuf instead of JSON. I can forgive OpenAPI/JSON but it honestly pisses me off that we ended up with MCP and JSON-RPC + JSON schema, and people think these are cool/good tools, and actively adopting them. Just piling on the slop

weitendorf··on Anthropic IPO valuation hinges on $190-200B 2028 revenue forecast
Anthropic’s revenue run rate just hit $50B, in a market that didn’t exist 5 years ago and was 10% of its current size 1 year ago. They are, by far, the fastest growing company in human history. The demand for “rockets” has been pretty high and growing ever since they started supplying it.

Anthropic and its investors/customers don’t need random internet commenter’s permission to decide whether or not something is worth doing or justified.

Personally I think when something doesn’t make sense to you, you should try to figure out why it makes sense to other people, and whether they might have different needs/constraints/incentives/skills/knowledge.

Why might people spend more on AI as it gets better, rather than less? Could they, perhaps, allow for entirely new kinds of capabilities and products that hacker news commenters have not yet seen? As they have done every year for 3 straight years (hacker news has been wrong in exactly the same way every time, btw)? Might some people prefer to spend $10/day to work with the most capable AI available, due to the amount they use it while doing their job, and its effect on output? Do some people use AI for more than just tinkering with open source harnesses? Could it be that when you don’t understand something, there really is a way to explain it, that isn’t “everybody else must be stupid”?

weitendorf··on Models Are Getting Dumber on Purpose
You 100% can finetune or adapt/build on top of models, and specialize them or extend their capabilities. That’s literally what post training is.

The problem is that “finetuning” was a 2023 AI FOTM associated with products/demos that were almost exclusively using it for LLM character role-play/output style purposes (ie not in actual systems where they served a more functional role).

This made people think you could train models without replay/real evals by yoloing it with SFT (this is partially an artifact of that era being much heavier on autoregressive training and not so much evals). You really can finetune and get results but you have to treat it like a small ML training run, with real evals, and more intentionality than just “more examples”.

You can find pretrained and -instruct models on huggingface that clearly demonstrate what specialization/staged training runs do.

I’d be very wary of conflating finetuning with specialization/extending a model’s capabilities in general.

weitendorf··on Anthropic IPO valuation hinges on $190-200B 2028 revenue forecast
I walk to work and ride a bicycle in my free time, don’t see why anybody would need a hatchback or semi truck or build rockets. Idiots
Page 1 of 15Next →