HNHacker News
TopNewBestAskShowJobs

calebkaiser

1,058 karma · joined October 25, 2019

https://twitter.com/KaiserFrose
submissionscomments
calebkaiser··on Pi 1.0
I don't use Pi very much for actual writing code. I have some specific use-cases where I do, but I'm admittedly just not a person who can reasonably juggle a lot of tools or a super personalized setup. I mostly just use Codex or CC from the terminal, though I've recently dabbled with some GUIs. Basically, once I find something that works, I'm pretty reticent to spend any time tinkering unless I feel a real need.

But what I do use Pi a lot for is as a base for agents. I much prefer it to using an agent SDK. I find its minimalism and extensibility to be a really nice substrate for new projects.

I could easily imagine someone getting to a really productive personal setup with it as well, for the same reasons as above.

calebkaiser··on Clef: Open-weight decision models, and new RL fine-tuning platform
There is a bunch of stuff to tease apart.

In general, training a general purpose classifier is something lots of people have worked on for a long time. Large Transformer models themselves are typically "generalists" already, so structured generation and constrained decoding have given you the ability to use an LLM as a general classifier for years. It's an incredibly common pattern for working with LLM judges or any sort of branched decision making workflow.

A lot of people who are a bit less familiar with the field saw the hype around Jev and presumed that the reason it was so exciting was that it was a fundamentally new interface for working with an LLM. And that additional excitement drove even more attention to Jev. But fundamentally, TypeSafe's announcement was that they found a particular architecture/training paradigm that resulted in a model for this particular interface that had incredible accuracy, very low latency, and for which they could offer inference at a super low cost.

I've not kept up with the flood of Jev clones that have been released, but I think this is just typical for any new component in deep learning that gets popular. There are an absurd number of open source autoregressive LLMs and fine tunes you can use. The thing that makes one more popular than the other is typically the general performance of the individual model.

But training a model for this purpose, or emulating the procedures described in Jev's papers, isn't something that would be beyond the capabilities of any lab. It's not an entirely alien architecture or approach.

The bigger question for TypeSafe as a company would be if other teams are producing Jev-like models that win on performance or cost. Like I said, I haven't followed the reports super closely, so no idea if that's the case or not.

calebkaiser··on DeepSeek Elastic Compute (DSec)
That's not quite right. A recurring trend in ML is figuring out how to get an elastic interface for the the particular quirks of ML workloads (source: I worked on an open source one years ago for inference). Recently, there's been a lot of interest in doing this for agent workloads. Google recently released something called ax that is similar in spirit. The core of it is this: https://github.com/agent-substrate/substrate

At a high level, agent sandboxes have peculiar needs. Agents are really bursty, but also long lived. You need low latency suspend/resume calls. Checkpointing and recovery have some particular considerations. And naive approaches are often really wasteful, but over optimizing without harming durability, isolation, or consistency in performance can get tricky.

This is the new hot infra topic for agent swarms, for the time being. Whether that's impressive or not is up to the reader I guess, but it's a pretty involved project nonetheless.

calebkaiser··on Ollaya – Ollama for open-source, Jev-style decision models
I think the answers are "a lot" and "a lot more"? But what I'm saying is that Jev's virality isn't because people didn't have access to classifiers before it. Jev's general claims about capability and performance vs. cost would have been a very big deal 5 years ago too. In a vacuum, the idea that you can get a general classifier that is very accurate across any domain and on any modality with minimal latency and a very low price point is wild.

In early 2016, Clarifai's core product was basically just an image classifer exposed via an API. And at that point, they'd raised $40 million--the same amount as TypeSafe.ai/Jev--and they were experiencing viral growth among developers + signing contracts with a bunch of flashy logos. The demand was so high that Amazon launched Rekognition and Google launched their similar APIs to compete.

The AI hype cycle and the number of people thinking about using AI certainly puts more wind at Jev's back, but even in an alternative universe where we don't have contemporary LLMs, Jev's core claims would be remarkable and there would be a big market for it.

calebkaiser··on Ollaya – Ollama for open-source, Jev-style decision models
I don't know if that first part is true? Classifiers were/are one of the dominant applications of classic ML and neural networks, especially in production. Even today, image classification, object recognition, language detection, segmentation models etc are still super common.

I think the hype with Jev is just that, while structured generation is great, LLM judges tend to kind of suck for precise classification. And the more powerful the base model, the more accurate they can get, but they get increasingly expensive/impossible to finetune. It was specifically the latency/price point Jev offered vs. the general accuracy it claimed that generated all the excitement. Plus the promise of cheap calibration (tuning).

"Jev exists because LLMs exist" is kind of a truism, as Jev apparently is literally a Transformer model.

calebkaiser··on AX – Google’s Open Agentic Orchestrator
Yeah I'm actually less hesitant to try out Google open source projects than I am new Google products. I have no idea if this is accurate or just my impression, but I feel like I've been burned by the "killed by Google" meme almost exclusively on their software products, whereas there are plenty of open source efforts from Google that I think of as stable.

In addition to the ones you listed, I'd add the V8 runtime, Jax, Protobuf. Even some of their projects that wound up declining in market share (Angular, Tensorflow--both losing share to projects that wound up at Meta, ironically) are still actively maintained and pushed.

But I'm sure there's also a huge graveyard of open source projects they abandoned that just never hit my radar. Still, at least with their open source stuff, you can fork in the worst case.

calebkaiser··on Prompts aren’t Real
I think this is basically the motivation behind MCP servers, so you're in good company.

But I've spoken with many people at companies who've decided to add an in-UI agent to their apps, and I don't think this trend will persist. Absolutely AI/agents will increasingly become a major component of interfaces, but the current common incarnation of "Clippy for X" feels like a kludged bridge between an app that was not designed from first principles for agents (and often was poorly designed for humans) and the impulse to be "AI native".

From my experience, some of the most successful niches for this sort of UI so far have been in apps that were already well suited to it. I'm thinking specifically about analytics/dashboarding/"this is a portal for you to query things easily" software. They already started with a lot of the elements you want: visible provenance of the agents actions via the queries it writes, a malleable interface that you already expect to be customizable and ephemeral, and most importantly, "navigation" that is genuinely difficult for many users (in the sense that "navigating" can mean "querying specific data"). The agent provides a ton of value to users and its actions are intuitive and legible.

But the agents you see right now in a lot of apps that do things like navigate you to the right page by... sending you a link, which you could have clicked from the navbar? Or worse, which hijack your navigation and throw you on a page you're unfamiliar with, and where you have no sense of place or how to make your way to/from? I don't think they're particularly long for this world.

calebkaiser··on I built non-autoregressive decision models with RL a year ago
Jev seems pretty cool! I just got access and have only gotten to do minimal experiments, but I love this general area of research and it fills a very real need.

I agree with you. I think the OPs pushback is emblematic of a larger reaction I've seen that is, at the very least, misinformed.

There are a lot of approaches that use a self-attention backbone for classifier-style outputs. You have structured generation libraries like SGLang and Outlines, but those basically give you guided generation on an autoregressive model. You also have a bunch of models that are non-autoregressive that try something similar. Older NLP stuff applies here, and there's newer stuff using diffusion transformers for this purpose.

But I don't think the Jev author has ever said that he's the sole human, alone in a vast sea of misguided researchers, who is interested in schema-guided classification? I think he said he found a novel way to train a model for this task that has much higher general intelligence at much lower cost than other approaches. Which is an exciting result with lots of applications if it bears out.

I think some people are just reflexively skeptical of anything that gets a lot of hype. Maybe that's fair. Things that are wildly successful and high impact also tend to get a lot of hype though, so it seems like a poor filter.

calebkaiser··on I built non-autoregressive decision models with RL a year ago
This is also a really common thing in ML specifically. We joke about getting Schmidthuber'd, which is when Jurgen Schmidthuber (sometimes correctly) announces that he or one of his colleagues actually proposed your thing 37 years ago in a Japanese linguists journal.

Statistical modeling, from simple classical stuff up to modern deep learning, just has this dynamic where the theory is rich and bottomless, but the actual components of implementation are pretty neat and compact. So for any given idea, there are probably 20,000 other people who have had the same intuition, just with subtly different application or implementation. Add in that depending on what your particular flavor of research is, you might name an almost identical implementation something completely different. And it leads to a huge amount of sour grapes whenever anyone's idea really garners attention.

If you listen to any podcast with a founder in the ML space who has been in it for long enough, they will invariably say at some point "We actually developed xyz over a year before OpenAI"

calebkaiser··on Almost Never Use AI to Write Anything Substantive
I understand the distinction you're drawing, but in my experience, it gets a lot blurrier as a project evolves. Caveat that I use LLMs all day and have for quite some time, so I'm coming from a positive perspective.

The danger for me in LLM code is the same as in writing, it's just that I'm not typically writing at the same scale as when I'm building something. The final piece when I'm writing is usually a message or a 1,000 word article at most. So I'm naturally going to analyze it quite intensely, because I can afford to. And I don't really use LLMs for this at all. I use them for things around writing (research, interrogating ideas, situationally specific stuff, mapping, visuals, publishing, etc.) And the code equivalent to an essay or message would probably be something like a single script, or the sort of thing I'd write as example code when I'm teaching. In those settings, again, LLMs can be helpful, but I'm still going to be really opinionated at a highly detailed resolution.

But a codebase is more comparable to a novel than an essay. Or more directly, the writing in a codebase is usually the documentation, which grows commensurately with the codebase. And the real LLM risk here is the drift that can happen over the course of many epics or "chapters" as the LLM writes "code that works but is imprecise and probably shouldn't work this way" or introduces weird new terminology that neither of us can precisely define. Worse, this usually becomes obvious down the line, and I have to parse through the verbose constructed world the agent has created to trace the issue back. That's a big cognitive tax, because I'm holding these weird parallel worlds of "How did the LLM's alien brain get here within the bounds of the contracts" and "What do I really want this to look like".

So I think it's fundamentally the same phenomenon, and we're all developing our skills around working with it in real time.

calebkaiser··on I think you should almost never use AI to write
I think about this a lot. I started my teenage-to-young-adult life in the literary world as a poet who loved programming, and made a living as a ghostwriter. At some point, I fell in love with mathematics and wound up in ML in research/engineering for the last 8 years or so. So, I've thought a lot about writing and ML and their intersection.

I think one of the underappreciated things about writing "substantive" work is that the work you see at the end isn't the first attempt. And I don't mean the first draft of the piece. I mean that almost always, writers iterate on the same topic many times, either with complete published pieces or abandoned drafts or even just conversations and sessions of unproductive daydreaming. It's a cliche that your best work typically also comes out fastest, but it's not because of divine inspiration, it's because you've whittled the big idea you actually care about down so much in your mind that you instinctively know exactly how to write it.

My experience has been that for people who don't work this way or don't write a lot, LLMs can give them this incredible feeling of leaping straight from inkling to "substantive" writing. And because they haven't built up those muscles or "taste", they don't immediately recognize that it's imprecise and hard to follow.

That's not to say they're not brilliant in their own right, just that they haven't spent a lot of time on this particular thing. Sort of like a very gifted programmer who doesn't have a ton of experience yet (speaking as someone who is gifted at nothing and frequently has to do things they're inexperienced at).

So I don't think the problem is that LLMs just write bad. It's that LLMs are so wonderfully powerful that they allow you to confidently leap forward to create something that is a little beyond your experience.

And that's why I have different reactions to heavily AI generated writing. When it feels like marketing at scale, it grosses me out. But when I feel like it's just someone who is excited to write an idea and maybe doesn't have a lot of experience doing it, I'm not judgemental. My hope is that it makes them more excited about writing, and that trying to make their next piece better will lead them inevitably to start thinking about where the last piece fell short. And if there's some slop along the way, eh, I'm not compelled to read it.

calebkaiser··on I spent $220 on Google app ads and 60% of the installs were robots
Which hyperscalers are doing that? I've heard of one startup trying this, XFRA, currently in early pilot phases. But I'd be pretty surprised to hear that AWS or Google are paying people to host GPU clusters in their home.
calebkaiser··on OpenAI’s Navier-Stokes release included a Lean 4 formal proof
Yeah, in essence. This is actually a pretty cool part of working in Lean. It's a somewhat normal convention to write something in a human readable way and then write a second optimized implementation with some kindness of correctness theorem connecting them. There was a whole open "competition" for writing a faster Lean kernel/proof checker that didn't sacrifice on soundness called Lean Kernel Arena. Fun reference point: https://kim-em.github.io/blog/2026-7-24-why-lean-is-faster-t...
calebkaiser··on Douglas Hofstadter: Analogy as the Core of Cognition [video]
I'd wager that the vast majority of the ML research community, especially anyone interested in "AGI", is familiar with Hofstader's work. And I don't think anyone working on contemporary language models would argue that they are somehow an assumption-less "pure" model--the particular inductive bias of the Transformer has been studied by a huge number of researchers and continues to be, and the same is true for things like training data bias.

I think the Hofstader's view of modern LLMs is actually a deeply human and touching one. Looking at his work over the years, his curiosity has always veered towards human thought. He could have written GEB with a focus on completely different examples of self-reference, but he chose three striking humans from history. When he's describing modern systems as "empty intelligence", I think there's a little bit of heartbreak in his perspective, because he sees them as fundamentally different from humans in a way that leaves the part he loves--the "I" in the loop--out of the equation. He gave an interview a few years ago where he explains his feeling as being "diminished" not in a "What will I do if I'm not the best at math?" kind of way, but more specifically as he puts it, that humans are "imperfect, flawed structures".

calebkaiser··on OKF Agent Memory – Git-native persistent memory for AI coding agents
Love seeing projects like this. The performance benchmarks are nice to see. Have you done any benchmarks against approaches like OpenAI's Symphony for things like token usage or task completion?
calebkaiser··on Corporate America is getting hooked on open-source AI
There will be companies doing this. I'm saying the labs are well positioned to be those companies, as they effectively are those companies right now.

The same dynamics that define the public cloud ecosystem are at play here. What AWS sells you is access to appropriate hardware and turnkey infra for your needs. Looking at the cloud industry over the last 20 years, I find it hard to believe that it is impossible to build a moat or a huge business around this.

calebkaiser··on Corporate America is getting hooked on open-source AI
I support and use open models as much as possible, but I'm not totally convinced that OAI or Anthropic have no moat, even as open models catch up to the frontier. Serving and inference are still hard problems when you're talking about a 2 trillion parameter model. Fine-tuning, if that remains a realistic need for businesses, is also a difficult infra problem at that scale. In the most bearish case, where there is no competitive advantage to using their models, big labs still have an advantage in this area.

Maybe there is some threshold where the price/quality math for your standard business tips in favor of smaller models and self-hosting the entire stack. I'd certainly love that.

calebkaiser··on The Emergent Symbolic Structure of Artificial Neural Networks
I haven't read this in depth yet, though I plan to. If this general line of research is interesting to you, I'd recommend checking out some of the lines of research it touches upon--they're really rich and fascinating, and some are pretty approachable mathematically even if ML research papers aren't usually your thing. The related works section here seems pretty well stocked, but mechanistic interpretability is a pretty interesting peephole into this general vein: https://transformer-circuits.pub/
calebkaiser··on I trained a small transformer in 1.5hrs and it beats many LLMs
It gets extremely blurry, because people commonly refer to any model that uses a component associated with the Transformer architecture as a Transformer (i.e. using some kind of QKV-esque attention mechanism). I think it's easier to think of it like this:

A large language model is just what it says--a very large statistical model trained for language tasks. This covers the spectrum of GPT-style models, but also those hard to classify ones, like Liquid's "Liquid Foundation Models", which can get up to 24 billion parameters and use grouped query attention, but are closely related to state-space models as well: https://huggingface.co/LiquidAI/LFM2-24B-A2B

Also, as others have pointed out, a Transformer isn't inherently a language model. So really they're sort of two different axes, one classifying the model size and task, the other referring to a specific architecture.

calebkaiser··on Nvidia agrees to acquire Hugging Face for $13B
I think maybe just click around Huggingface's site a bit? They don't run a compute-intensive business on low monthly plans. You're describing frontier labs that sell access to their enormous models with subsidized subscriptions, but that's just an entirely different company/model than Huggingface. They've been around in their current form since around 2018. They provide a GitHub like service for hosting and sharing models primarily, and they also provide infra for optimized compute (for training and inference) that you can purchase through them, but you pay as you go for the compute and they have their premium baked into the price.

Their lead advocate just posted this describing things: https://x.com/mervenoyann/status/2092924706508698025

calebkaiser··on Nvidia agrees to acquire Hugging Face for $13B
I just feel like you're saying things based on a general vibe about "AI companies" but didn't pause to look up the particular company being discussed.

Huggingface reached profitability 2 years ago. They've reportedly just started to touch the cash they raised 3 years ago. They have a tiered pricing model with metered pricing on resource heavy services. Seems pretty stable?

Your point about their user base is even weirder. They play essentially the same role for the ML community that GitHub does for software engineers, so I mean, yeah I'd pretty much expect their relationship with users to be "exactly like" GitHub's for the most part? They're where everyone has published their models for the last 6 years at least, going back to pre ChatGPT and the recent AI boom. There aren't realistically any other major model hubs, certainly not with anything near their footprint.

calebkaiser··on Nvidia agrees to acquire Hugging Face for $13B
I'm curious why?
calebkaiser··on Nvidia agrees to acquire Hugging Face for $13B
It's not too dissimilar from GitHub, but geared towards ML. They have a 9/mo pro plan for individual users for upgraded storage/usage, and an enterprise version of Hub that larger orgs can pay for. I think the enterprise has some contract minimum + 50/mo per seat. https://huggingface.co/pro

They also have inference endpoints with metered prices, and their spaces product (though i'd imagine this is a smaller portion of revenue).

calebkaiser··on Nvidia agrees to acquire Hugging Face for $13B
Solid point. I'm sure Nvidia went into this deal expecting completely flat growth and no other benefits to their core business. Sorta like how Meta never increased Instagram's revenue from $0 and is still waiting for it to pay off that billion dollar acquisition price.

Or like GitHub, which was generating something like 200 million in ARR and had never hit profitability when Microsoft bought it for $7.5 billion back in 2018. I'm sure it has come as nothing but a happy surprise to Microsoft that GitHub generated $1 billion in 2023. They had initially penciled it in for 38 years til ROI.

calebkaiser··on Nvidia agrees to acquire Hugging Face for $13B
Yeah, great point. Nvidia could potentially be overpaying. Not sure how that equates to Huggingface being "a file download mirror with a couple of side features dangling off".
calebkaiser··on Nvidia agrees to acquire Hugging Face for $13B
I dunno. You can argue over whether they're overpaying, but it's not like Huggingface is Clinkle. They hit $150 million in ARR this year, they have tons of runway, and according to reports, have just started to even burn the money they raised a few years ago.

I get that it's fun to be glib about the stupidity of tech elites and investors in general, but Huggingface have been pretty open about their financials and are, in my opinion as a practitioner in the field, one of the most responsible orgs in our space. They've been a pillar of open source ML for years now and have made a very positive impact on our ecosystem.

Nvidia is getting a real business generating revenue, and the center of the universe for open models. Both seem like pretty valuable attributes, from Nvidia's perspective.

calebkaiser··on Nvidia agrees to acquire Hugging Face for $13B
They hit 150 million in annual recurring revenue this year.

Pretty nice dangling side features apparently.

calebkaiser··on OpenAI Jalapeño: Better than Nvidia Blackwell
What? Both of those companies absolutely serve LLMs, and both of them would love for serving LLMs to be an even bigger part of their business. Not only that, AWS is Anthropic's primary compute partner for training and inference. They literally use the newest generation of the Trainium chips I mentioned before: https://www.anthropic.com/news/anthropic-amazon-compute

Chips are another axis for improvements in training and inference. Orgs large enough to explore the space have been doing it for at least a decade now. This is just a silly line of reasoning based on the faulty assumption that somehow, looking for increases in efficiency in training/inference means teams have reached some theoretical limit in model capability.

calebkaiser··on OpenAI Jalapeño: Better than Nvidia Blackwell
Google rolled out TPUs in 2015. AWS released Inferentia and Trainium chips in 2020.

If companies working on ML-specific chips was evidence that large transformer models have fully saturated their potential, the field would have been done circa GPT-2.

calebkaiser··on Moderna reports first positive Phase 3 for mRNA neoantigen therapy in melanoma
Sorry--I was unclear. I was specifically responding to the OP's "CEOs these days want to join the AI hype wave" claim. My point was that the "hype wave" OP is referencing is really about deep learning, not what we think of as traditional statistical learning. Even the dedicated medical AI startups of that era, like BenevolentAI, were still doing traditional ML in 2014. So if Moderna really was trying to join the "AI hype train" and announce a deep learning result with this treatment, they'd be announcing that they successfully trained a deep neural network for drug discovery 12 years ago.
Page 1 of 8Next →