HNHacker News
TopNewBestAskShowJobs

garrinm

284 karma · joined September 4, 2019

submissionscomments
garrinm··on “Next-token predictor” is the wrong mental model for LLMs
Originally I had those parts written in math with probability functions and the likes (its closer to my background). Then I remembered who is my target audience... but now that I see exactly who is my target audience I'm thinking I'll should have snuck a pelican in there. All jokes aside I appreciate the comment and I'm glad that rewrite paid off!
garrinm··on “Next-token predictor” is the wrong mental model for LLMs
To be more specific there’s no ground truth tokens to predict. There a verifiable answer in RLVR. But the tokens are explored. Not predicted as there’s no true token to predict.
garrinm··on “Next-token predictor” is the wrong mental model for LLMs
I try to make 3 claims in the post, it was a bit clumsy I'll admit that.

1. At inference time, LLMs emit one token at a time given the prior tokens. This looks like prediction and I concede that.

2. During pre-training, LLMs predict the next token and compare to the actual next token in the training data. This is the classic setting for ML predictions. And I think its meaningful, the model really is predicting what the ground truth next token will be in the data.

3. During post-training, in the case of RLVR, there is no ground truth next token. In pretraining, the question is "what token actually came next?". In RLVR, the question is "what sequence of actions gets a high reward?"

And the whole point is that thinking about the RLVR is important. A mental model that stops at 1 or 2 is incomplete and doesn't capture what drives LLM tokens.

garrinm··on “Next-token predictor” is the wrong mental model for LLMs
Yes I understand the analogy was a bit loose. I'm comparing what happens at "inference time" in chess engines to what happens at train time in LLMs. In hindsight AlphaGo Zero was the perfect analogy, but I missed that opportunity.

The analogy with chess still works, but there's an extra step to think about. In both cases there is some kind of search over possible future trajectories. A chess engine explicitly searches branches of the game tree and evaluates which moves lead to good outcomes. In RL for an LLM, you sample rollouts, evaluate the resulting trajectories, and use those evaluations to update the policy.

The extra step with the LLM is that you don't keep doing that whole search at inference time. You use the rollouts to update the weights, so in some sense the useful information from that search gets compressed into the model.

But if you accept that the model is, in some loose sense, storing what it learned from those rollouts in its weights, then at inference time they are doing a similar job: taking some input state (prior tokens or a board position) and choosing the next action.

garrinm··on “Next-token predictor” is the wrong mental model for LLMs
In the article I made 3 claims, and I agree it was a bit clumsy.

1st I say that "working forwards" in the sense of outputting one token at a time could be some form of prediction, I don't argue against that. This is what LLMs do at inference time.

2nd I say that to me what really constitutes a prediction is the pre-training. Here it's the classic setting for the word prediction in ML. The model outputs a prediction of the ground truth label: the next token.

3rd I argue that in RL there is no ground truth next token, so prediction doesn't apply here anymore.

Back to your question then: you're asking points 3 and 1 are different. Working backwards from a set of win states is basically what RL does in training. Working forward from the current state is what inference does. To me there is a distinction worth thinking about. First between the mechanism at inference time and at train time. Then between what happens in pre-training vs. RL post training.

garrinm··on “Next-token predictor” is the wrong mental model for LLMs
I think the point is more that in RL there's no ground truth to predict. So when training a model with RL the idea of "predicting" doesn't fit anymore. I'll make some edits I see that I wasn't very clear.
garrinm··on “Next-token predictor” is the wrong mental model for LLMs
It does in pre training, but not in RL post training. And not at inference time. Reading over all these comments I get the feeling my mistake was not clearly delineating inference time and train time.
garrinm··on “Next-token predictor” is the wrong mental model for LLMs
I think that’s fair, I didn’t actually run the whole thing through an AI. it was more targeted edits, but each time it does erode at my writing. But at the same time, I don’t think it’s a good reason to dismiss this. Because I did spend several hours writing it, and I did put a lot of thought into it, and it was not in any meaningful way generated by AI.
garrinm··on “Next-token predictor” is the wrong mental model for LLMs
The distinction I perhaps didn’t make clearly enough is that I’m not really debating the concept of prediction at inference time, although, as I pointed out elsewhere, I think that’s the less interesting interpretation of what “prediction” means.

What’s more interesting to me is its application at training time. In reinforcement learning, there is no ground-truth next token to predict.

So if you’re comfortable calling Deep Blue a “next move predictor,” then I think it’s perfectly consistent to call an LLM a “next token predictor.” But I think it’s more useful to think of Deep Blue as evaluating the value of possible moves. roughly, how likely they are to lead to winning.

And I think effectively the same distinction applies here.

garrinm··on “Next-token predictor” is the wrong mental model for LLMs
It was written by a human. There are AI edits but it’s very much a human composition. Perhaps a bit sloppy.
garrinm··on “Next-token predictor” is the wrong mental model for LLMs
Yes, I think that’s a good explanation. There are really two sides to it.

There’s the mechanical, inference time, autoregressive, one-token-after-another side, which I’m not going to argue isn’t prediction. I just think that’s a relatively uninteresting use of the word “prediction,” because it’s effectively a system predicting its own output.

The more interesting question is what happens at training time. As you describe, reinforcement learning allows the model to learn to output things that it never could have learned simply by predicting what appears in the training corpus.

More concretely, in reinforcement learning there are no ground-truth next tokens to predict.

In supervised machine learning, “prediction” usually means there is some ground-truth label that will eventually be revealed. The model predicts what that label is, the difference between the prediction and the truth gives you a loss, and you learn from that.

But in reinforcement learning, there is no ground-truth action waiting to be revealed. The model chooses an action, observes the consequences, and learns from the reward. To me, that’s a meaningfully different thing from prediction.

garrinm··on Show HN: Stack Error – ergonomic error handling for Rust
Thanks for the insight, I wasn't aware of `track_caller`. I'll definitely be looking into this. I was scratching my head trying to figure out how to make file and line number usage consistent and customizable, this looks like the answer!

You're also right that this will pretty much eliminate the need for macros.

That's also a very key insight about Display vs. Debug printing. I'll be looking into that as well.

Thank you for the thoughtful reply.

garrinm··on Show HN: Stack Error – ergonomic error handling for Rust
Anyhow still makes things easier for application development. The main drawback is that the resulting error type doesn't implement std::error::Error, so it's not suitable for library development (as pointed out in the anyhow documentation). Stack Error is a bit less ergonomic, but suitable for library development.
garrinm··on Show HN: Stack Error – ergonomic error handling for Rust
I played around a bit with SNAFU a couple of years ago, but I'm haven't worked deeply with the library so there might well be some features I'm not aware of.

I think SNAFU is more like a combination of anyhow and thiserror into a single crate, rather than Stack Error which leans more heavily into the "turnkey" error struct. Using the Whatever struct, you get some overlap with Stack Error features:

- Error message are co-located.

- Error type implement std::error::Error (suitable for library development).

- External errors can be wrapped and context can easily be added.

Where Stack Error differs:

- Error codes (and URIs) offer ability for runtime error handling without having to compare strings.

- Provides pseudo-stack by stacking messages.

Underlying this is an opinion I baked into Stack Error: error messages are for debugging, not for runtime error handling. Otherwise all your error strings effectively become part of your public interface since a downstream library can rely on them for error handling.

garrinm··on Stack Error – Pragmatic error handling for Rust
Stack Error is a pragmatic error handling library for Rust that provides helpful messages for debugging, and structured data for runtime error handling.

Features:

- Informative error messages: stack error messages and optionally add file/line context to your messages. This helps convey not just what went wrong, but also how it went wrong, making debugging faster.

- Programmatic error handling: include optional error codes and URIs for robust runtime handling.

- Library-Friendly: define custom error types easily while staying compatible with Rust’s error ecosystem.

If designing good error structures for your projects slows you down, and you need something more library-friendly and structured than anyhow, Stack Error might be what you’re looking for.

garrinm··on Functional semantics in imperative clothing
This gets discussed in the Roc community. They are exploring designing a language without higher kinded polymorphism.

Here's a snippet from the Roc FAQ.

> It's impossible for a programming language to be neutral on this. If the language doesn't support HKP, nobody can implement a Monad typeclass (or equivalent) in any way that can be expected to catch on. Advocacy to add HKP to the language will inevitably follow. If the language does support HKP, one or more alternate standard libraries built around monads will inevitably follow, along with corresponding cultural changes. (See Scala for example.) Culturally, to support HKP is to take a side, and to decline to support it is also to take a side.

https://www.roc-lang.org/faq.html#higher-kinded-polymorphism

garrinm··on Show HN: Chat with GPT about medical issues, get answers from medical literature
Yes, Clint is a proof-of-concept, meant to showcase this use case of LLMs more than anything else.
garrinm··on Show HN: Chat with GPT about medical issues, get answers from medical literature
Unfortunately you can ask Clint to tell you just about anything. But fortunately it will at least try to tell you that some things are less plausible than others.
garrinm··on Show HN: Chat with GPT about medical issues, get answers from medical literature
Clint should be used only to research information. It provides links to resources. It uses Retrieval Augmented Generation which is less prone providing incorrect information, though it can still happen.
garrinm··on Show HN: Chat with GPT about medical issues, get answers from medical literature
Clint should not be used for diagnosis. Only for personal information. Like an web search but more interactive.
garrinm··on Show HN: Chat with GPT about medical issues, get answers from medical literature
Thank you! That's great feedback. Clint is very much a proof-of-concept, I'm sure this idea can be taken much further when done properly. But to get a "pretty neat" from an RN feels like an accomplishment xD.

I see what you mean about going down the wrong path. I can explore a couple of modifications to help it have better context of the conversation as a whole.

garrinm··on Show HN: Chat with GPT about medical issues, get answers from medical literature
Clint could be short for CLINician Tool. I didn't put much effort into this name xD.
garrinm··on Show HN: Chat with GPT about medical issues, get answers from medical literature
Hi, thanks for the comment.

I put this project pretty quickly and I don't want to pretend there is tremendous depth behind any of the decisions I made :/.

For now the only source is the Stats Pearl book published on ncbi.nlm.nih.gov (the only place this is mentioned is here: https://github.com/clint-llm/clint-cli/blob/main/README.md#u...). It contains about 11,000 peer reviewed articles about anatomy and conditions: https://www.ncbi.nlm.nih.gov/books/NBK430685/. The copyright terms are CC BY-NC-ND 4.0. I might add some Wikipedia articles to this in the future.

I chunk the documents by section, and embed only the first 2048 tokens that fit in the OpenAI embeddings. I'm using OpenAI for embedding as opposed to something like all-minilm-l6-v2 because I don't want to have to ship a model to the clients (transfer times could be large and supporting this would increase the complexity of the library).

I didn't experiment with different chunk sizes, and I suspect something smaller would be more beneficial as you point out. But it would also complicate the logic, and most choices I made in this project were to remove complexity and get this done quickly. If I revisit this I might chunk by paragraph on your advice :).

RAG is indeed what is being used. But it a few different ways. The diagnoses are refined using a pretty straightforward RAG prompt: consider these notes ... consider this diagnosis ... can you improve on it etc.

But in a way the entire program is RAG-based. In most prompts some documents are added to the system message for context. It's not clear that the information in the documents is always used, but based on a bit of experimentation it seems to improve various responses.

I have no plans to fine tune. I'm not sure how beneficial would be fine tuning here. The model needs a fair bit of general knowledge to reason about descriptions of symptoms. Fine tuning could over-specialize it. And hallucinations could come up even with fine-tuning, so you would probably want a RAG-like prompt to get it to focus on real details.

This this is very much a hobby, so I haven't dug deep enough to look into other models. But I'd be _very_ curious to see how GPT 3.5 with RAG compares to vanilla MedPALM. In my experience GPT 3.5 can reason quite well about with the right documents in the context.

garrinm··on Show HN: A decentralized semantic web built atop Activity Pub
I agree the description wasn't quite ready for prime time. I will work on the documentation and adding examples.
garrinm··on Show HN: A decentralized semantic web built atop Activity Pub
It's using Activity Stream (https://www.w3.org/ns/activitystreams/v1) which can be represented in RDF or JSON-LD formats. In this case, JSON-LD is used, but this can be easily converted to RDF.
garrinm··on Show HN: A decentralized semantic web built atop Activity Pub
To me it's still unclear if having a message traverse the entire graph is good or bad. Something like this is required otherwise this will really just be a chat app, and probably not a very good one at that. What I mean is that if you can see only messages directly addressed or shared with you, well that's just Signal.

The current implementation traverses the graph of follows, which has some nice properties. Even if you had 10x as many bots as users on a platform, they can all post and share and like etc. But if no one from your social graph is following those bots, you won't see any of their content. Of course all it takes is one person in your network to follow a bot, and now you're exposed to all that spam. In that case you might want to unfollow this non-discerning friend. And the threat of being unfollowed will create incentives for people to make meaningful follows in the first place.

Some more granular tools will need to be built to help refine this. For example, you might not want to see things more than 3 steps away in your graph. There is also the concept of flagging messages which is not currently implemented, but will allow one user to stop a message from spreading to their followers.

As for the echo chambers, I'm not sure yet what this will look like. In the physical world there are soft echo chambers. People make like-minded friends and join like-minded organizations. To some extent this is good, when comparing to the alternative where everyone has to listen to anyone which would create a lot of conflict. Taken to the extreme though it does seem to make people intolerant to differing ideas.

The question I am asking myself (and for which I do not have an answer) is: will a platform like Chatter Net encourage people to progressively discover new ideas (the more people you follow, the more variety of content you will receive), or will it give the tools for people to lock out any competing views.

garrinm··on Show HN: A decentralized semantic web built atop Activity Pub
I have not seen this, thank you for sharing! I will have a look, it does seem interesting.
garrinm··on Show HN: A decentralized semantic web built atop Activity Pub
There are lots of similarities indeed. No idea exists on its own, and it is not surprising that there are many others trying to solve the same problems and arriving to similar conclusions. I would consider Chatter Net to be a parallel experiment exploring similar ideas.

I think I can also address the elephant in the room: Bluesky is built by a team of (very likely talented) individuals. Chatter Net is, well, Chatter Net xD. The atproto project is very interesting to me, and part of the reason for pushing Chatter Net so early was to make sure its have a chance to be heard before the space becomes too crowded.

At a high level I can spot a few differences in the approaches:

- Chatter Net is smaller project focused much more on the data model, whereas atproto is more of a holistic solution specifying not just the data model but also how it should be shared and stored to some extent.

- Chatter Net builds on the Activity Pub (and JSON-LD) standard whereas atproto introduces a new data format.

The core ideas atproto are very similar to the ideas in Chatter Net. Those ideas are extended to touch on some complicated topic that Chatter Net is just starting to address (their big world vs. small world discussion).

However, I'm not entirely sure yet how decentralized Bluesky (atproto) will be. In security the system is only as strong as its weakest link, and I think ultimately whether Bluesky is decentralized, federated, or run by a consortium will depend on:

- how personal data servers are created / managed - how easy it is in practice for a user move between apps, servers etc. The signing key is not controlled by the user, but they do control a recovery key. If a platform changes direction or becomes otherwise incompatible with a user, I can imagine this system working out quite well, or it becoming akin to asking for your post history from one platform, and then manually uploading it to another and having to rebuild your network etc.

garrinm··on Show HN: A decentralized semantic web built atop Activity Pub
For authenticating messages, Chatter Net uses Linked Data Proofs (https://www.w3.org/TR/vc-data-integrity/#proofs). The JS client uses this implementation: https://github.com/digitalbazaar/crypto-ld. And the Rust server uses this one: https://github.com/spruceid/ssi.

I hadn't seen ucan yet, I think the space of JWT adjacent protocols is growing, and I'll be interested to see where it all goes!

garrinm··on Show HN: A decentralized semantic web built atop Activity Pub
Unfortunately not at this time, this is very early days for the project.
Page 1 of 2Next →