836 karma · joined December 17, 2025
I mean you could show people random markov chain gibberish in 1996 and they would swear they found intelligent meaning in it
I think that has been well proven so far
Just scrolled now and saw like 5
How is that different than https://www.nobodywho.ai/posts/jev-in-25-lines/
Couldn't you always get probabilities for a fixed set of tokens without training a new model?
In the sense that if I hire a competent person, I can tell them something and they will deterministically do it
For an LLM I have no mechanism to even update the weights
All you can do is play around with context, which is like taping a post-it note to someone's desk
Imagine you had an employee where you had to tape 1000 post-it notes to their desk
There is no RNG involved when I decide to push vs pull the unlabeled door to my building every morning, it becomes deterministic because its baked into memory
You can put stuff in context to deal with this but you can't do that for everything, its not practical and you would blow the context window
Lame, my interest drops to 0 whenever I find out a project is vibe-coded
Because why would I be impressed if I can just go prompt it myself?
If its vibe-coded its just another display of claude code's abilities, which its like...duh at this point
That means the only thing interesting is the idea itself, which in this case...
No you're falling victim to the common programmer fallacy that "my use case is everyone's use case"
These things exist because people had different use cases and priorities over the years
As humans we don't have our memory reset multiple times per day
You don't actually use the "next token" that the model chooses
That's what the author is doing in this part
token_ids = [model.tokenize(text=label.encode(), add_bos=False)[0] for label in labels]
choice_logits = numpy.asarray([logits[token_id] for token_id in token_ids])
logprobs = choice_logits - numpy.logaddexp.reduce(choice_logits)
probabilities = numpy.exp(logprobs)
This works because the model is always producing probabilities for all tokensSo its still fairly involved and low level but a lot less manual typing
I also don't use auto-mode when they write code, I review the diffs immediately as they come and approve/deny as I find it stops the agents going too far down the wrong path.
I think the term automated programming captures it best
This is such an insane take I see all the time from self-driving boosters
If a self driving car glitches out and crashes in some edge case pathological scenario we don't just accept that as totally fine because its hidden under big statistics
The reason why a crash happened does matter, its not just about aggregate statistics
As a thought experiment if I have a perfect self driving system but I add some code that purposefully crashes 1 in 10 million rides are you ok riding in it since the aggregate statistics look good?
I think large amounts of compute are big strategic asset for a country and governments will partner with the labs
The thing I produce does not replace demand for the original though?
I can’t recite the original for a million people
* You own it forever
* Nothing to charge
* Don't have yet another screen to stare at
* Its the definitive way to collect books (similar in the sense of vinyl for music)
* Going to used bookstores is awesome
* Seeing your whole collection physically is awesome and can be a conversation starter
* Graphic novels/comics are still better physical