HNHacker News
TopNewBestAskShowJobs

chpatrick

2,533 karma · joined May 5, 2014

submissionscomments
chpatrick··on I fed 24 years of my blog posts to a Markov model
> Sure, a LLM is a "markov chain" of state space size (# tokens)^(context length), at minimum.

Okay, so we're agreed.

chpatrick··on I fed 24 years of my blog posts to a Markov model
I think you're confusing Markov chains and "Markov chain text generators". A Markov chain is a mathematical structure where the probabilities of going to the next state only depend on the current state and not the previous path taken. That's it. It doesn't say anything about whether the probabilities are computed by a transformer or stored in a lookup table, it just exists. How the probabilities are determined in a program doesn't matter mathematically.
chpatrick··on I fed 24 years of my blog posts to a Markov model
The KV cache doesn't affect it because it's just an optimization. LLMs are stateless and don't take any other input than a fixed block of text. They don't have memory, which is the requirement for a Markov chain.
chpatrick··on I fed 24 years of my blog posts to a Markov model
It's not n sometimes, k tokens some other times. LLMs have fixed context windows, you just sometimes have less text so it's not full. They're pure functions from a fixed size block of text to a probability distribution of the next character, same as the classic lookup table n gram Markov chain model.
chpatrick··on I fed 24 years of my blog posts to a Markov model
I'm not saying it's an n-gram Markov model or that you should store them as a lookup table. Markov models are just a mathematical concept that don't say anything about storage, just that the state change probabilities are a pure function of the current state.
chpatrick··on I fed 24 years of my blog posts to a Markov model
Sure many things can be modelled as Markov chains, which is why they're useful. But it's a mathematical model so there's no bound on how big the state is allowed to be. The only requirement is that all you need is the current state to determine the probabilities of the next state, which is exactly how LLMs work. They don't remember anything beyond the last thing they generated. They just have big context windows.
chpatrick··on I fed 24 years of my blog posts to a Markov model
Not sure why that's contorting, a markov model is anything where you know the probability of going from state A to state B. The state can be anything. When it's text generation the state is previous text to text with an extra character, which is true for both LLMs and oldschool n-gram markov models.
chpatrick··on Zebra-Llama – Towards efficient hybrid models
Then they'll be able to use those datacenters much more efficiently.
chpatrick··on It’s time to free JavaScript (2024)
Sure, the type system isn't perfect but it still beats 90% of mainstream languages in use today.
chpatrick··on It’s time to free JavaScript (2024)
Everything is already written in JavaScript. If we had WASM from the start and dozens of languages with different APIs we wouldn't be better off.
chpatrick··on It’s time to free JavaScript (2024)
TypeScript is a really decent language though, I wouldn't feel happier or more productive using Fortran or whatever. Its type system is actually really powerful which is what matters when it comes to avoiding bugs, and it's easy to write functional code with correct-by-construction data. If you need some super optimized code then sure that's what WASM is for but that's not the problem with most web apps, the usual problem is bad design, but then choice of language doesn't save you. Sure TS has some annoying legacy stuff from JS but every language has cruft, and with strict linting you can eliminate it.

It's also better if there's one ecosystem instead of one fragmented with different languages where you have to write bindings for everything you want to use.

chpatrick··on The Death of Arduino?
And a dev board only costs a couple of dollars on AliExpress.
chpatrick··on What if you don't need MCP at all?
It's also kind of an impossible request given that AI tools are for generic unpredictable input.
chpatrick··on What if you don't need MCP at all?
What does deterministic mean in this case?
chpatrick··on What if you don't need MCP at all?
Why are they nondeterministic? You can use a fixed seed or temperature=0.
chpatrick··on The Case That A.I. Is Thinking
Sufficiently smart auto complete is indistinguishable from thinking, I don't think that means anything.
chpatrick··on The Case That A.I. Is Thinking
If a human uses their general knowledge of electronics to answer a specific question they haven't seen before that's obviously thinking. I don't see why LLMs are held to a different standard. It's obviously not repeating an existing answer verbatim because that doesn't exist in my case.

You're saying it's nothing "special" but we're not discussing whether it's special, but whether it can be considered thinking.

chpatrick··on The Case That A.I. Is Thinking
Because I gave them a unique problem I had and it came up with an answer it definitely didn't see in the training data.

Specifically I wanted to know how I could interface two electronic components, one of which is niche, recent, handmade and doesn't have any public documentation so there's no way it could have known about it before.

chpatrick··on The Case That A.I. Is Thinking
I've certainly got useful and verifiable answers. If you're not sure about something you can always ask it to justify it and then see if the arguments make sense.
chpatrick··on The Case That A.I. Is Thinking
I've definitely had AIs thinking and producing good answers about specific things that have definitely not been asked before on the internet. I think the stochastic parrot argument is well and truly dead by now.
chpatrick··on Claude for Excel
They're not great at arithmetic but at abstract mathematics and numerical coding they're pretty good actually.
chpatrick··on A small number of samples can poison LLMs of any size
Does it work on humans too?
chpatrick··on Qualcomm to acquire Arduino
You can use the exact same Arduino environment with ESP32 for a fraction of the price. A D1 mini dev board with wifi costs $5 (!) on AliExpress.
chpatrick··on Qualcomm to acquire Arduino
ESP stuff is so damn cheap and capable now I'm not sure what you would use Arduino for these days.
chpatrick··on Fire destroys S. Korean government's cloud storage system, no backups available
"secured against cyberattacks or crisis situations with KSI Blockchain technology"

hmmmm

chpatrick··on Beginner Guide to VPS Hetzner and Coolify
Or you can use a more reliable host like Hetzner.
chpatrick··on Beginner Guide to VPS Hetzner and Coolify
Except when their datacenter burns down...
chpatrick··on ProofOfThought: LLM-based reasoning using Z3 theorem proving
Sure, sufficiently advanced dominoes.

https://xkcd.com/505/

We're already at the point where LLMs can beat the Turing test. If we define thinking as something only humans can do, then we can't decide if anyone is thinking at all just by talking to them through text, because we can't tell if they're human any more.

chpatrick··on ProofOfThought: LLM-based reasoning using Z3 theorem proving
Do you understand human thinking well enough to determine what can think and what can't? We have next to no idea how an organic brain works.
chpatrick··on ProofOfThought: LLM-based reasoning using Z3 theorem proving
Everything happens in an opaque super-high-dimensional numerical space that was "organically grown" not engineered, so we don't really understand what's going on.
← PreviousPage 6 of 34Next →