Everything in life is variable reward. You invite a friend over, they might accept or they might not. Drive to work, traffic might be good or might be bad. You ask a colleague to finish a task, they might do it or might not or might do a good job or might not.
Everything is variable reward. Is everything gambling?
Sugar is the original point of the reward system!
Do hard work (takes time), get dopamine for successful completion.
Find berries, taste sweet (hopefully safe), eat all, get calories. Doordash Krispy Kreme instead = few too many calories.
(I’m no Luddite in the sense popularly thought of them pre-‘22 [1], though we have to watch skill atrophy)
[1] regressionist? Decelerationist, too loaded perhaps. Someone remembers or knows the word…
regarding your examples, i think the difference is that with ai, you’re literally sitting in front of a machine, pressing a button, and (almost instantly) getting a result that, if not desired, can immediately be tried for again. you even spend “tokens” to do this, and at least in my native language, “token” brings to mind the coins you’d stick in a slot machine
If you sit at a slot machine and pump quarters into it, each “turn” is independent. You spin and you win or lose. It’s pure chance and there is no destination. You execute the exact same action over and over and hope random chance brings you more money.
If you sit down in front of a coding harness, the progress is incremental and directed. You ask for a thing, the LLM produces something that is hopefully close to what you wanted. You give it more direction to prod it closer to the end state you want. You are not executing the same action, but incrementally nudging it in the right direction. I’ve literally never restarted from the same initial state with the same prompt and hoped for a different result and I don’t know why anyone would. Rarely I’ve thrown away the progress made and started over but always with a very different prompt that includes learnings from the failed attempt.
most people don’t use ai this way though, and i still feel like the end-psychological reward mechanism is very, very similar to gambling regardless of how well one utilizes it (and this is even more obvious with image generation as you chase that perfect output)
perhaps it’s better to compare it to gacha than slots?
How do they use it? Surely no one is just repeating the same prompt over and over (except as a Ralph loop perhaps, which is automated). I’m really struggling with the notion that most people just throw the same prompt repeatedly hoping it eventually works. Because that doesn’t sound like gambling. It sounds crazy (and frustrating).
> and i still feel like the end-psychological reward mechanism is very, very similar to gambling regardless of how well one utilizes it
In the sense that you get a dopamine reward when you succeed, sure, but I get the same reward when I code by hand and achieve a successful result.
> and this is even more obvious with image generation as you chase that perfect output
This is fair, because sometimes with image generation the same exact prompt will produce very different output. This is becoming less true as the models get better and it becomes more effective to direct image generation iteratively than to keep starting from scratch with a barely tweaked prompt.
And sometimes the debugging has already veered completely off course at the beginning so it's futile, but each time the fleeting hope that just a few thousand more tokens will magically fix it tempts you to keep going a little longer.
as an example, i was using chatgpt a few weeks back to help me remember the name of a painting i’d seen about a decade ago. i could recall the general shape of the subject and that it was europeanish, but nothing else. after seven turns or so it finally got it, and honestly, the relief of finally remembering the name felt like, well, hitting a jackpot
i’ve had a similar feeling of success after trying to get it to give a comprehensible answer when asking it for a solid counter-argument to philosophical questions. it is indeed often crazy and frustrating
> but I get the same reward when I code by hand
i have only done very simple coding work with llms, so that may be why our ideas differ about the feeling of reward. this is where the comparison to gacha makes more sense than slots, since when you’re coding, you still get a reward each turn whilst chasing the final/desired result
I understand the joy of success but I fail to see how this is gambling. I could have an equivalent conversation with a friend (more likely about a movie in trying to remember than a painting, but still) and get the exact same type of iterative “no, not that one, it was more like X” and feel elated when my friend finally realizes I’m taking about a scene from Hot Tub Time Machine.
This isn’t gambling in any meaningful sense.
i wasn’t trying to convince you, i’m just explaining that this is gambling in a meaningful sense to some people, especially with how turn-based and unpredictable the whole system is. whether it is actually gambling (semantically, legally, ontologically?) isn’t really interesting imo. what’s interesting is that it’s structured similarly and feels nearly identical to some people
but then, i also feel like there’s an element of gambling in the example of you talking with your friend, though i think it would be better illustrated if it was a conversation with a random person
I’m working from a place where my employer pays for my tokens so I’m also not spending anything except my time. Maybe if I were, it would feel more like gambling.
Not everyone has the desire to work around the system, and many are diametrically opposed to the concept of AI. They get this perception that it's a slot machine because of that inconsistency, and then do the human thing of assuming that other people must just be flawed if they're different from them. They're "addicted to gambling".
Obviously, things have changed. Open models can still be like that, but are often so fast and cheap at iterating it doesn't matter. SOTA models aren't perfect, but are to the point that they're generally much better than the average developer.
But once that perception set in and the meme spreads, it's really hard for some to break out of it. Especially at the pace AI development has been moving. It's just that simple.
well, no. If you work overtime and get paid overtime, you are not gambling and that is not a variable reward.
Humans engage more with rewards that are intermittent and variable. Like Futurama's scene from 'The Scary Door' where the character says "A casino where I'm winning, I must be in heaven! A casino where I always win, that's boring, I must really be IN HELL!". A constant predictable reward is boring, less engaging. So if you know you get no overtime, but sometimes your boss rewards you with $5 coffee voucher, sometimes a free pizza dinner, sometimes double-time pay for the time worked or a half-day off, now you might be gambling 1hr overtime for an intermittent variable reward.
> "Drive to work, traffic might be good or might be bad."
Good traffic is not a "reward" for driving to work(!) and you have to drive to work regardless so you are not risking anything [you might be risking your life, but you are not making a choice which can reward you with good traffic]. You might say that going a different route is a choice and a gamble which could reward you with good traffic, but traffic engineering does not work that way because if there was a consistently low-traffic route, everyone else would take that route until it was no faster than any other route. Traffic will generally be the predictable and similar every day, plus 'arriving at work early' is not much of a reward.
Perhaps but predictable outcome is a very desirable quality. No one wants a hammer that sometimes drives nails and sometimes doesn’t. All of the current harness engineering work is about squeezing predictability out of the LLM.
> Good traffic is not a "reward" for driving to work(!)
Like hell it’s not. I drove into work last Friday and there was no traffic because of the holiday weekend. It was amazing. Had me considering whether Friday should be one of my standard RTO days.
In what way was amazing no-traffic "a reward"? What system was rewarding you for what change in behaviour?
> In what way was amazing no-traffic "a reward"? What system was rewarding you for what change in behaviour?
What does this mean? Are you asking me to describe the dopamine system or are you implying that rewards have to be driven by some external system’s goal?
How is it a environmental catastrophe level, datacentre requiring, bullshit model is so much worse than something that runs on my workstation and doesn't cost us a ha itable planet?
Either they're fucking with us serving 8b models at scale or china really has the AI race in the bag so much so that they can openly release what the USA jealously guards.
They moan and complain about china copying from them (while doing the same), but if that's the case in full, then why are the chinese models better? you dont copy bad work and come out ahead.
(By the way this is not to say that all nomads are low-skilled or even low-sw-skill, not by a long shot. I’ve met plenty of CS-degreed nomads who I treat as the wizened experts that they are, from whom deferentially elicit their pearls of hard-won software project wisdom. It was one of these even who, non-derisively, recognized the ‘pulling the slot machine lever’ reflex among vibecoders after I posed the apparent addictive tendency. On later describing this to a longtime friend deeply conversant in social science, he immediately responded with a nod and the phrase ‘variable reward response’.)
I got so sick of all this at some point that I slowly stopped doing anything that wasn't my job. But then AI got better and better and I realized it was the ultimate unblocker. When that dreaded malaise started creeping in signaling it was a project's end because I didn't want to waste any more of my life dealing with bullshit orthogonal to what I was trying to do, I'd give it to the AI. It felt like a miracle the first time this worked, and it still does. If we were previously equipped with shovels to dig through bullshit, we now have a fully automated Bagger 288.
The reward schedule now isn't variable anymore; the chance that I finish something in a good state is 100%. I can focus on the parts I actually enjoy - architecting the broader system, making the parts mesh together in a sensible way that's easy to work with and has some mathematical elegance to it, hand coding the bits I want to be really specific about (but now without the endless frustration of bugfixing or import errors and edgecases being immediately discovered, thanks to the AI).
I was working on a side project recently. I had spent months designing the data model in my spare time, thinking through how to make it as elegant and durable to change as possible in the long term, since (if I launched it) the repercussions for getting it wrong would be significant.
Once I had a working design, it probably would have been several more months to build a working prototype and start testing it.
Instead, Claude knocked out the prototype for me in an afternoon. And it immediately became clear that it didn't work: not because the data model didn't solve all the problems I wanted it to solve, but because it didn't fit the shape of how I quickly learned a normal person would need/want to interact with the product. I was so focused on the long term, that I never thought about what the first five minutes of a user with hands on the thing would need. And the changes needed would be significant.
Maybe there's some variable reward mechanism. But I sure was glad to be able to pull that particular slot machine handle and learn that than waste even more of my time on what was a dead end.
Also, my coming from being a non-coder, I have a lot of appreciation for the possibilities for project failure one way or another due to a data model or project schema being wrong, even though I still only have a superficial understanding of what either of those concepts even are. . My question is, does your conception of data models in the abstract come from a formal academic course, like an algo’s & data structures course, or from trade-knowledge acquired through the practitioner grapevine?
As far as I can remember, this has been fairly common in tech (or at least where I've worked) for some time, though it's possible that my memory here is a bit fuzzy. I will say that my spouse has often commented that "Claude talks like you", which I suspect is because it is trained on a lot of language specific to the tech world.
>My question is, does your conception of data models in the abstract come from a formal academic course, like an algo’s & data structures course, or from trade-knowledge acquired through the practitioner grapevine?
I have worked for over a decade on the telemetry for a specific major product that most people have probably used, and shepherded it through a major re-architecture, so most of this is from my career. I came in right after it was built and witnessed the pain as things had been layered on over a long period of time as the product evolved, so I've just witnessed all the pointy bits where naive early decisions can come back to bite you later.
At this point, what do the words even mean? Your own patience and available time are always completely random and fairly distributed across a large enough data set?
> Maybe I'd waste hours down the wrong rabbit holes trying to find a library that worked for my use case. Maybe I'd waste a day trying to get an API to do something it turned out it couldn't do. Maybe I'd have to redo my entire approach because of some factor I hadn't considered.
Our ignorance isn't random chance. As we research and experiment, we reduce the problem area.
Predicting the time a task will take is impossible. Something that sounds like a 5 minute script can turn into a month of banging your head against unknown unknowns. I lose my patience when the afternoon I allocated is getting overrun by nonsense and I'm missing out on other things I wanted to do or household maintenance.
> Our ignorance isn't random chance. As we research and experiment, we reduce the problem area.
Every thought we have has random chance to be wrong despite our conviction that it's correct. Descartes' Evil Demon plays his tricks on all of us. How many times have you typed some line of code only to realize it was obviously wrong afterwards? Even for simpler matters we "hallucinate" all the time. I was deep in thought trying to help someone come up with an acronym the other day and felt convicted that "Goal Oriented Augmented Retrieval" worked for GOAL until I said it aloud.
Our thoughts and actions are consistently wrong some portion of the time because our meat computers are not perfect positronic brains running prolog. We put cereal in the fridge and say "you too" to the waiter. Every thought we put down or action we take is a gamble on the soundness of the thought or action.