HNHacker News
TopNewBestAskShowJobs

D-Machine

830 karma · joined November 22, 2021

submissionscomments
D-Machine··on Experts Have World Models. LLMs Have Word Models
> You, on the other hand, claim that there are infinite firsthand sensory experiences (maybe we can call them qualia?) that fall in between the cracks of language and are rarely communicated (though we use for that a wealth of metaphors and synesthesia) and can only be understood by those who have experienced them firsthand.

Well, not infinite, but, yes! I am indeed claiming much world models are patterns and associations between qualia, and that only some qualia are essentially representable as or look like linguistic tokens (specifically, the sounds of those tokens being pronounced, or their visual shapes if e.g. math symbols). E.g. I am claiming that the way one learns to e.g. cook, or "do theoretical math" may be more about forming associations between those non-linguistic qualia than, say, obviously, doing philosophy is.

> I'm not sure they constitute such a big part of our thought and communication

The communication part is mostly tautological again, but, yes, it remains very much an open question in cognitive science just how exactly thought works. A lot of mathematicians claim to lean heavily on visualization and/or tactile and kinaesthetic modeling for their intuitions (and most deep math is driven by intuition first), but also a lot of mathematicians can produce similar works and disagree about how they think about it intuitively. And we are seeing some progress from e.g. Aristotle using LEAN to generate math proofs in a strictly tokenized / symbolic way, but it remains to be seen if this will ever produce anything truly impressive to mathematicians. So it is really hard to know what actually matters for general human cognition.

I think introspection makes it clear there are a LOT of domains where it is obvious the core knowledge is not mostly linguistic. This is easiest to argue for embodied domains and skills (e.g. anything that requires direct physical interaction with the world), and it is areas like these (e.g. self-driving vehicle AI) where LLMs will be (most likely) least useful in isolation, IMO.

D-Machine··on Experts Have World Models. LLMs Have Word Models
The multimodality of most current popular models is quite limited (mostly text is used to improve capacity in vision tasks, but the reverse is not true, except in some special cases). I made this point below at https://news.ycombinator.com/item?id=46939091

Otherwise, I don't understand the way you are using "conscious" and "unconscious" here.

My main point about conscious reasoning is that when we introspect to try to understand our thinking, we tend to see e.g. linguistic, imagistic, tactile, and various sensory processes / representations. Some people focus only on the linguistic parts and downplay e.g. imagery ("wordcels vs. shape rotators meme"), but in either case, it is a common mistake to think the most important parts of thinking must always necessarily be (1) linguistic, (2) are clearly related to what appears during introspection.

D-Machine··on Experts Have World Models. LLMs Have Word Models
Sort of, but the images, video, and audio they have available are far more limited in range and depth than the textual sources, and it also isn't clear that most LLM textual outputs are actually drawing too much on anything learned from these other modalities. Most of the VLM setups are the other way around, using textual information to augment their vision capacities, and even further, most mostly aren't truly multi-modal, but just have different backbones to handle the different modalities, or are even just models that are switched between with a broader dispatch model. There are exceptions, of course, but it is still today an accurate generalization that the multimodality of these models is kind of one-way and limited at this point.

So right now the limitation is that an LMM is probably not trained on any images or audio that is going to be helpful for stuff outside specific tasks. E.g. I'm sure years of recorded customer service calls might make LMMs good at replacing a lot of call-centre work, but the relative absence of e.g. unedited videos of people cooking is going to mean that LLMs just fall back to mostly text when it comes to providing cooking advice (and this is why they so often fail here).

But yes, that's why the modality caveat is so important. We're still nowhere close to the ceiling for LMMs.

D-Machine··on Experts Have World Models. LLMs Have Word Models
> Funny, because riding a bicycle or speaking a language is exactly something people don't have a world model of. Ask someone to explain how riding a bicycle works, or an uneducated native speaker to explain the grammar of their language. They have no clue

This is circular, because you are assuming their world-model of biking can be expressed in language. It can't!

EDIT: There are plenty of skilled experts, artists and etc. that clearly and obviously have complex world models that let them produce best-in-the-world outputs, but who can't express very precisely how they do this. I would never claim such people have no world model or understanding of what they do. Perhaps we have a semantic / definitional issue here?

D-Machine··on Experts Have World Models. LLMs Have Word Models
You clearly have no idea about the basics of what you are talking about (as do almost all people that can't grasp the simple distinctions between transformer architectures vs. LLMs generally) and are ignoring most of what I am saying.

I see no value in engaging further.

D-Machine··on Experts Have World Models. LLMs Have Word Models
Let's be more precise: LLMs have to model the world from an intermediate tokenized representation of the text on the internet. Most of this text is natural language, but to allow for e.g. code and math, let's say "tokens" to keep it generic, even though in practice, tokens mostly tokenize natural language.

LLMs can only model tokens, and tokens are produced by humans trying to model the world. Tokenized models are NOT the only kinds of models humans can produce (we can have visual, kinaesthetic, tactile, gustatory, and all sorts of sensory, non-linguistic models of the world).

LLMs are trained on tokenizations of text, and most of that text is humans attempting to translate their various models of the world into tokenized form. I.e. humans make tokenized models of their actual models (which are still just messy models of the world), and this is what LLMs are trained on.

So, do "LLMS model the world with language"? Well, they are constrained in that they can only model the world that is already modeled by language (generally: tokenized). So the "with" here is vague. But patterns encoded in the hidden state are still patterns of tokens.

Humans can have models that are much more complicated than patterns of tokens. Non-LLM models (e.g. models connected to sensors, such as those in self-driving vehicles, and VLMs) can use more than simple linguistic tokens to model the world, but LLMs are deeply constrained relative to humans, in this very specific sense.

D-Machine··on Experts Have World Models. LLMs Have Word Models
> In practice it would make heavy use of RL, as humans do.

Oh, so you mean, it would be in a harness of some sort that lets it connect to sensors that tell it things about its position, speed, balance and etc? Well, yes, but then it isn't an LLM anymore, because it has more than language to model things!

D-Machine··on Experts Have World Models. LLMs Have Word Models
> It has little to do with the reality of that language or the correctness of your model of it, but rather with the need to train realtime circuits to do some work.

To the contrary, this is purely speculative and almost certainly wrong, riding a bike is co-ordinating the realtime circuits in the right way, and language and a linguistic model fundamentally cannot get you there.

There are plenty of other domains like this, where semantic reasoning (e.g. unquantified syllogistic reasoning) just doesn't get you anywhere useful. I gave an example from cooking later in this thread.

You are falling IMO into exactly the trap of the linguistic reductionist, thinking that language is the be-all and end-all of cognition. Talk to e.g. actual mathematicians, and they will generally tell you they may broadly recruit visualization, imagined tactile and proprioceptive senses, and hard-to-vocalize "intuition". One has to claim this is all epiphenomenal, or that e.g. all unconscious thought is secretly using language, to think that all modeling is fundamentally linguistic (or more broadly, token manipulation). This is not a particularly credible or plausible claim given the ubiquity of cognition across animals or from direct human experiences, so the linguistic boundedness of LLMs is very important and relevant.

D-Machine··on Experts Have World Models. LLMs Have Word Models
It isn't a misnomer at all, and comments like yours are why it is increasingly important to remind people about the linguistic foundations of these models.

For example, no matter many books you read about riding a bike, you still need to actually get on a bike and do some practice before you can ride it. The reading can certainly help, at least in theory, but, in practice, is not necessary and may even hurt (if it makes certain processes that need to be unconscious held too strongly in consciousness, due to the linguistic model presented in the book).

This is why LLMs being so strongly tied to natural language is still an important limitation (even it is clearly less limiting than most expected).

D-Machine··on Experts Have World Models. LLMs Have Word Models
Completely relevant, because LLMs only "somewhat model" humans' "somewhat modeling" of the world...
D-Machine··on Experts Have World Models. LLMs Have Word Models
LLMs are language models, something being a transformer or next-state predictor does not make it a language model. You can also have e.g. convolutional language models or LSTM-based language models. This is a basic point that anyone with any proper understanding of these models would know.

Even if you disagree with these semantics, the major LLMs today are primarily trained on natural language. But, yes, as I said in another comment on this thread, it isn't that simple, because LLMs today are trained on tokens from tokenizers, and these tokenizers are trained on text that includes e.g. natural language, mathematical symbolism, and code.

Yes, humans have incredibly limited access to the real world. But they experience and model this world with far more tools and machinery than language. Sometimes, in certain cases, they attempt to messily translate this messy, multimodal understanding into tokens, and then make those tokens available on the internet.

An LLM (in the sense everyone means it, which, again, is largely a natural language model, but certainly just a tokenized text model) has access only to these messy tokens, so, yes, far less capacity than humanity collectively. And though the LLM can integrate knowledge from a massive amount of tokens from a huge amount of humans, even a single human has more different kinds of sensory information and modality-specific knowledge than the LLM. So humans DO have more privileged access to the real world than LLMs (even though we can barely access a slice of reality at all).

D-Machine··on Experts Have World Models. LLMs Have Word Models
Not sure about that, I'd more say the Western reductionism here is the assumption that all thinking / modeling is primarily linguistic and conscious. This article is NOT clearly falling into this trap.

A more "Eastern" perspective might recognize that much deep knowledge cannot be encoded linguistically ("The Tao that can be spoken is not the eternal Tao", etc.), and there is more broad recognition of the importance of unconscious processes and change (or at least more skepticism of the conscious mind). Freud was the first real major challenge to some of this stuff in the West, but nowadays it is more common than not for people to dismiss the idea that unconscious stuff might be far more important than the small amount of things we happen to notice in the conscious mind.

The (obviously false) assumptions about the importance of conscious linguistic modeling are what lead to people say (obviously false) things like "How do you know your thinking isn't actually just like LLM reasoning?".

D-Machine··on Experts Have World Models. LLMs Have Word Models
The amount of faith a person has in LLMs getting us to e.g. AGI is a good implicit test of how much a person (incorrectly) thinks most thinking is linguistic (and to some degree, conscious).

Or at least, this is the case if we mean LLM in the classic sense, where the "language" in the middle L refers to natural language. Also note GP carefully mentioned the importance of multimodality, which, if you include e.g. images, audio, and video in this, starts to look like much closer to the majority of the same kinds of inputs humans learn from. LLMs can't go too far, for sure, but VLMs could conceivably go much, much farther.

D-Machine··on Experts Have World Models. LLMs Have Word Models
It's also important to handle cases where the word patterns (or token patterns, rather) have a negative correlation with the patterns in reality. There are some domains where the majority of content on the internet is actually just wrong, or where different approaches lead to contradictory conclusions.

E.g. syllogistic arguments based on linguistic semantics can lead you deeply astray if you those arguments don't properly measure and quantify at each step.

I ran into this in a somewhat trivial case recently, trying to get ChatGPT to tell me if washing mushrooms ever really actually matters practically in cooking (anyone who cooks and has tested knows, in fact, a quick wash has basically no impact ever for any conceivable cooking method, except if you wash e.g. after cutting and are immediately serving them raw).

Until I forced it to cite respectable sources, it just repeated the usual (false) advice about not washing (i.e. most of the training data is wrong and repeats a myth), and it even gave absolute nonsense arguments about water percentages and thermal energy required for evaporating even small amounts of surface water as pushback (i.e. using theory that just isn't relevant when you actually properly quantify). It also made up stuff about surface moisture interfering with breading (when all competent breading has a dredging step that actually won't work if the surface is bone dry anyway...), and only after a lot of prompts and demands to only make claims supported by reputable sources, did it finally find McGee's and Kenji Lopez's actual empirical tests showing that it just doesn't matter practically.

So because the training data is utterly polluted for cooking, and since it has no ACTUAL understanding or model of how things in cooking actually work, and since physics and chemistry are actually not very useful when it comes to the messy reality of cooking, LLMs really fail quite horribly at producing useful info for cooking.

D-Machine··on Experts Have World Models. LLMs Have Word Models
Fun play on words. But yes, LLMs are Large Language Models, not Large World Models. This matters because (1) the world cannot be modeled anywhere close to completely with language alone, and (2) language only somewhat models the world (much in language is convention, wrong, or not concerned with modeling the world, but other concerns like persuasion, causing emotions, or fantasy / imagination).

It is somewhat complicated by the fact LLMs (and VLMs) are also trained in some cases on more than simple language found on the internet (e.g. code, math, images / videos), but the same insight remains true. The interesting question is to just see how far we can get with (2) anyway.

D-Machine··on India's female workers watching hours of abusive content to train AI
From the article you linked:

"Some workers report nightmares in which the violent images they reviewed replay in gruesome loops. Others experience intrusive flashbacks while riding the bus or shopping for groceries. Over time, many describe a numbing of their emotions-a flattening of joy, sorrow, or empathy-because the only way to cope is to feel nothing at all."

and bla bla bla. "Some", "Others", "many"... Wikipedia itself couldn't generate a better example of "weasel words". Like I said, we all know some people are traumatized, but all that matters is the amount. What if some is less than 1%? Less than 5%? More than 50%? The answer matters, but you are not providing answers to this.

Also, learn to basic science. Anecdotes are not data.

No, I am not watching a documentary for a collection of anecdotes, I clearly explained why documentaries don't count as serious sources of info when trying to accurately quantify things.

EDIT: And if you really know that traumatization from images on computer screens is "well-studied", you can surely link to one to three of such studies, rather than lame documentaries telling cherry-picked sob-stories.

D-Machine··on TikTok's 'addictive design' found to be illegal in Europe
Generalizations are just doing better than random guessing by just paying attention to base rates, or, equivalently, describing the majority / bulk of distributions. This isn't meaningless, it is just simple statistical summary in common language.

Also, with modern tools like sentiment analysis and other semantic analysis tools enabled by LLMs, it is pretty trivial to do empirical tests of claims like this (further proving how empty the claim that generalizations about a thing are "meaningless").

Being for paternalistic, censorious policies that restrict software behaviour so broadly is very obviously not "Hacker" (even if you agree with these policies), and this would have been far more surprising 10 years ago to see on HN. This is a "generalization" of the ideological shifts of HN, yes, but a meaningful one (which of course people are free to agree or disagree with).

D-Machine··on TikTok's 'addictive design' found to be illegal in Europe
> Do you have family? My cousins, aunts, even my mom is on it.

Yes, I too am deeply traumatized from having family and friends... shudder... using Social Media.

Anyway, not sure of the relevance of the question. Not sure how "time spent" on a thing is proof of its badness either, but then, people comparing TikTok to heroin are clearly not generally interested in things like clarity and quantification.

D-Machine··on TikTok's 'addictive design' found to be illegal in Europe

    > I don’t get it, is this some kind of gotcha?
Only for people who think comparing TikTok to heroin is some kind of gotcha.

    >Have you walked down the skid row of any large city? Heroin and well other drugs now are a problem, saying otherwise is delusional. Those people need help.
For sure. Two minutes from where I live, at a main intersection, they hang emergency Naloxone injection kits, in public, where anyone can grab them, on the trees and walls of buildings. I presume so addicts can save each other in cases of accidental overdoses.

Of what relevance was this all to TikTok again? And why are we comparing scrolling a phone app to literal actual heroin? Even when, empirically and factually, heroin is in fact not addictive for the majority of people?

Comparisons between TikTok and heroin are deranged and simplistic, but this is made all the more embarrassing when you realize that a dance with heroin is in fact more likely than not to just be... not the thing everyone is afraid of?

D-Machine··on I'm going to cure my girlfriend's brain tumor
And sometimes that advocacy is harmful, desperate, arrogant flailing—against the reality one knows is true with overwhelming likelihood—manifesting as "advocacy" or "will" that destroys so many chances for fully experiencing the reality of the precious, remaining, time one has (or one has with one's partner).

NOTE: This is not me disagreeing at all, just your point moved me to make the obvious counterpoint, having been through all this myself very literally and very recently. I know firsthand how important the advocacy is, but also how often it causes nothing but harm. There is a real tricky balance between agency vs acceptance when you've truly lost control of things, like in these cases.

D-Machine··on TikTok's 'addictive design' found to be illegal in Europe
Yup, hence why only a reckless person or fool would try it.

But, since only a minority of people get addicted to heroin (i.e. the evils of heroin are overstated), and since no one is actually seriously arguing that viewing TikTok is as risky (23-38% chance after exposure) as trying heroin, or has as bad side effects, I think it reveals that comparisons to heroin use in arguments against TikTok are hyperbolic and disconnected from reality, by empirical data.

D-Machine··on I'm going to cure my girlfriend's brain tumor
Yup, you got it. I'm a survivor (so far) from a relapsed cancer myself. People have no idea of the kind of insanities you are willing to pursue in such desperate situations. Grace and forgiveness is the right approach here.
D-Machine··on TikTok's 'addictive design' found to be illegal in Europe

    > You also probably don't use heroin. Everyone knows it's a bad idea
About 1–12 months after using heroin, only 23%–38% become addicted [1]. Occasional and controlled heroin users do in fact exist and are documented [2]. And, most famously, the use of heroin by American soldiers during the Vietnam war was largely situational [3].

So what "everyone knows" here is not very impressive. I still very strongly believe you'd be a fool or at least reckless to try heroin, but it really isn't the bogeyman people want it to be.

[1] https://jamanetwork.com/journals/jamapsychiatry/fullarticle/...

[2] https://www.drugsandalcohol.ie/3906/

[3] https://ajph.aphapublications.org/doi/pdf/10.2105/AJPH.64.12...

D-Machine··on India's female workers watching hours of abusive content to train AI
As I've said, images can be horrific, and some people can be traumatized by them. That must not be dismissed.

However, It is also important to carefully and properly quantify these things and not sensationalize. You've linked a 50+ minute documentary, without comment, that seems to prove that one person hired to curate content can become traumatized by that process. I can't be certain that is what it is about, because I will not waste time watching documentaries (the vast majority of which are outright propaganda or incredibly biased, while pretending to be objective), but still, I've no doubt the general claim is true, since I never claimed or believed otherwise.

But you've not provided meaningful statistical or scientific evidence properly quantifying such harms in general.

D-Machine··on EU bans infinite scroll and autoplay in TikTok case
Yup, agreed. I don't trust these companies for a second, but I also don't want overreach.

Also, though this isn't super relevant, I am not actually personally too worried if they are really using behavioural psychologists. Most psychology is such junk science and most research psychologists struggle with such basic math and stats that I am extremely skeptical they'd be gaining anything from having a person like that on hand.

I think the reality is that the addictiveness is engineered with simple A/B and other kinds of testing and data, and this kind of engineering / research is better done by people with other qualifications. This would be the thing to look for and demand documentation and evidence about, if we were serious about limiting this kind of thing.

D-Machine··on Forcing Rust: How Big Tech Lobbied the Government into a Language Mandate
Getting some strong ChatGPT vibes from the overall sectioning and some stylistic flags, e.g. the "This isn't X, it's Y" meme appears many times as an intro to paragraphs or sections, e.g. "This isn’t a conspiracy. It’s something more mundane and more durable: structural incentive alignment". There are lots of (spaced) em-dashes, and the overall rhythm, tone, and length of things is very ChatGPT.

The way the references are sort of lazily clustered also to me strongly looks like they could have been pasted from a ChatGPT Extended Thinking sidebar: if you were doing this kind of research organically, you'd just more clearly be able to link your references to each clause that is is relevant, with appropriate enumeration. I would bet with about 90% certainty that this was mostly done via ChatGPT with Extended Thinking.

There is also not much discussion and/or consideration that, well, perhaps Rust is being adopted because it actually does provide strong guarantees in combination with things like being modern, having cargo, being fast, and etc. bla bla bla. The proposed alternatives of ADA, modern C++, and hardware optimization just feel laughably out of touch.

It's a weak post overall, but I do really appreciate the documentation of some of the financial ties involved in supporting / boosting Rust.

D-Machine··on TikTok's 'addictive design' found to be illegal in Europe
Yes the argument quality has become poor too, and there is less nuance than before.

And like obviously society benefits from some paternalism for things like this, and we really do need to see good, concrete recommendations be proposed.

Imagine if HN were discussing the merits of things like "interrupt every X minutes of infinite-scrolling with reminders / popups / forced X/10 minute breaks" or other actual concrete, balanced solutions to such problems. This would introduce the nuance and make things interesting again.

But it is increasingly the case I find I need to look elsewhere for such discussions.

D-Machine··on My AI Adoption Journey
The difference is that I have a few sites and resources I already know that are NOT useless. With an LLM output, I have to check and verify every time, and, since LLMs are based on junk, almost always produce junk. But with the trusted sites, I do not have to check, and almost always get something decent and/or close to authentic!

The difference is between a trusted source that is good most of the time, vs. an LLM recipe that is trash 99% of the time.

EDIT: If you haven't visited any / all of the sites / sources I mentioned, check them out! They are really good, especially SeriousEats if the recipe is from Kenji Lopez. Maybe just avoid AmazingRibs, unless you have uBlock installed: they were way ahead of their time, but haven't updated in forever, and clearly have become desperate...

D-Machine··on TikTok's 'addictive design' found to be illegal in Europe
HN is becoming increasingly pro paternalism and censorship (largely invalidating the "Hacker" in its name). I'm not sure if this is natural / organic or some kind of astroturfing, but it sure is odd and doesn't really match the rest of the overall (historical) ethos of the site.
D-Machine··on TikTok's 'addictive design' found to be illegal in Europe
Hard not to think of the "hard times create strong people, strong people create good times, good times create weak people, weak people create hard times" meme here.
← PreviousPage 11 of 20Next →