HNHacker News
TopNewBestAskShowJobs

_t89y

0 karma · joined March 24, 2013

submissionscomments
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
That's a really interesting thing to point out. NLP doesn't even work on language anymore. If it was adjacent to information retrieval before it is now a subfield of information retrieval. As long as it's grounded in Firth Mode natural language understanding, as it's called, can't really be a semantics.

I tried to create a Kaggle (TensorFlow Hub, TensorFlow Quantum) competition for motivating alternative formalisms but was unable to publish it because all Kaggle competitions must be evaluated with information retrieval metrics. Talk about a one-track mindset!

Today work in NLP advances by ``leaderboards'' and dubious, language-specific evaluation datasets that the same authors stand to benefit from when their proprietary model is praised for doing well on the evaluation criteria they invented a few months back. It validates the price hike for access to their proprietary models.

These formalisms that do work are at odds with Firth Mode, the preferred representation for Google (Stanford, OpenAI), so I guess we should be thankful they're still in the book. If you're interested in language, though, I'd suggest picking up a different book.

_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
They've taken off because they have utility in information retrieval systems. They work for getting info into Google (Stanford) Knowledge Panels. I don't think it really goes any further than that. They are most useful to the few orgs that went from dominating NLP research to controlling it outright by convincing everyone scale is the only way forward and owning scale. Alternatives to word embeddings aren't even considered or discussed. They are assumed as a starting point for pretty much all work in NLP today even though they are as uninteresting today as they were when word2vec was published in 2013. They do not and will not work for language.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
There is an inherent structure in language. Embeddings do not and will not capture it. It's why they do not work. Their ability to form grammatical sentences with high accuracy is part of the illusion that you have been understood.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
Easy there, Firthmiester. I'm familiar with the canon. If getting some desirable behavior in your application is good enough for you then feel free to ignore what I'm saying.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
https://news.ycombinator.com/item?id=39680852
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
The Lambek calculus. Categorial grammars. Meanings are proofs. Not clusters of directional magnitudes in space.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
Having a mid-century theory of natural language semantics isn't necessarily a bad thing. You just have to pick the right one.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
Modeling language in a latent space is useful for certain kinds of analyses and certain aspects of language. It has its place as an empirical tool. That place is not the nuts and bolts of language itself. There are more suitable formalisms for this than directional magnitudes and BPE tiktokens.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
Uh oh. LOL. Got some angry Firthers out there.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
You can't fine-tune for understanding or reasoning. You can't "get better performance" on understanding. You're either equipped for it or you're not.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
For describing semantics in natural language? Pretty much anything else.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
Works for what? The leaderboards? BPE tiktokens, BPE GPT-2 tokens, SentencePiece, GloVe, word2vec, ..., take your pick, they all end up in a latent space of arbitrary dimensionality and arbitrary vocab size where they can be mapped onto images. This is never going to work for language. The only thing the leaderboards are good for is enabling you to charge more for your model than everyone else for a month or two. The only meaning hyperparameters like dimensionality and vocab size have is in their message that more is always better and scaling up is what matters.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
They make a mess of language. They are not a suitable representation. They are suitable for their efficiency in information retrieval systems and for sometimes crudely capturing semantic attributes in a way that is unreliable and uninterpretable. It ends there. Here's to ten more years of word2vec.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
It is true. And if you want to say anything about meaning this isn't even the right math.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
It is meaningless to talk about cosine similarity of sentences, or words, at all. Choose whatever mapping you want. You'll still be in Firth Mode.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
No understanding. Embeddings are a semantically vacuous representation and similarity is a semantically vacuous interpretation.
_t89y··on Is Cosine-Similarity of Embeddings Really About Similarity?
It's definitely not about semantics or language. As far as language is concerned similarity metrics are semantically vacuous and quantifying semantic similarity is a bogus enterprise.
_t89y··on Let's Build the GPT Tokenizer [video]
Thanks for this perspective on the tradeoff between accuracy and efficiency and the insight that an adequately pre-trained model should be in a position to recover lost information from bad tokens.

Tokenization, the gateway to word embeddings, is a means to an end. I'm not suggesting that better tokens are needed or that BPE tokens should be replaced with something else. I'm suggesting that aiming for a distributional semantics is setting the bar pretty low and that there are better places to end up than These Things Are Over Here And Those Things Are Over There Let's Combine Them And See What Happens. I'm expressing disbelief that these representations have been taken at face value and that there has been practically no discussion of applying alternative formalisms which may be more expressive.

Modeling language in a latent space only makes sense for certain aspects of language and certain kinds of analyses. Crucially, you have to have meaningful primitives to begin with. This line of thinking that an understanding of language and an understanding of the world is somehow going to emerge from mapping character spans onto a latent space and combining them with dot product attention is pretty half baked. These systems remain in Firth Mode™.

_t89y··on Let's Build the GPT Tokenizer [video]
2009: Porter stemming with NLTK

2013: LDA with MALLET

2015: spaCy

2018: BERT

2023: GPT-4

2024: every person is an NLP expert in four lines of LangChain code

_t89y··on Let's Build the GPT Tokenizer [video]
Thanks for your reply.

That's my first point. In 10 years we have word2vec, GloVe, GPT-2 and... tiktoken. lol. It's as if directional, numeric magnitudes in an embedding space of arbitrary dimensionality have magically captured or will magically capture the nuances and expressivity of language. Optimization techniques and new strategies for domain adaption are what matters, particularly for mobile devices, on-device ASR and short-form videos.

I don't think robust is a good characterization of clusters of semantic attributes in space or a distributional semantics of language. I'd say crude and without understanding are more accurate descriptions. Capturing semantic properties sometimes is not the same thing as having a semantics.

By targeted improvements you must be referring to domain adaptation and by the default option you must be referring to attention over BPE tokens? You can move directional quantities around in directional quantity space all day. If it results in expected behavior for your application that you weren't getting before that's great. If that's all you want to get out of these models then indeed there's nothing to do here. I'm not after improvements so much as I'm after something that works.

_t89y··on Let's Build the GPT Tokenizer [video]
Thanks for your reply.

It's exactly like lexers for compilers. This parsing strategy coupled with the decision to then map the results into an embedding space of arbitrary dimensionality is why these models don't work and cannot be said to understand language. They cannot reliably handle fundamental aspects of meaning. They aren't equipped for it.

They're pretty good at coming up with well-formed sentences of English, though. They ought to be given the excessive amounts of data they've seen.

_t89y··on Let's Build the GPT Tokenizer [video]
It’s pretty wild how little discussion there's been about the core feature of these models. It's as if this aspect of their development has been solved. Basically all NLP publications today take these BPE tokens as a starting point and if they are mentioned at all they’re mentioned in passing.
_t89y··on Jeff Dean: Trends in Machine Learning [video]
Thanks for your reply.

>>> I think it would help if you link "NeSy computation engine". I'm actually not familiar with this (not in the symbolic world, but interested. Just never had time, so if you got links here I'd personally appreciate it). I can find the workshop but not the engine.

I linked to it in my previous comment. I'm referring to the ``NeSy computation engine'' described here. I didn't know there was a ``NeSy'' workshop and this paper was my first encounter with the term.

I think it's interesting that you mention symbolic world like it is separate from some other world. There's the AI that was and the AI that is today. There's the AI over there and the AI over here. Whenever you hear someone mention symbolic in the context of AI go ahead and grab a chair because immediately after this they are going to talk about cyc and John McCarthy for at least 20 min. If you're lucky they might throw some Prolog in there.

I don't think this is productive and I don't think there is another symbolic world. There is just the world. There are certain things in the world for which a numeric, directional representation makes sense. There are other things for which it makes no sense at all. It's my view that primitives in language are one of these things. Additionally, there are certain places where it makes sense to consider these representational approaches and other places where it only makes political sense. Lastly, there are symbols - atomic primitives - and there are ``symbols,'' objects with vectors in them and who knows what else.

What's striking to me about this paper is the coverage of formal grammars and semantic parsing entirely within the context of domain adaptation. Definitely the best part is the coverage of compositionality (https://ncatlab.org/nlab/show/compositionality) in the context of composing computational graphs. This is striking to me because all of these things (except domain adaptation) are essential to any reasonable theory of meaning but they are covered as if they've been repurposed for the practical application of populating Google Knowledge Panels, which I believe is exactly what happened. Check out the definitions of semantic parser and symbol.

>> Domain adaptation across verticals >>> This is also a bit vague and so I'm not sure what you _specifically_ mean.

Crude semantic attributes pulled from character sequences and mapped onto the latent space of images have utility in business contexts if the mapping for some term sufficiently distinguishes it from the mapping of another term that has the same surface form. It ends there. GloVe was a half-baked representation of meaning in language when it was adapted from word2vec in 2014. GPT-2 grabbed the torch in 2019. It still doesn't work. Well, it sometimes works for adapting a general model to a specific domain such as a business vertical, but only in a crude and superficial way. Note that almost no ML research today discusses this representational issue at all, and that almost all ML research takes this representation as a starting point. If you decide to publish hyperparameters in your paper, such as in an appendix, hyperparameters related to vocab size and the dimensionality of your embedding space often aren't even worth mentioning. That's fine, I guess, because they don't mean anything anyway, but not talking about this, in my view, is not fine.

Check out the Mamba paper for example. Like most of ML research today the focus is on optimization. The representation problem has been solved so there's no need to talk about it: we map everything onto the latent space of images because short-form video content rules the day and that's how dude is gonna hit his 7T: advertising ([link redacted]).

>>> I think the problem is that math is taught by a game of telephone.

I think that, for language, the ML research community is, by and large, not even using the right maths.

>>> But that said, I still think vectors can do a lot. Especially since vectors and functions are interchangeable representations. Though I think we need to do a lot more to ensure that networks are capable of learning things like equivariance and importantly abstract concepts.

Thank you so much for highlighting the important of equivariance. I think this is a crucial concept for work at the cross-modal interfaces, especially in the context of the Curry-Howard correspondence, or, more recently, the Curry-Howard-Lambek correspondence. Right now the ML (CV) research community is labeling nouns with bounding boxes... lol. If that doesn't illustrate the fact that multimodal work is a vision-first enterprise I don't know what will.

>>> I think a bit too exaggerated but hey, I've been known to say that ML research is captured by industry and we're railroading everything. And that it is silly we publish papers on GPT when we don't have the models in hand as it just becomes free work for OpenAI and we can't verify the works because OAI will change things.

Check out the evaluation criteria in that ``NeSy'' paper, especially the metric that's supposed to tell you something about what the system was designed to do. I'm sure OpenAI is happy to have this info about their system.

>>> But I also don't know what you mean by "IR".

Ten years ago I considered NLP adjacent to information retrieval. Today I consider it part of information retrieval. There's very little work published today that suggests otherwise.

>>> Honestly I don't know what you mean by this. But if you are saying that the divide we create like NLP vs CV is dumb, then I'm all with you.

It is not my intention at all to create or highlight any divide. If there is indeed a known divide between CV and NLP I don't know anything about it, I don't want to know anything about it and it's not surprising.

>>> I also think it's silly how we call generative models. Aren't all models generative?

Generative refers to a situation where you begin with a finite set of things and productively form any number of well-formed expressions from these things.

>>> That includes me, and even I have a hard time parsing what you're saying and it doesn't help with the side snipes like URB-E scooters.

I'll take potshots at the Paul Grahams and Steve Jobs of the world every day and not lose any sleep over it. If they take their AirPods out of their ears maybe they'll hear me coming.

>>> But if I'm right and you need more than scale, then we better keep working because I'd rather not have another AI winter.

All I have to say about scaling is that, for language, I hope it's clear by now that more data and more params is not going to improve the situation. I can see how this is almost never the case for vision.

Damn it somebody said AI winter again. You aren't going to start talking about cyc and McCarthy for 20 min now are you?

>>> I also think it is quite odd for these companies to not be hedging their bets a little and more strongly funding other avenues.

The formula works.

>>> At this point, all I'm trying to get people around me in ML to understand is how nuance matters. That alone is a difficult battle. I'm just told to throw compute at a problem and data with no concern to the quality of that data. It does not matter how much proof I generate to show that a model is overfit, as long as the validation loss doesn't diverge, they don't believe me.

I'm interested in learning more about what you mean by nuance.

Probably just needs more compute and data. Just throw some synthetic data in there and call it.

_t89y··on Jeff Dean: Trends in Machine Learning [video]
Thanks for your comment.

Please let me know what's not clear. Still no takers on my ML is CV comment below.

Research on on-device ASR and computer vision is primarily driven by the same organizations that stand to benefit the most from it. It's nearly impossible to talk about machine learning today without talking about computer vision. Just look at the daily papers from Hugging Face or any other outlet. Machine learning is basically synonymous with computer vision and natural language processing is basically synonymous with information retrieval. Research is corporate research. With very few exceptions.

Two popular lines of current ML research, multimodal and this latest neuro-symbolic re-hash, are not about furthering our understanding of what we currently can and cannot do. Not about doing science. They are about maintaining the status quo. They are about short-form video content and Google Knowledge Panels.

Multimodal ML research is a vision-first enterprise. This doesn't make sense for a number of reasons. Here are two: the latent space of images and the representations therein cannot adequately capture the nuances and expressivity of language; language is a more fundamental cognitive process than vision.

And what do you get for language in current ML research? How about a ``neuro-symbolic semantic parser'' that is neither a semantic parsers or symbolic. lol https://arxiv.org/abs/2402.00854v1 What's it good for? Computation graphs, domain adaptation, Google Knowledge Panels.

This is a directed attack on machine learning research taking concepts that could be pursued with scientific merit, such as multimodal perception and neuro-symbolic parsing, but are instead turned into marketing hype and leveraged by the powers that be for the things that keep them in power. My audience is anyone participating in this research.

Y Combinator... that's the mob of onewheels and URB-E scooters dodging human waste in the Tenderloin on their commute from Nob Hill to the Mission, right? Maybe you are referring to another audience.

_t89y··on Jeff Dean: Trends in Machine Learning [video]
lol. Three computer vision researchers dislike this comment. Do any of you want to respond to it?
_t89y··on Jeff Dean: Trends in Machine Learning [video]
Domain adaptation across verticals is the only driver of innovation. Check out the NeSy computation engine. Its ``semantic parsing'' is domain parsing and its symbols are numeric. Scale if you want. It works for images. That's all that matters, right?

Mapping language onto the latent space of images gets you crude semantic attributes sometimes. If you have to push for multimodal out of the gate maybe start with articulatory perception. These LVMs aren't going to cut it.

ML research isn't meant to further your understanding of anything. You can't separate it from corporate interest and land grabbing. It's the same thing re-hashed every year by the same people. NLP is pretty much a subfield of IR at this point.

I love how fast my comment was blacklisted to the bottom of this thread. lol

_t89y··on Jeff Dean: Trends in Machine Learning [video]
Exactly. Google.
_t89y··on Jeff Dean: Trends in Machine Learning [video]
From the GPT-2 paper: ``Also thanks to all the Googlers who helped us with training infrastructure.''
_t89y··on Jeff Dean: Trends in Machine Learning [video]
It's because they don't understand language. You may have been mislead by their ability to generate language.
_t89y··on Jeff Dean: Trends in Machine Learning [video]
There's a typo in the title of the talk but don't worry I fixed it: Trends in Computer Vision

``In recent years, ML has completely changed our view of what is possible with computers''

In recent years ML has completely changed our view of what's possible with computer vision.

``Increasing scale delivers better results''

This is true for computer vision.

``The kinds of computations we want to run and the hardware on which we run them is changing dramatically''

Optimizations on operations for computer vision isn't exactly dramatic change. Who is We?

Trends in Machine Learning, 2010: semantic search for advertising. Trends in Machine Learning, 2024: semantic search for advertising, short-form video content.

[link redacted]

Page 1 of 2Next →