Language homogenization at Harvard
inteoryx.com
inteoryx.com
It also wouldn't surprise me a huge amount if the idea space was constricting, but just looking at the cosine distance is not enough to establish that. There probably are meaningful ways to exploit and analyze the representations inside language model transformers to better capture the idea-space geometry, but that's a big research project.
It’s reinforcement learning over decades
I agree with your general ramble, but found this interesting in context.
The analysis also seems sensitive to the mapping of words to categories. Some kind of robustness analysis to this sensitivity would also be interesting to see.
Actually it IS evidence for idea homogenization. Words represent ideas. Unless you are claiming that the same word is used to represent different ideas, which would be even more confusing than having different words representing the same idea. Therefore fewer unique words => fewer unique ideas.
All named theorems do this in math, continually compressing the lexical space as a way of enabling further out ideas to be even expressed. Granted, math may be the except that proves the rule, but it is an important one.
It is for the word thinkers.
The quality of skating has drastically improved in just a decade. In 2010 the Olympic champion had no quad jumps. In 2022 the easiest jump from the top 3 competitors in their short program was a triple axel - the rest were all quads.
My theory about this is that the advent of smartphones and ubiquitous video has made it much easier to constantly see what your competitors are doing. This not only pushes you to work harder and try new jumps, but also lets you see what new techniques work for other athletes. A couple decades ago people wondered if a quad jump was even possible, now people train going into it that a triple is no longer the expected limit.
In other words, if we could represent skating routines as vectors, would the average cosine distance between all those vectors be increasing or decreasing?
What you see there is what you see in any sport - evolution of it toward better performance.
It's exciting to see transfer learning's effects in near real time, as a lot of that has occurred over my lifetime. I just wish meta learning was more accessible and common knowledge. As it stands, I find myself continually having to do my own research to uncover ways to improve my own improvement.
Consult with a someone in the field.
The trick were really invented not that far ago.
This can naturally lead to lower average "distances" between proposals. In a simplistic example, lets assume 100 proposals existed in the "good old times" and they were different at random. Let's further assume in the "bad new days" people use those old ideas, change/improve upon them just slightly ("add noise"), but some old ideas are less often picked up than others. Say for example the least attractive old proposal is picked up twice, whereas the most attractive old idea is picked up and changed in 100 new proposals. Then, there is a lower average distance between proposals, all the while the total range of ideas has increased.
Because it's Friday and I am waiting for my oven to finish cooking my food, I wrote a small simulation. It's probably full of mistakes and I may have made terrible mistakes in my assumptions, but I thought it's fun:
Yes, that was indeed the point I was trying to convey! In an even sillier example, assume word vectors X, then calculate "proposal by proposal" similarities (i.e. inverse distances). Then duplicated X and concatenate [X,X], recalculate "proposal by proposal" distances (now for twice as many proposals)---those distances must now be less on average because each proposal has at least one "zero distance" neighbor. HOWEVER, why would you assert that the overall "idea space" has been reduced?
I disagree. Words are used to convey ideas, so if the space of words is shrinking, one should assume that the space of ideas is shrinking. It's possible for this not to be the case, but if word-space is shrinking then the burden of proof should be on those who claim that idea-space is not shrinking.
Maybe we could use embeddings / nlp analysis to determine whether idea-space is shrinking. Or just get a bunch of people to read abstracts from different time-periods and rate how similar they are to one another in their semantic content.
1. Language needs to be somewhat homogenous. At an extreme, if the languages spoken by two academics are entirely heterogenous, they (literally) won't understand a word the other is saying!
2. I recently investigated why there were 300 (!!) words I did not know the meaning of in a single George Orwell novel. Google N-Gram Viewer shows around 60-80% of these words to have been in common usage in 1934 when the novel was written, but not in common usage for some decades now [1]. Using these words today would increase lexical diversity, but at the expense of communicative effectiveness!
Jocosely is the only one I had a question over, until I realised it had the root "jocose".
Sometimes multiple words represent the same idea, so in that case homogenization is "good". It is also possible that multiple ideas are represented by the same word, but that is "bad" in an academic context because it leads to ambiguity. Therefore fewer unique words IS actually good evidence that the idea space is decreasing, unless you are saying academic literature is becoming more and more ambiguous.
Words transmit ideas. The difference is notable I’m this data set in particular. I posted a longer version elsewhere in the thread but to summarize.
During this time period the nsf became more strict about explicitly addressing their review criteria. Faculty became trained to use the languag explicitly and connect their ideas to the language of the review criteria. I’m effect, you started to get an ‘api’ where, irregardless of what the idea in the grant is, it’s is clearly and explicitly connected to key nsf terminology. That is a de novo narrowing of the language space no matter the underlying ideas. It’s purpose is definitively about improving communication, not narrowing ideas.
From what I can tell the paper didn’t filter anything in the data. Had they filtered “broader impacts” and “intellectual merit” I would be hard money they would get different results.
The paper graph shows a decrease from like .1 to .075 over 30 years
The blog post shows a decrease from .35 to .1 over 100 years
However, the crimson had a notable drop around 2000 - on the order of the entire decrease in the research study - and then looks fairly stable.
It’s almost like there are policy and effective communication reasons for narrowing your vocabulary to that appropriate for a target audience.
If you look at the graph from the underlying study, there’s a bit of a “shoulder” around 1997. That is telling. 1997 was when the NSF introduced a new, clear and explicit, set of review criteria - broader impacts and intellectual merit [0]. That change alone is likely a cause of significant (Meaningful?) linguistic narrowing. To get NSF funding, researchers now had to explain why their research matters using explicit language that aligned with specific strategic objectives of the funding agency.
Then you add a layer of Goodharts law. During this time period, two other things were also happening. First, University’s increasingly began to rely on external funding - especially public university’s. Second, the field of “faculty development” was increasingly formalizing and offering training and support on things like grant writing. Those trainings include a lot of focus on using normalized, almost shibboleth like, language in grant applications. Ensuring that there is religion of key words so that it is easy for the reviewers to establish that a particular grant application addresses the required review criteria.
So the data set used here is part of the problem - if they had used submitted rather than funded they would likely see different results. Not because ideas are narrowing but because grant applications are basically a human manifested api, and someone tried to actually standardize it.
[0] https://stem.colostate.edu/a-history-of-the-broader-impacts-...
The studies could be showing that we use a smaller dictionary today.
A graph that may get to the heart of your question is something like "Unique word percentage over time" or maybe "What percentage of articles use unique words".
You'd probably see the same effect if the articles don't get any longer, but more of them get written every year. Unique word count will always go up with words produced, even if most words produced are formulaic boilerplate.
Also, the author may be over-interpreting the result: it's odd that very different fields are showing roughly the same change in cosine difference, and it is a measurement over individual words. I think a deeper analysis is needed to figure out in more detail what is going on.
In my point of view this is actually good news; Since I'm not a native english speaker and since this language has become the world's lingua franca it's very nice being able to understand an article just because the author's didn't waste time looking for bombastic synonyms just to sound smarter. I applaud this and projects like Simple English Wikipedia which allows us, non-native english speakers, keep learning.-
Esoteric synonyms, perhaps? :-)
How much of the homogenization is due to “broader impacts” and “intellectual merit”. Which became the default lingo of nsf grants in about 1997
Now he terms them "In order to stay ahead of China we should..."
Basically, communities seem to get into vocabulary and sentence structure ruts. Deviation means your reviewers are going to give you more shit, at least indirectly.
Suppose, for example, that The Crimson (not Harvard btw, it's a student newspaper) runs 10x as many articles this year as last. It's possible you're going to get a huge reduction in cosine distance just by virtue of a few authors producing a lot more content.
At a minimum, we need mean, variance, and number of samples. This doesn't tell you anything about "Harvard", it just tells you about the students who the Crimson choose to publish. There are lots of structural reasons within Harvard that the Crimson has probably stopped being a unified voice of the student body -- but again, that's not what the article purports to show.
Regarding the flaws in the article (phew, I'm glad I didn't quite reach criminal level) - I'm curious which assumptions you think are horrible or what data shoddy. I don't think, for example, that I'm assuming the latent space model is "correct" as you say. I don't think I really have any significant assumptions about the technique or the meaning behind it. I read about the technique in the linked paper and reproduced it in my blog with a different dataset and found a similar result. It's strange to me that the signal produced by this technique is as consistent as it is across the 120 years of data. Beyond that, I'm pretty explicit that I don't know what it means or why it happens.
Regarding the "arbitrarily fit" line - as I say explicitly in the post, that's a regression plot to illustrate the trend.
Regarding the possibilities that The Crimson has more articles per year - it's true that's possible. It's not reality, they run about (for a generous definition of "about") the same number of articles every year. The articles do get longer over time. Either way, it's not clear to me what impact this should have on average cosine distance.
There are a lot of things that I looked at that didn't make it into the blog post. Without including them, then perhaps it looks like I'm cutting corners. If I did include them then I think the blog post would be shooting off in many directions. For example, I considered that political violence might be related - like maybe, in times where there's lots of political violence elite institutions come together and their language becomes more similar. That didn't really pan out though. I graphed a bunch of things that ultimately I decided didn't contribute very much and did not include.
Another way of thinking about it is in the original article Rasmussen (the original author) says "Look at this elite writing in NSF grants. The cosine distance is decreasing over time." I then say "Here is some elite writing - student newspaper at an elite school. Is the cosine distance decreasing there over time too?" And, it is. That's what the blog post is trying to say.
Now, maybe the latent space is "incorrect" - although Rasmussen and I use different embeddings that find a similar trend. Maybe it's not meaningful to use cosine distance in this context. But, it does seem like something has to cause it. Whatever it is and whatever it means, it doesn't look like the kind of thing that happens entirely by chance because it is consistent in different datasets and over many years.
- there's (unsurprisingly) no significant diversity-word change from 1900 to 1940 but a very significant distance drop
- there's a big diversity-word change around ~1990 with no concomitant distance change
Let me quote from the end:
"Another argument against connecting distance and diversity is that distance is on a long running decline from 1900 even for the first four decades while diversity words were basically flat. When diversity words pop in the 90's there isn't an immediate reaction in cosine distance, it's only about a decade later, in 2000, that cosine distance takes a steep drop."
That seems awfully similar to the two points you've raised here.
What I do find a bit distasteful is that you jump in with "your diversity hypothesis is bunk" and accuse me of trying to fit a narrative - without even reading what you're commenting on.
If someone complains about not mentioning variance - it's still implicitly visible by the cloud of dots representing each year around the regression line.
The author should ask the question: What does my "model" actually assign here? Instead, this very core of the work is ignored and the focus rests completely on the numbers put out of a black box...
If the model simply assigns a closer similarity to more modern words (i.e., it would evaluate older words as "weird"), we would expect exactly this outcome, no?
I can put a chart of lexical similarity over a chart of the number of employed airplane mechanics and probably pontificate that by golly, we need to stop fixing airplanes for the sake of intellectual diversity!
And that would be about as well-reasoned.
"The diversity words are new (new-ish). Anything else newish would be inversely correlated with cosine distance too. For example, I repeated the same experiment with "tech" words like Google, YouTube, iPhone, browser, and so on. Those tech words correlated at -0.6 with cosine distance"
Or alternatively, NSF favors grant applications that mention it and toe the party line, so everyone begins to add more and more garbage to their application to optimize for grant acceptance.
Generally if something is so profound that it disrupts the status quo we quickly see it take over. For example the smartphone. But there still isn't a scientific consensus on diversity despite the well-known liberal lean of college campuses. And companies pay lip service and it almost seems that the most successful companies implement DEI programs _after_ they become successful, implying that it provides no competitive advantage.
It does show lexical homogenization. It does show an increase in their shoddily assembled list so-called diversity words. The concluded implications and pretty much every other part is hand waving, assumptions and political insinuation.
Like here where the author disclaims the shoddiness of using word frequency as a measure of politicization because the nuance and subtlety of bias make quantitative measurement difficult (he should apply for a grant to research that)… but then draws the entirely unsupported conclusion that politicization is actually worse than their only data implies:
“Note that word counts is a somewhat crude way of measuring politicization. Bias, particularly in the social sciences, is often subtle, and can apply to the kinds of questions that get asked and the standard of evidence used to accept or reject a hypothesis. Thus, the fact that so many grants contain terms that are in most contexts clearly associated with left-wing political causes likely underestimates the degree of politicization in science funding.”
And the author puts that assumption to work a couple of paragraphs and charts later:
”This report presents direct evidence that scientific funding at the federal level has become more politicized and less supportive of novel ideas since 1990.”
The opening paragraphs don’t even indicate a connection beyond “Look at this trend which I assume means Y. Now look at that trend which I assume means Y. Now look at them together. Vaguely similar, eh? Awfully suspicious, eh?.”
“Taken together, the results imply that there has been a politicization of scientific funding in the US in recent years and a decrease in the diversity of ideas supported.”
How about a real “diversity word” selection criteria instead of just showing some tenuously relevant clustering? How about contextual analysis or explanation of the terms? Language and public discourse have changed a lot since the 90s. To inform and help validate their word selection analyzing ~30 years of data, they used the “DEI terms” diversity, equity and inclusion, but the DEI acronym wasn’t used until about 10 years ago— it could be a coincidence, but could indicate those terms weren’t always an accepted standard — is there evidence other terms weren’t used instead? Is there any evidence those are the best terms to use now? Did 90s political grants use euphemisms or code words for political ideas to send more neutral? Were the terms you selected used by political activists as commonly in 1993 as they are now? Would the declining prevalence of second wave feminists and original civil rights movement activists, to use two prevalent examples of very political movements ubiquitous in academia long before 1990, have been equally political but using different words? How many of those words are central to the research topic and how many are positioning unrelated research as good for society or aware of current trends or more beneficial than they are? What other words or topics showed similar/different results for comparison? So-called inclusion words removed, was the lexical similarity static? During that same period, a small number of ubiquitous spelling and grammar checkers gained sophistication and adoption; writers who disregarded or disabled them probably retired and others probably changed how they wrote; the Internet sharpened the curves of trends and memes and also probably exposed people to better writing practices; writing trends have probably come and gone; educational curricula become more standardized; the Internet facilitated access to far more grant/business writing examples; there may have been particularly influential events or pieces of writing that changed things for other reasons— how did they consider or reason about factors like that? Do Internet-focused terms which would have gained prevalence in the same time frame follow a similar trend line? Any medical terms? BPA? GMOs? Genetic testing? Any other points of comparison at all?
Anyone interested in actually advancing human understanding rather than creating political cudgels could poke holes in this garbage all day. I don’t see how anyone could look at that and think “huh — thoughtful analysis” rather than “huh — that’s some elaborate work to give a tenuous air of legitimacy to a claim they didn’t test presented in a manipulative piece which won’t get much traction beyond a politically sympathetic subreddit.”
You know it can't be right--too good to be true.
> For example, if we wanted to rate "poppycock" we might say that it's a 9 for "old fashioned" and an 8 for "Funny". Poppycock would be 9, 8. Another word, like snail, isn't especially old fashioned or funny. Though, the word has been around a while - we could call it a 5, 2.
This is a sense of "old-fashioned" that I've never heard of. As far as I'm aware, "old-fashioned" is entirely defined by not being in current use. The age of the word is irrelevant, but if it were relevant, and "poppycock" - a word from the 19th century - were a 9 out of 10, "snail" - a word from Old English - would be something more like a 1500 out of 10.
A better way to think about it is that these measurements are my subjective opinion on old timeyness and sillyness via an undefined and intuitive process for assigning values.
It's more interesting than you might think! Turns out snail is a diminutive form of snake[1], from a root referring to creeping over the ground.
The other aspect of judgments of old-fashionedness is that things that are really old-fashioned, like being named Etheldreda, will tend to be rated as less old-fashioned than things that are pretty recent but out of current fashion, like being named Gertrude. The old old things are too forgotten to be "old".
My other question would be "why 'cosine similarity' rather than 'correlation'?". Same thing, but people are a lot more familiar with the term 'correlation'.
[1] OE snaca preserved the /k/ sound, but OE snægl voiced it, and the /g/ then predictably turned into a Y sound (compare "yard" / "day"), giving us the I of modern snail.
As a native English speaker, "poppycock" sounds more "old-fashioned" (maybe "quainter") than "snail," even though "snail" might actually be older.
But see also my observation about the English female names Gertrude and Etheldreda. Which one is more old-fashioned?