Turing awardees republished key methods and ideas without credit
people.idsia.ch
people.idsia.ch
I can see how citations can get very complex for the purposes of scientific credit, traditionally there's a self-policed aspect to it in every academic community; however, in this century as advanced research gets more complex and globalized, it might be time for a more rigorous and neutral process of doing so.
Schmidhuber is only the worst of many abusers of the citation system. Everyone knows that anonymous reviewers are more likely to approve of people who cite their work. Someone is politically powerful? Better cite their paper, even though you don't really think it's a good paper and you didn't use it.
Schmidhuber is cited far more often than he should be, because people just don't want to deal with this crap, and the cost of sticking in one more citation is very low.
The audacity of deciding on behalf of readers whether information is valuable to them when publishing an academic paper is quite an interesting concept to me.
I'm not published in journals, but I do have a "professional" writing background. There's only a few ways I can imagine to read into your statement. Either you think you're above citation, or so much smarter than the reader you're qualified to decide on their behalf whether it's relevant, or simply hiding it for nefarious reasons.
Any of these options, or even the perception of them, is precisely the type of behavior fostering anti intellectualism and prejudice for academia.
It's also quite ironic you repeatedly mention "politics" while your entire argument seems to hinge on seeking to deny empowering further those already in power. Isn't that exercising political capital?
These are very good points.
> ... the worst of many abusers of the citation system.
Is this ad hominem argument by user "lacker" meant to distract from the omissions of the awardees Bengio, Hinton, and LeCun? Especially the first two got tons of citations for work that should have credited Schmidhuber's lab: the analysis of vanishing gradients in neural networks, the principle of generative adversarial networks, attention in neural networks, distilling neural networks, speech recognition with LSTM neural networks, self-supervised pre-training, and more.
The disputes with LeCun are more recent and of lesser magnitude IMO.
The reason that this is not the norm in academia is because there is a substantial number of researchers who care very, very much about academic integrity and even more so about getting their work properly recognised, just like good old You_again.
Btw, "everyone knows" is a running joke in Game of Thrones, it's what the Dothraki always say to show that they're pre-scientific barbarians who believe whatever they like, or at least that's how I get the joke.
But as I said before, we should take the long view. Globalized scientific research ought to have a more neutral way of assessing citations. The 20th century way of doing it may not be the right way any more. I don't see anyone else suggesting a problem of this scope.
You haven't even read the paper, have you? Otherwise you'd see that it's Hinton and Bengio who are cited far more often than they should be. Just look at disputes B1, B2, B5, H2, H4, and H5 to see how they republished parts of his work again and again without citing it. No honest scientist can approve of something like that.
Another aspect that makes it difficult is that different people even in the same field use different terms for the same thing. As an example, engineering optimal control folks like to optimize cost functions, economists might refer to that as the utility instead, and RL folks like to use rewards (which is just a negative of a cost and functionally equivalent). That makes it difficult to run a search. I’ve come across a whole slew of new papers to read on topics just by changing out key terms for synonyms.
A reviewer may suggest other works to include in your discussion. Sometimes the recommendations are appropriate, other times it’s a backhanded way to fish for citations and the works suggested aren’t actually relevant. I would say harping over not including one is poor form unless it is so similar that is important for you to explain the novel aspect in your paper.
> BTW, I committed a similar error in 1987 when I published what I thought was the first paper on Genetic Programming (GP), that is, on automatically evolving computer programs[GP1][GP] (authors in alphabetic order). At least our 1987 paper[GP1] seems to be the first on GP for codes with loops and codes of variable size, and the first on GP implemented in a Logic Programming language. Only later I found out that Nichael Cramer had published GP already in 1985[GP0] (and that Stephen F. Smith had proposed a related approach as part of a larger system[GPA] in 1980). Since then I have been trying to do the right thing and correctly attribute credit.
Source: https://people.idsia.ch/~juergen/deep-learning-miraculous-ye...
He put you through the wringer for not citing him in the announcement of the title?
- Gary Marcus
- Juergen Schmidhuber
- Pedro Domingos
- Max Tegmark
- Eliezer Yudkowsky
Some context for people unfamiliar with ML research: the author, Schmidhuber, is well known for claiming that he should get credit for many ML ideas. Most ML researchers think that:
- He doesn't deserve the credit he claims, in most if not all cases.
- There's a few cases where his papers should have been cited and weren't. That's fairly common.
- People do not get much credit for formulating an abstract idea in a paper or implementing it on a toy problem. Credit belongs to whoever actually makes it work.
- Credit assignment in ML is not perfect but roughly works.
That is according to whom? Is it a rule you just came up with or accepted practice? And if it's accepted practice, in what community is it accepted practice? Because where I publish and review there's really no such rule and credit belongs to the people who deserve credit for the work they've done that was useful to others.
Certainly agree. The point is that coming up with the idea, writing it as an equation, or an architecture diagram in a paper, is a small fraction of the effort that goes into making the idea work in a model showing good performance on real life datasets.
For example, just taking a random paper that Schmidhuber claims should give him credit for GANs, https://people.idsia.ch/~juergen/FKI-126-90ocr.pdf hopefully you can easily see that a lot of work would be needed to turn this into a realistic image generation model. And that is, even if you admit that the idea is strongly related to GANs, which I'm not convinced of but won't spend time on.
> Credit belongs to whoever actually makes it work. >> That is according to whom? Is it a rule you just came up with or accepted practice? And if it's accepted practice, in what community is it accepted practice?
It is accepted practice in the ML community. If it weren't, Schmidhuber wouldn't be complaining.
Deep learning is a relatively unexplored field and there are many open mathematical and scientific questions to ask that involve only model equations or contrived datasets. Novel theoretical results are not just about some architecture idea but about proving facts that can be useful for understanding how the model class would perform in different scenarios. Which in turn can help shape the search space for applied work.
Additionally, I don't think credit assignment should be so discrete. 100% agree that vomiting out vague ideas shouldn't grant claims to credit, but academic science much too often gives only a single author the "real" credit.
Incidentally, in other fields the person who actually makes it work very well may not be the person that receives this credit. Like biology can involve a lot of hard manual work (that isn't really intellectual) in order to realize a project plan. It varies how much of the credit those people receive, and I'm not even sure how much they should receive. This topic is extremely nuanced.
"Introductory theoretical work in GAN was done by Schmidhuber [1], but it was not until large experimental efforts [2,3,4] on image generations that the power of GANs was revealed."
> "the inventor of an important method should get credit for inventing it. She may not always be the one who popularizes it. Then the popularizer should get credit for popularizing it (but not for inventing it)." Nothing more or less than the standard elementary principles of scientific credit assignment.[T22] LBH, however, apparently aren't satisfied with credit for popularising the inventions of others; they also want the inventor's credit.[LEC]
I don't have to deal with citing papers, but I once had to deal with people pitching me ideas, wanting me to sign an NDA, in exchange for 50% of the revenue after I did all the actual work. Just out of curiosity, I signed one once. It was a fart app, IIRC. They thought a fart app needed an NDA, and that I'd then go do all the work and give them 50% because they "had the idea". It was so laughably sad.
If you think these ideas are valuable, I have a beautiful clock for you. It is right twice a day. You'll have the same problem: you won't know when it's right. You'll need someone else's work to tell that.
There's a spectrum of ideas, from groundbreaking to "dime a dozen". In tech startups, and in almost all of computer science, most ideas are a dime a dozen, and the value is in the execution.
But clearly, some ideas are groundbreaking. Einstein rightfully gets the credit for an on-paper hypothesis that wasn't proved until decades later via a chain of critical discoveries and experimental innovations by other people. It's legit to call it Einstin's relativity, and not Mossbauer/Hay's relativity.
This ... idea that Schmidhuber is an ideas man who's never done any real work is Hinton's allegation, and it's clearly designed to misrepresent both Schmidhuber and his work in order to discredit his complaints. And I'm sorry to say that people on HN have fallen for it hook, line and sinker, I guess because that's what social media says.
Btw, the point I make, that you don't get published in machine learning without beating some benchmarks and establishing a new state of the art, I can attribute that to none other than Hinton himself, in an interview with Wired, whence I quote, by the by:
>> What we should be going for, particularly in the basic science conferences, is radically new ideas. Because we know a radically new idea in the long run is going to be much more influential than a tiny improvement.
https://www.wired.com/story/googles-ai-guru-computers-think-...
So that's the guy accusing the other guy of being nothing but an ideas man and that you don't need to cite someone who first came up with an idea, saying that "new ideas" are important.
But that's just Hinton presenting things just the way he likes. Now ideas are important, now they're not, as he pleases.
(1) Ideas are a dime a dozen, and making it work or bringing it to fruition, is the important thing
and
(2) The idea is the important thing; the specific implementation by someone doesn't matter, as they're just doing what the idea creator or discoverer laid out for others to follow.
Sometimes I feel like there's a fundamental paradox there that arises a lot in numerous areas of work, business, and economics.
/s
That deserves a source. Especially for "all cases"; I don't think anyone who understands machine learning could read some of his earlier papers and still think Ian Goodfellow invented GANs.
ex. in the article: "Goodfellow eventually admitted that my PM is adversarial...but emphasized that it's not generative. However, [it] is both adversarial and generative (its generator contains probabilistic units)...It is actually a generalized version of GANs."
When you're at "actually, probabilities means generative, and actually you know what, even my initial claim was too specific: turns out its a generalized version of GANs", all in service of arguing a paper should have been cited in another paper, years after the other paper has been published, there's not much room for sympathy.
Why are some people here even debating the generally recognised rules of scientific publishing mentioned in the paper:
> The deontology of science requires: If one "reinvents" something that was already known, and only becomes aware of it later, one must at least clarify it later, and correctly give credit in all follow-up papers and presentations.
His lab is excellent and was easily Europe's best deep learning lab for decades before it blew up.
Some of his complaints are valid too. European labs often get ignored, and he has been sidelined despite being one of the most important people in deep learning himself.
But man doesn't know when an argument runs out of gas. His claims get grander with every passing year.
He would've just been the 'get off my lawn' grandpa of deep learning, but he somehow comes across as even more insufferable than that.
I wonder if 2023 schmidhuber was created because the polite one from a decade ago was ignored. A sort of evil phase, if you will.
I feel bad for him. He did get passed over of some deserved awards and recognition. But he reeks of resentment and thats never a good look.
It's a terrible look. We're in the middle of one of the biggest gold rushes in tech history and he's wasting time complaining when he claims to be one of its pioneers? That effort is much better invested in building stuff but I suspect he's fallen into the classic PI trap of writing grants all the time and leaving the real work to the rest of the faculty, atrophying his skills too much to do anything now that the industry is moving so quikcly.
His recent papers are still cutting edge. He's already solved self-improving AI: https://arxiv.org/abs/2202.05780 , it just needs to be scaled up.
Things also vary from paper to paper. Sometimes the first author just did the actual work for somebody else, and sometimes they also made significant intellectual contributions. (If the first author is listed as the sole corresponding author, it usually indicates the latter.) Sometimes the last/senior author just brought the money in, sometimes they were primarily mentoring the first author, and sometimes they were the driving force behind the project.
The PhD student needs the street cred a lot more than a tenured PI.
I once had a paper where I shared first authorship with four other people. That implies that the project was large enough that different people were in charge of different subprojects. Which in turn implies that the senior author must have made major contributions by being in charge of the entire project.
I know senior authors who are last on a paper just because they're senior and realize they (1) didn't really contribute much at all (they might have just inserted themselves on a paper), and (2) invoke the old meaning of last author knowing that the meaning has changed, so they end up being "the senior author" who gets a lot of credit just because they're senior.
Even producing the data takes on new meaning in an age of open science where datasets are distributed freely. What's the difference between citing an original study paper to give credit to the study PIs, and having them as a last author? Should someone who generates a dataset be last author on every paper using that data?
Sigh. Academics is so broken.
In ML, NLP, and many Humanities too, the supervisor (supervising professor/postdoc, lab head, PI) is put last regardless of contribution. The rest of the author list is ranked in descending order for contribution. Often the last author's contributions are very limited to obtaining funding or proof-reading.
Now this practice is controversial with many venues stating that obtaining funding and only supervising is not a valid reason for authorship, but in my experience this practice doesn't die out.
This is just plain wrong. No working version without the idea.
I independently arrived at[0] something close to Max Tegmark's idea of the Mathematical Universe, almost nobody noticed and fewer still cared because I published it as a LiveJournal blog post whereas he fleshed it out into a whole book.
I didn't get credit because I didn't do the hard work that deserves credit, I had the flash of inspiration and stopped after a few paragraphs of mediocre student philosophy.
[0] and possibly predated, but I lost track of the date format when shifting from LJ to WP: https://kitsunesoftware.wordpress.com/2018/08/26/mathematica...
Anyone who cares about academic integrity should at least not attack someone complaining of plagiarism. That sort of attack is the academic equivalent of blaming the victim. If even half of Schmidhuber's accusations have a basis that's still a major academic scandal of epic proportions.
What about cases when accusation was unreasonable and not matching reality?
Which paper was that; I've always found his writing relatively clear. Are you sure it wasn't just the case that you didn't understand the paper?
B: Priority disputes with Dr. Bengio (original date v Bengio's date): B1: Generative adversarial networks or GANs (1990 v 2014) B2: Vanishing gradient problem (1991 v 1994) B3: Metalearning (1987 v 1991) B4: Learning soft attention (1991-93 v 2014) for Transformers etc. B5: Gated recurrent units (2000 v 2014) B6: Auto-regressive neural nets for density estimation (1995 v 1999) B7: Time scale hierarchy in neural nets (1991 v 1995)
H: Priority disputes with Dr. Hinton (original date v Hinton's date): H1: Unsupervised/self-supervised pre-training for deep learning (1991 v 2006) H2: Distilling one neural net into another neural net (1991 v 2015) H3: Learning sequential attention with neural nets (1990 v 2010) H4: NNs program NNs: fast weight programmers (1991 v 2016) and linear Transformers H5: Speech recognition through deep learning (2007 v 2012) H6: Biologically plausible forward-only deep learning (1989, 1990, 2021 v 2022)
L: Priority disputes with Dr. LeCun (original date v LeCun's date): L1: Differentiable architectures / intrinsic motivation (1990 v 2022) L2: Multiple levels of abstraction and time scales (1990-91 v 2022) L3: Informative yet predictable representations (1997 v 2022) L4: Learning to act largely by observation (2015 v 2022)
I think there is a reason why the ACM Turing awardees have never tried to defend themselves by presenting facts to the contrary: because they can't.
This might get interesting:
> The "Policy for Honors Conferred by ACM"[ACM23] mentions that ACM "retains the right to revoke an Honor previously granted if ACM determines that it is in the best interests of the field to do so." So I ask ACM to evaluate the presented evidence and decide about further actions.
He was famous well before that for having one of the best labs in Europe, and many key papers in early deep learning. He's only famous for the beef among people who aren't familiar with his earlier work and the huge contributions it made to the field.
It sounds like he published some theoretical musings back in 1990s without any real practical implementation that did anything useful and since then has run around accusing AI researchers who actually produced concrete research and techniques to get are actually in use today of plagiarism.
On the other hand, if his claims are verified, some of these research “discoveries” are simply common sense and would occur to most people working on the subject. HLB were awarded to a good extent because they worked on deep learning at the right time. Deep learning became hugely practical, and outperformed the state of the art in many applications. HL also worked for major companies.
He published working models back then, the problem was compute power was very limited. In the past decade deep learning took off, people took his models, renamed then and ran them on vastly more powerful computers, to great success, then failed to cite him.
Don't you know that billions of people are using his work on a daily basis on their smartphone? CV: https://people.idsia.ch/~juergen/cv.html
> I wanted to popularize neural language models by improving Google Translate. I did start collaboration with Franz Och and his team, during which time I proposed a couple of models that could either complement the phrase-based machine translation, or even replace it. I came up (actually even before joining Google) with a really simple idea to do end-to-end translation by training a neural language model on pairs of sentences (say French - English), and then use the generation mode to produce translation after seeing the first sentence. It worked great on short sentences, but not so much on the longer ones. I discussed this project many times with others in Google Brain - mainly Quoc and Ilya - who took over this project after I moved to Facebook AI. I was quite negatively surprised when they ended up publishing my idea under now famous name "sequence to sequence" where not only I was not mentioned as a co-author, but in fact my former friends forgot to mention me also in the long Acknowledgement section, where they thanked personally pretty much every single person in Google Brain except me. This was the time when money started flowing massively into AI and every idea was worth gold. It was sad to see the deep learning community quickly turn into some sort of Game of Thrones. Money and power certainly corrupts people...
Reddit post: "Tomas Mikolov is the true father of sequence-to-sequence" https://www.reddit.com/r/MachineLearning/comments/18jzxpf/d_...
But yeah, you honestly have to wonder what it is like day-to-day working with someone with this level of delusions of grandeur.
It's not delusions of grandeur; anyone with half a brain who read his early papers would see he clearly came up with the idea of a GAN well before Ian Goodfellow.
Has he explained his reasons for thinking there is some conspiracy? Otherwise it reflects badly on his assessment of himself, or possibly his mental state.
---
to schmidhuber
When you publicly claim that someone else's idea that is remotely resembling your own is stolen from you.
---
schmidhuber
1) To interject for a moment and explain how one's recent popular idea is a few transformations away from your 1991 paper.
2) To miraculously produce fifty years of relevant literature after someone claims to trace the origin of an idea in a particular work.
---
schmidhubered
Being "schmidhubered" looks something like this:
1) Invent something brilliant that no one cares about. Experience derision.
2) That thing becomes popular years later. Someone else is given credit for inventing it. That person appears in the New York Times and is declared smartest person alive.
3) Go on a campaign explaining the situation and how you are the rightful inventor and thus the rightful Smartest Person Alive.
4) Everyone accuses you of being a sore loser and no one takes you seriously.
5) A verb is named after you.
---
[a] https://www.urbandictionary.com/define.php?term=to+schmidhub...
[b] https://www.urbandictionary.com/define.php?term=schmidhuber
[c] https://www.urbandictionary.com/define.php?term=schmidhubere...