The reanimation of pseudoscience in machine learning
cell.com
cell.com
Clueless ML researchers claim to read this and that from brains. Do they know or care that muscle and eye artifacts overlap most of the EEG frequency range? Do they realize that skin conductivity changes when people react to events? It's often easier for ML model to lean from side channels and skip brain waves altogether. ML can learn to move cursor on the screen by measuring unconscious muscle tension or tracking eye movements trough scalp electrodes.
ML is great for EEG signal analysis, but you do have to know what you are doing.
https://www.scientificamerican.com/blog/scicurious-brain/ign...
> It's often easier for ML model to lean from side channels and skip brain waves altogether.
My favorite example of this is a machine learning model for visually recognizing cancers that seemed to do extremely well... Until they realized it had actually learned that cancerous samples were more likely to have a measuring stick in the photo.
> Clueless ML researchers claim to read this and that from brains. [...] ML is great for EEG signal analysis, but you do have to know what you are doing.
1) If you know what you are doing, it's possible to use ML algorithms to control cursor etc. from EEG coming from brain.
2) It's also possible for ML algorithms to enable paralyzed person (neck down) to to control a mouse or play a video game through scalp eeg sensors using eye and muscle artefacts.
Both are good, but the latter is clueless and would work better with better positioning of electrodes. Track eye and muscle tension directly. Paralyzed people care only that it works.
I think the comments here are arguing past each other. If a machine learning model can correctly classify $attribute by photos in 70% of the cases, that is an academic result.
Should it used for the ever popular pre-crime detection scams or for misclassifying feminine looking men as homosexuals? No, that is nasty and is forbidden in the EU. At least for companies, the security state of course does whatever it wants without any repercussions.
Should it be used in principle? No, classifying humans by a machine algorithm is dehumanizing.
Should research be done at all? If the subjects are volunteers and it is confined to a university, it should be allowed. But researchers should not be surprised it they are confronted with obvious parallels to phrenology. This literally is phrenology.
Phrenology was classified a pseudoscience not because it could lead to socially bad outcomes but because it didn't work; it had no statistical/empirical grounding. If it turns out deep learning really can reliably predict things about people's personality from their faces, that doesn't make it a pseudoscience.
The paper calls out inferring political leanings as an example of pseudoscience. Give me an American’s age, gender, and race (which I can roughly identify by looking) and I’ll tell you, better than chance, whether they’re a trump supporter.
I remember a case where Amazon used a resume filtering bot that systematically rejected female candidates, because of bias in the training data. So we might go from "you can't be free because your skull is too small" to "you can't get the job because the computer says so".
I think we cannot discount the degree to which ideology played a part BOTH in the promotion AND rejection of phrenology. And when the racist ideologies eventually became taboo, anti racist ideology won.
If we consider the actual science, it was probably highly tainted by a desire to show that certain human lineages were superior to others.
But there DOES appear to be a correlation between brain volume and IQ. When controlling for "race", this correlation is typically reported at 0.3-0.4, meaning brain volume accounts for 9-16% of the variance.
However, if we reject "race" as a social construct, and include people of all "races" in our analysis, the correlation goes up to about 0.6, or 36% of the variance [1].
[1] https://emilkirkegaard.dk/en/2023/03/modern-neuroscience-con...
What makes it pseudoscience is that it's not theory-driven. These are statistical models that recapitulate distributions in their training data. It's Brian-Wansink-style p-hacking [1] at a massive scale.
[1] https://www.vox.com/science-and-health/2018/9/19/17879102/br...
Science is primarily empirical, and there's nothing inherently wrong with an effective theory that works, until we find a better, principled theory.
> It's Brian-Wansink-style p-hacking [1] at a massive scale.
Sometimes. Models that generalize are in fact generating theories though.
I see utility in a machine that can help me notice when I'm having a bad day (or a very bad day ie. when having a stroke) or ... help me figure out what the fuck is going on with my totally hypothetical mental illness.
Arguably the current generation of school-age kids (genZ? genA?) already lost it. Because even if they don't have a phone, others do. (Though hopefully this trend can be reversed. Classic bullying at school is bad enough now with cyberbullying becoming the "norm" things are definitely not looking great.)
To be clear ML research has “paper mill” problems but we should be careful that we don’t imply that there are only “rare successes”
There are many many amazing results published at ICLR, NeurIPs, ICML every year that are important developments that are not only research successes but also commercial and open source success stories. For example LoRA and DPO are two recent incredible development and these are not “rare” - this years ICLR had many promising results that will in turn be built on to produce the next “transformer” level development. Without this work there are no transformers.
Even transformers themselves were a contribution whose impact only became valuable through the work of many researches improving the architecture and finding applications (for example LLMs were not a given use case of transformers until additional researches put the work in to develop them)
Man, the author did not beat around the bush, they're straight up naming names! Props to them
Politicians often make the mistake of assuming that if they do something in order to get X, the outcome will be X. Researchers who are more familiar with the methodology than the subject matter often make similar mistakes. The data that is supposed to measure X never actually measures X. It measures something related but subtly different. If you want to make conclusions from the data, you need to understand those differences. You need to understand what the data is exactly and the process how it was collected. Including the details you think are irrelevant but aren't. The last part is particularly problematic for people who are not subject matter experts.
But probably we underestimate the value of our "contemporary contextual richness", ie. relationships/correspondences that are not apparent (not yet known) yet turn out to be important and valuable and easy to comprehend are mostly only possible because we spend our life (mostly pretty successfully) in this extremely complex and ever-changing environment.
AI/ML/LMMs first would need to get up to speed, I guess, to be able to have these insights and be able to provide them at the right time. (Otherwise ... it's probably already in the training data. Or not that deep. Or too deep.)
- these days, everything is called machine learning... AlphaGo is a great AI achievement, but I don't really consider it ML. it's classic AI augmented with NNs iirc. However I'm willing to concede that it's (a|my) taxonomy issue.
- however, being on the receiving end of Stockfish and friends (chess engines) I see "just moves, no insights". Even, for example, the insight that pushing the h-pawn is often better than previously perceived was caused by humans reverse engineering the results.>Machine learning (ML) is a field of study in artificial intelligence concerned with the development and study of statistical algorithms that can learn from data and generalize to unseen data and thus perform tasks without explicit instructions.
So it's about learning via statistics & data. If you don't use data or statistics it's not ML. Chess engines definitely do not fit there. They descend the game tree and evaluate it with an algorithm like minimax. No statistics. No learning. No training (although the static evaluation function might have been trained/tweaked via NNs). I don't know the details of AlphaGo, but I'm guessing it's similar: the concept of the game is hardcoded (game tree,....) while the evaluation of a position is done via NN. the training can be done via games against itself.
As far as I know that's the whole thing that makes it better than the older ones is ML.
https://miro.medium.com/v2/resize:fit:4000/format:webp/1*0pn...
The authors are NOT arguing that ML models cannot work in science, that the examples they mention for pseudoscience have fundamental methodological problems, nor even that classifying people based on ML is incorrect because correlation is not causation.
Rather, they are warning the community about a problem with the epistemic value of results from applied ML, and argue that the problem is a cultural one. In other words, how do we _know_ that ML results are valid science? People working in ML itself are well aware of the garbage-in garbage-out problem of these models, but when these methods leave the field of machine learning and are applied in other sciences the experts in those fields do not know enough about these issues. So, because of the great success of ML models in the past, results from ML classifiers are taken as objective grounds to support dubious hypotheses.
Part of their argument is that because extrapolation from data is seen as having little or even no bias, and because experts of most fields are not also ML experts, such publications with dubious hypotheses bypass the self-regulating mechanism of science, whereby bullshit research is called-out during review.
And this problem, they also mention, is aggravated by the fact that data-driven methods can produce results faster than classical theory-based method hooking into bigger problem of academia that the number of publications is considered a proxy for success.
So given the above I think it is easy to see that this could become dangerous since today science is considered a reliable tool to inform policymaking and political decisions at large. The authors use phrenology as an example to discuss the ethical implication of the issue because it is the most blatant example of socially corrosive pseudoresearch.
What happens if X is bad but right? If you need X to be wrong for the sake of your system of morality then you twist and frame and p-hack and don't publish if you get contrary evidence and in the limit you have destroyed the trustworthiness of your own institution and quite frankly this is where we sit today.
The only principled and sustainable position is that X is bad regardless of whether or not it's true. Then at least you're never in a position where you must live by lies.
It is an *unsustainable* position to hold that something is bad regardless of whether or not it is true. Basing your morality on potentially false information is a guarantee for conflict.
Likewise, if your morality never changes, especially when you learn and grow from experience, you’re stunting your growth as an individual and causing problems for society as a whole. Imagine a child refusing to update their understanding of morality as they age.
On this specific case, if it happened that some physical feature of people predicted violent behavior or whatever, we would probably have to push "violence treatment" into a medicine area. So the fact that it is wrong is extremely important to how we react to it.
The fact that we can now discover the genes behind a particular propensity does not make the actions you take any or more or less of a personal responsibility than they were a century ago.
There are good arguments on either side of the debate but discovering the underlying physical mechanism doesn't make any difference.
Even if we discover, say, a violence gene, we can't usefully test for it because it may be present and inactive or there may be other violence genes as yet undiscovered. It doesn't tell you anything at the individual level.
Genetic fatalism is currently fashionable but it is not a fact.
That doesn't change the fact that if it was true, the morality around it would have complete different dimensions, and none of what is talked about it would make any sense.
The solution to this problem is either moral flexibility (which I do not suggest) OR a moral stance which is not dependent on the material facts.
We hold these truths to be self-evident: that all men are created equal.
"We don't like research that claims ML seems able to classify people according to identity groups from photos".
Now, that is obviously a controversial area with huge implications ethically and politically and the clear potential for abuse of such tools. A well-argued check on the strength of such claims would be welcome.
But throwing a load of terms more contemporarily associated with culture-war diatribes like "physiognomy", "phrenology" and "nazi" etc. into what should presumably be a calm scientific analysis is not a persuasive call to invest the time in working out what the actual argument is within the paper.
I think the paper would be more effective if considerably less verbose, with a more balanced tone and a clearer exposition of its core argument.
I challenge that the first two terms are characteristically associated with culture-war diatribes — they are associated with the claimed ability to determine personal characteristics from the shape of one's head or face, which is what this paper is about. As to the third one, it is used only three times in its literal historical meaning to cite the Nazi regime as a well known example of misuse of pseudoscience by the state.
Could the paper make its point more clearly? Perhaps, but I don't think mentioning nazis (etc.) is a problem.
These are pretty regular terms describing well-known and discredited fields and ideologies and I don't see them being "thrown around" in the paper. Maybe the paper is bad or unclear in other ways but this is a strange demand for "balance". It is not culture war to say 'phrenologist' or 'physiognomy' when talking about these actual things.
Good writing should communicate a concept clearly. For some reason, science writers feel the need to swallow and partially regurgitate the dictionary. “Look how clever I am: therefore my conclusions are valid”.
I’ve spent ten minutes reading, and I still can’t figure out whether they’re saying that physiognomy should not be explored or does not work. They might have a good point but by christ are they asking the reader to work to dig it out.
It is important to resist that dynamic at every level, so it is probably worth supporting the papers authors in pointing it out. The risk of pseudoscience taking on a racial tinge and leaking out into the real world is always present.
* Inferring sexual orientation: Linking «self-reported sexual orientation labels» with «[...]scraped their data from social media profiles, claiming that training their classifiers on “self-taken, easily accessible digital facial images increases the ecological validity of our results.”[...]». Social media profile photos are by their very nature socially influenced, with open sexual orientation being an important cue to display.
* Personality psychology: Training and test datasets came from the same pool of «participants [who] self-reported personality characteristics by completing an online questionnaire and then uploaded several photographs». This heavily suggests that the participants were aware when choosing the photos that this was a "personality type" experiment, and may even have made their own awareness of their personality more salient by doing the test first and then uploading the photographs.
* “Abnormality” classification: General critique of lack of transparency as to how the true labels were determined.
* Lie detection: The ability to detect the facial differences between people following two different experimental instructions does not equate to lie detection.
* Criminality detection: At least they used official ID photographs instead of self-selection-biased photos like the first example... but consider this: what conclusions would their same model reach if it used official ID photos of US populations? The confounding factors of class and ethnicity are obvious.
"Epistemic" has become a bit of trigger word for me. More often than not its usage contributes nothing but social signalling.