For example, in the Mac Giolla and Kajonius study [1] the data came from an online questionnaire.
I'm finding it very hard to criticise the methodology of this sort of study without [edited out] being extremely rude and dismissive to a whole field of research. But, I really don't understand how all this is supposed to make sense. I look at the plot in the middle of the Scientific American article that is marked "Agreeableness" on the x-axis (with a score from 1 to 5). I look at the papers cited. There's plenty of maths there, but what is anything calculating? What is "agreeableness" and why is it measured on a scale of 1 to 5? It seems to be a term that has a specific meaning in psychology and sociology, and that ultimately translates to "a score of n in this standardised questionnaire".
But, if you can just define whatever quantities you like and give them a commonly used word for a name- then what does anything mean anymore? You can just define anything you like as anything you like and measure it anyway you like- and claim anything you want at all.
At the end of the day, is all this measuring anything other than trends in filling up questionnaires? Can we really draw any other conclusions about the differences of men and women than that the samples of the studies filled in their questionnaires in different ways? Is even that a safe conclusion? If you take statistics on a random process you can always model it- but you'll be modelling noise. How is this possibility excluded here?
All the statistics are quantifying the answers that respondents gave to questionnaires and how the researchers rated them - without any attempt to blind or randomise anything, to protect from participant or researcher bias, as far as I can tell (I'm searching for the word "blind" in the papers and finding nothing). In the study I link below, participants found the website by internet searches and word-of-mouth. The study itself points out that this is a self-selecting sample, but what have they done to exclude the possibility of adversarial participation (people purposefully filling in a form in a certain way to confuse results)?
There is so much that is extremely precarious about the findings of those studies and yet the Scientific American article jumps directly to d values. Yes, but d-values on what? What is being measured? What do the numbers stand for?
This is just heartbreaking to see that such a contentious issue is treated with such frivolity. If it's not possible to lay to rest such hot button issues with solid scientific work- then don't do it. It will just make matters worse.
__________________
[1] https://onlinelibrary.wiley.com/doi/epdf/10.1002/ijop.12529?