What I found was "How to throw away data that doesn't support your desired conclusions," for the most part. "Actuarial Science," a different field, had some useful techniques but not many. They're most interested in ensuring the bad data doesn't get into the tables in the first place; but at least they are doing "data on data" comparisons and not "data to expectations"
We're building "AI" right now but think about the inputs those see: The very first step is to throw away the statistically too common "stop words" ...
What exactly are you referring to here? This seems like a wildly misguided characterization of statistics, which I am sure cannot be based in expertise or practical applied experience.
> We're building "AI" right now but think about the inputs those see: The very first step is to throw away the statistically too common "stop words"
This is a fundamental misunderstanding of what a "stopword" is and how it's used.
Words like "the" are hard to utilize within with a bag-of-words model specifically. Removing them is not something people do/did because they are clueless monkeys. The goal is to improve the signal-to-noise ratio.
For example, traditionally spam filtering uses a very crude variety of bag-of-words model called "Naive Bayes", in which we assume (wrongly of course) that word choice is completely random, and that the only difference between spam and not spam is that random distribution of words. Are you really going to argue that the word "the" is critical to that process? If you can build a better NB spam filter by including stop words, by all means go ahead and do it. But both linguistics and decades of success in the field are against you.
On the other hand, words with grammatical function like "the" are absolutely important and relevant to the overall structure and meaning of a document. Therefore, training pipelines for modern deep-learning-based LLMs like GPT don't remove stop words (as far as I know at least), because the whole idea of a stopword doesn't make sense in a model like that.
I want to be respectful here, but it sounds like you took a cursory look through three vast literatures, without the perspective of having actually used any of this stuff in real life, and drew some invalid conclusions.
Thanks!
Many people in these fields agree my conclusions are invalid. I say the same about theirs.
I'm the guy who builds the experiments on a team of user researchers. There are all sorts of things that seem intuitive to an outsider but are poo-pooed by practitioners as unethical. For instance, you might run a study that doesn't have enough participants to have a statistically significant conclusion. An outsider would deploy it to more participants to see if the trend becomes significant with more data. A trained researcher will cringe at that proposal.
So far as I can tell, researchers consider the experiment final as soon as you peek at the data. If you want any changes - more data, different demographics, etc - you have to throw out everything and start over. Even though it's logically interchangeable, the data you've already collected is considered spoiled, because they don't want allegations of tampering/data grooming.