279 karma · joined January 31, 2016
All real-world analysis (especially for observational studies) rests on certain assumptions. It is always true that these assumptions might be wrong, but it is important to think about whether or not the assumption is plausible. It seems plausible that on average, a lawyer who is able to get better outcomes when they don't settle is also able to get better outcomes when they do settle.
Furthermore, even if most cases are settled, the rare cases that do go to trial can have an outsized impact. Usually people settle because a bad judgment is devastating (as well as not wanting to pay legal costs).
Furthermore, on average the effects you are mentioning will wash-out, unless there is a systematic bias whereby lower ranking firms and higher ranking firms settle in different manners.
Science is funded by the public, and done for the public. Good science reporting is very important to ensure that science continues to get funded. Too often scientific papers are written in a way that makes them incomprehensible to anyone outside of the field, whether that is through pressure to use the least amount of words possible or use of technical jargon.
Gradient boosting doesn't get nearly enough hype as compared to things like neural nets. The significant majority of winning solutions to Kaggle competitions for a non-image or text-processing dataset will use xgboost to do gradient boosting as part of the ensemble model. Furthermore, it is a really easy method to understand and use while still being state-of-the-art.
2) Same as 1), depends on the model. With good hyper-parameters xgboost can handle correlated features well, while other models may struggle.
3) With a good model (again like xgboost), feature engineering is usually the best use of your time. Removing "bad" labels and "noise" in the data is especially dangerous, as if you are not extremely careful you can make your model worse. If you can identify why the label is "bad" then you can remove or correct it, but you need a reason why you wouldn't have these bad labels on your test dataset. Removing outliers can help your model, but it is risky. In contrast smart feature engineering is low risk and can provide large gains if you see a pattern the model could not see. Feature selection can be important as well, and is generally pretty quick assuming you have good hardware, so you might as well do it, especially if you have some knowledge about which features you expect to be not that useful.
Like if you google R for loop, the first result http://www.r-bloggers.com/how-to-write-the-first-for-loop-in... is much worse than the equivalent first result for python: https://wiki.python.org/moin/ForLoop
The assumption that a white applicant and a black applicant should be roughly equal is a strong prior. I would need to see convincing data to counteract this assumption.
Another major reason governments fund research is to produce people with PhDs. This is why the apprenticeship structure of academia is so sticky, and why in order to get tenure professors have to graduate students.
I could definitely believe that black people are less likely to tip (or avoid paying altogether), as black people are poorer (on average). Black people are also much more likely to commit violent crime (on average). Note that cab drivers are likely to not be white (on average), at least in my experience. I did some research and it appears to be common for cab drivers to avoid picking up black passengers - this is pretty strong evidence that cab drivers have had negative experiences with black passengers.
It is racist for cab drivers to avoid picking up black passengers, but we shouldn't pretend that they don't have a reason for it.
This is clearly happening, based on the anecdote of cab drivers avoiding picking up black people as well as telling the OP he is the first black person who they had a positive experience with. Some stereotypes are based in reality, and it isn't racist to acknowledge this.
This was the most surprising to me in his list of things he has grown accustomed to.