HNHacker News
TopNewBestAskShowJobs

selectron

279 karma · joined January 31, 2016

submissionscomments
selectron··on Why Law Firm Rankings Are Useless
It is a legitimate question.
selectron··on Why Law Firm Rankings Are Useless
How much better?
selectron··on Why Law Firm Rankings Are Useless
This is a great point about making sensible assumptions. Too often I see evidence that people think that data analysis should be devoid of assumptions, and any assumptions completely invalidate the analysis. In reality, almost all analysis will have some assumptions, and the good questions to ask is how plausible are the assumptions and what is the plausible impact of what happens if the assumptions are wrong.
selectron··on Why Law Firm Rankings Are Useless
Why should we expect there isn't a correlation? They do mention it in their analysis. At the end of the day, just because there might be a systematic bias in your result doesn't mean there is a systematic bias.

All real-world analysis (especially for observational studies) rests on certain assumptions. It is always true that these assumptions might be wrong, but it is important to think about whether or not the assumption is plausible. It seems plausible that on average, a lawyer who is able to get better outcomes when they don't settle is also able to get better outcomes when they do settle.

Furthermore, even if most cases are settled, the rare cases that do go to trial can have an outsized impact. Usually people settle because a bad judgment is devastating (as well as not wanting to pay legal costs).

selectron··on Why Law Firm Rankings Are Useless
It is always easy to criticize a data-driven analysis by saying its assumptions could be wrong. In the real world, all analysis is based on assumptions, some of which you can always claim might not be correct. But you have to really present an argument as to why and by how much the assumption is likely to be wrong, you can't just state that the assumption might be wrong. The assumption that cases which don't settle are not at all indicative of how well a lawyer performs is a very bold claim, much bolder in my mind (admittedly I am not a lawyer) than the claim that there should be some correlation between lawyer performance and results in cases that did not settle. It is possible that lower ranking lawyers settle less often, I'd like to see the data on that.

Furthermore, on average the effects you are mentioning will wash-out, unless there is a systematic bias whereby lower ranking firms and higher ranking firms settle in different manners.

selectron··on Job Descriptions Should Be Better
The main thing I want from job descriptions is a salary range. The fact that companies don't post salaries is a strong counter-point to how companies complain about how hard it is to hire software engineers.
selectron··on How to Become a Data Scientist, Part 2
The problem is that the term data analyst has come to mean data reporter. Similarly "business analyst" generally involves tasks that are best solved in Excel. The "science" in data science is about testing predictions. But I agree data science is a terrible phrase.
selectron··on Ask HN: Spending my free time in video games. It's eating me. Suggestions?
You don't have to be productive all the time. It is important to have some time to relax and have fun. There are far worse things you could be doing than playing too many video games. You can try replacing video games with a more productive activity, but make sure it is something you enjoy doing and don't feel the need to quit gaming entirely.
selectron··on We Should Not Accept Scientific Results That Have Not Been Repeated
> Taking papers at face value is really only a problem in science reporting and at (very) sub-par institutions/venues. > WRT the former, science reporters often grossly misunderstand the paper anyways. All the good reproducible science in the world is of zero help if science reporters are going to bastardize the results beyond recognition anyways...

Science is funded by the public, and done for the public. Good science reporting is very important to ensure that science continues to get funded. Too often scientific papers are written in a way that makes them incomprehensible to anyone outside of the field, whether that is through pressure to use the least amount of words possible or use of technical jargon.

selectron··on Gradient Boosting Explained in 3D
The explanation glosses over a few important details. Gradient boosting works by adding some small weight to the instances the model is incorrectly predicting. The amount of extra weight these instances get is a parameter that is tuned with validation - because this parameter can be 0, if you are doing correct cv gradient boosting trees is usually superior to random forests. You also do need to tune the number of trees you use in gradient boosting or else you will overfit.

Gradient boosting doesn't get nearly enough hype as compared to things like neural nets. The significant majority of winning solutions to Kaggle competitions for a non-image or text-processing dataset will use xgboost to do gradient boosting as part of the ensemble model. Furthermore, it is a really easy method to understand and use while still being state-of-the-art.

selectron··on Data Science Competitions 101: Anatomy and Approach
Feature engineering and model ensembling are usually what separates the top competitors.
selectron··on A Report on the Flawed 2016 Democratic Primaries
Hand counting of votes seems like a no-brainer, regardless of whether there was a conspiracy this election.
selectron··on Introduction to Zipline: A Trading Library for Python
This statement is too general. You could of said the same thing about chess, there are chess Grandmasters who devote their lives to studying the game yet computers play chess at a much higher level than any human.
selectron··on Approaching Almost Any Machine Learning Problem
I would say that table is really quite valuable. Kaggle problems come from all types of companies, so it doesn't make sense to say that it is "overfitted patterns that he's adopted in his own realm". With that said, validation on your own dataset will trump general knowledge, so you shouldn't view these parameters as hard and fast rules. But the parameters in that table will provide a useful starting point, and if you stray too far from them that is a warning sign that you might be overfitting.
selectron··on Approaching Almost Any Machine Learning Problem
For image competitions you are right. Neural networks are often in winning teams ensembles, but they require a lot more work than something like xgboost (gradient-boosted decision trees). For a dataset that isn't image processing or NLP, xgboost is in general much more widely used than neural nets. Neural nets suffer from the amount of computing resources and knowledge needed to apply them, though given infinite knowledge and computing power they are probably on par with or better than xgboost. And if you need to analyze an image they are great.
selectron··on Approaching Almost Any Machine Learning Problem
1) It depends heavily on the model. Something like xgboost (gradient boosted decision tree) will handle irrelevant features fairly well, while other models (like linear models, especially without lasso regularization) will have much more trouble. In virtually all cases adding noise will decrease model performance.

2) Same as 1), depends on the model. With good hyper-parameters xgboost can handle correlated features well, while other models may struggle.

3) With a good model (again like xgboost), feature engineering is usually the best use of your time. Removing "bad" labels and "noise" in the data is especially dangerous, as if you are not extremely careful you can make your model worse. If you can identify why the label is "bad" then you can remove or correct it, but you need a reason why you wouldn't have these bad labels on your test dataset. Removing outliers can help your model, but it is risky. In contrast smart feature engineering is low risk and can provide large gains if you see a pattern the model could not see. Feature selection can be important as well, and is generally pretty quick assuming you have good hardware, so you might as well do it, especially if you have some knowledge about which features you expect to be not that useful.

selectron··on Ask HN: Machine learning, AI – struggling learners
There is no way machine learning will be a necessary skill for software engineering, if that is your motivation I would not spend time learning it. However, if you still want to learn it you should first study statistics, for instance http://www-bcf.usc.edu/~gareth/ISL/.
selectron··on Ask HN: R or Python? I am a newbie
My advice is Python, but it depends on what your background is and what you want to do. If this is your first language and you have a stats background, R is a solid choice. If you already know another language, R has a lot of flaws that are quite frustrating. Perhaps the worst thing about R is how hard it is to google answers to as opposed to Python.

Like if you google R for loop, the first result http://www.r-bloggers.com/how-to-write-the-first-for-loop-in... is much worse than the equivalent first result for python: https://wiki.python.org/moin/ForLoop

selectron··on The Mythos of Model Interpretability
Interesting. After watching the show Billions, and reading up on how much money hedge fund managers make on fees (seems totally ridiculous), I wonder how common is illegal insider trading for hedge funds? No matter how good your model is, you won't beat someone with information your model doesn't have.
selectron··on Tech Companies and Diversity Hiring
To really understand if companies are biased or not, you also need to know the percent of applicants to these companies who are black. If only 2% of applicants to Google are black, I would expect only 2% of new hires at Google to be black.

The assumption that a white applicant and a black applicant should be roughly equal is a strong prior. I would need to see convincing data to counteract this assumption.

selectron··on So Many Research Scientists, So Few Openings as Professors
I agree completely. I also think there is way more luck involved than people want to admit - a lot interesting results are unexpected, and there are so few jobs the timing has to work out for you. I have known plenty of great post docs who couldn't get a Professor job, and plenty of Professors who seem pretty mediocre.
selectron··on So Many Research Scientists, So Few Openings as Professors
Sorry I wasn't clear - the attitude of going into industry being seen as a failure is common, especially among older professors. This attitude is changing somewhat, but is still definitely there.
selectron··on So Many Research Scientists, So Few Openings as Professors
If you just want ROI, you are better off spending your money elsewhere. This is evidenced by the lack of money most companies put into scientific research. Further the gains of science are in general hard to profit off of. (http://www.therichest.com/business/they-could-have-been-bill...)

Another major reason governments fund research is to produce people with PhDs. This is why the apprenticeship structure of academia is so sticky, and why in order to get tenure professors have to graduate students.

selectron··on So Many Research Scientists, So Few Openings as Professors
The problem is that there are plenty of graduate students willing to work for peanuts.
selectron··on So Many Research Scientists, So Few Openings as Professors
The goal of research (at least for basic science) isn't to make money, it is to increase knowledge about the universe. This knowledge is a public good, so it makes sense that private industries motivated by profit do not support fundamental research. Scientific progress is a rising tide that lifts all boats.
selectron··on So Many Research Scientists, So Few Openings as Professors
I can confirm that this idea is pervasive in physics (both experimental and theoretical)
selectron··on On Being a Black Man
That has to be one of the worst abstracts I've ever read. I have no idea whether or not black people tip less having read the abstract.
selectron··on On Being a Black Man
The some stereotypes are based in reality comment was meant to indicate that the fact is there are real differences between the average behaviors of different groups of people. So when shown clear evidence that groups behave differently, it isn't racist to point out that groups behave differently. It would be great if we could have a discussion about real differences in how groups of people act without being called racist.

I could definitely believe that black people are less likely to tip (or avoid paying altogether), as black people are poorer (on average). Black people are also much more likely to commit violent crime (on average). Note that cab drivers are likely to not be white (on average), at least in my experience. I did some research and it appears to be common for cab drivers to avoid picking up black passengers - this is pretty strong evidence that cab drivers have had negative experiences with black passengers.

It is racist for cab drivers to avoid picking up black passengers, but we shouldn't pretend that they don't have a reason for it.

selectron··on On Being a Black Man
> Why would there be a reason black people would give cab drivers a harder time than other skin colors? There isn't any good reason.

This is clearly happening, based on the anecdote of cab drivers avoiding picking up black people as well as telling the OP he is the first black person who they had a positive experience with. Some stereotypes are based in reality, and it isn't racist to acknowledge this.

selectron··on On Being a Black Man
> Being told by cab drivers that they’re the first Black person they’ve ever had a positive encounter with

This was the most surprising to me in his list of things he has grown accustomed to.

Page 1 of 3Next →