Machine Learning for Developers
xyclade.github.io
xyclade.github.io
Lots of software doesn't work. Is there a substantial difference between putting an overfitting model in production, and putting a poorly tested program in production?
That depends "how" it breaks. As a novice coder myself, I've had things go wrong that I don't notice or can't identify, and it looks like my program is running fine.
I think that's the parent's point: it might be stupid to implement crappy macho learning models into production, but it isn't worrisome. It's expected.
Your comment is a bit obtuse, developers are not creating parsers and compilers. Elevators and calculators are very robust technologies that already work.
What I worry about is a new wave of engineers and developers thinking they understand statistical models and then proceeding to work at the big banks and have their models blow things up. If PhDs can make such disastrous non-robust models, how on earth is a random developer who took a summer course not going to do the same?
Now if the banks actually failed on their own, then by natural selection the less skilled would be out of jobs and people would stop trying to "short-cut" gaining this type of knowledge. But that's not what happens. Academics keep writing papers and hyping up specific techniques for which they can give conferences on, and the taxpayer bails out the idiots at the top.
I worked in the finance sector as a (lead) developer on a production use trading system for quite some years. None of the developers had formal math or statistics knowledge to the extent required to develop this system.
This didn't really bother anyone in the least despite the fact a mistake could cause a loss of millions of dollars in practically no time at all...
The reason for this wasn't because anyone was ignorant enough to think the programmers knew what was going on. It was because it was a finance company that also employed plenty of mathematicians, statisticians, physicists, and other's with the proper math/stats background. The programmers wrote the code, but the math/stats people wrote the business rules and the formulas and extensively tested that the system worked correctly against a large enough variety of models with expected outcomes that they were able to have sufficient confidence that the system reward/risk measurements were appropriate.
So my answer would be no; I don't find this all that troublesome. We can't be experts at everything and smart companies realize this so they should be creating teams with the correct skillset to be successful.
Realistically, the outcome will be the same as it is now: those firms whose models don't reflect reality will blow up, those whose do will get bigger, a few will get too big to fail off some very confidently-expressed models and make a lot of people mad at them, and eventually the market will straighten out who's lying and who's not. Won't be painless, but then, capitalism never is.
I'm picturing one of those dystopic films/novels where the main character is deleted/fired/jailed as a result of an algorithm error. Yes, in real life the trends will overcome the bad models. But just think of the potential consequences for harm on an individual basis!
It sucks, it's not fair or just, and everything would run more smoothly if we were omniscient beings living in a completely egalitarian society. Unfortunately, that's not the reality we live in. In the meantime, we accept it as simply fate or vulnerability, and muddle through as best as we can.
Think of the possibilities of machine learning for detecting precrime!
1. There isn't much correlation between quantity or even quality of papers you publish and the quality of your code. Meaning, writing cleaner code is not going to help you get that postdoc or faculty position.
2. Doing research is full of stops and starts and branches that fail and approaches that get thrown out. It's a waste of time to write clean code since you know it'll most likely be thrown out. When you do get an approach that works, you publish your paper and move on.
What is 'good'? In 'software development' 'good' is usually connected to writing clear, maintainable, test covered code. In most scientific research it means something completely different. I think on HN most adhere to the former definition of good and in that sense most researchers (especially in physics, but also CS / ML) are not 'good' according to that definition (because you need quite a lot of years of experience in a corporate setting usually) and actually even bad. But the code works and implements the concepts in their papers so they are 'good' in that respect. That is more rapid prototyping to make a POC to show it works, after which you properly rewrite it.
As a Junior, we were building toy programs that do Operations research type of work - solving linear equations via various matrix operations, design optimal queue processes based on poisson process.
Assuming a software engineer is a CS undergrad, he/she most likely has good footing to learn more by themselves.
E.g - no probability class that doesn't require analysis 1 and 2 is truly a probability class.
In my understanding "machine learning" is just a buzz word for the good old fashioned data-mining, which is still a part of applied mathematics/statistics. Only because it involves computers it doesn't belongs to CS.
So what you have written sounds for me like "doing applied statistics without much understanding of mathematics and probability". And yes, I am worried about it.
Basically, if this is going to become a thing, then there is no stopping it.
Right now it's the glory days of ML when nobody much has the ability to judge success. Unlike software engineering broadly, where these glory days just keep going, ML is all about measuring success. People will detect failures.
The real risk is when people systematically underestimate the risk like the copula thing occurring with the subprime market. That was anything bug untrained people using models—they would not have been as dangerous as they were if they weren't so damn good to begin with. This is a robustness failure, not a poorly trained workforce failure.
You're kidding, right? Java has been extremely popular for ML for a long time. Not to cast any shade on Python, but I'd say Java and Python are roughly equivalent in this regard. Both have good libraries for various ML tasks and both are very popular in the domain.
For reference, a quick search on mloss.org finds 84 projects identified as "Java" and 105 identified as "Python". So while Python has a small edge in sheer numbers I think that supports my assertion that they roughly equivalent in terms of their popularity for ML.
And at least one seemingly fairly current NN library:
https://github.com/ivan-vasilev/neuralnetworks
An an older "pre deep learning" NN library called Neuroph.
http://neuroph.sourceforge.net/
and another older one called JOONE:
http://sourceforge.net/projects/joone/files/joone-engine/
So in general, the answer is "yes" as to whether or not people are doing Neural Network / DL work in Java. I can't tell you how much such work is happening, or really compare Java/Scala to Python, etc., at that level of granularity though.
And just for a little bit more perspective: IBM Watson is (or was) apparently largely Java based:
http://www.drdobbs.com/jvm/ibms-watson-written-mostly-in-jav...
So, respectfully, lots of people use languages other than Python for ML, and I doubt if Python is even the largest deploy base of ML.
Most ML is an iterative process, and the final model that's used in production is just the tip of the iceberg of the development work that went on. Python works as well for exploratory programming there as it does for any other domain.
There's plenty of heavy duty machine learning libraries and implementations for big data platforms. Just spark alone has a fairly high quality one:
Deleted comment
What do you use to do a search that specific?