Not sure how you can trust anything else in the post with that obvious of a strawman argument. I'm not going to spend the time trying.
Not sure how you can trust anything else in the post with that obvious of a strawman argument. I'm not going to spend the time trying.
It's more common in the young, but also has a much better prognosis.
In particular I've seen my share of code written by non-developers, and they all make the same classes of mistakes one makes when you have less than a couple thousand hours of focused development work under your belt. Well we can't do that because X, or we don't need that because Y. Oh, okay. Is that why I have to come tow you out of the proverbial mud every couple of months?
Could you humor me and try a few of these things out? I think it'll help.
What people fear most is automated systems that are constructed on wishful thinking, not healthy skepticism. A lot of our processes are about proving to ourselves what we think we already know. There is plenty of code under the AI umbrella for which very conventional unit tests could be written. If you can't use our practices for ML, you'd better drop what you're doing and start developing parity, instead of bullshitting yourselves and everyone else about your exceptionalism.
Code that makes decisions for or about people needs more ceremony and sobriety around it, not less. If you can't do at least as well as a startup (which is a pretty low bar), then we are really in some deep doodoo. But not to worry. The good news is that when the machines take over and we have to hide underground, you'll be the first ones we eat, so you won't suffer long.
If you're iterating on your approach to a model, it really is a drag to do a bunch of bullshit in order to run an experiment that's likely to fail and just be thrown out anyways.
If you're shipping that model to production and affecting people's lives, it should really be held to a different standard than the "fuck around and see what happens" experiments.
And this is before we even get into unexplainability of deep neural nets.
Stop behaving like you can fake it and then make it later. The code has already shipped and they’re pressuring you to move on to the next problem. Anything you haven’t nailed down by the time the customer is happy is not going to get nailed down. The time you get is the time it takes. So if you genuinely haven’t gotten enough time, then you need to slow down and take the time now, because there is no after.
There’s only been two times my code could kill someone, and you bet your ass I didn’t deliver it a moment faster than I thought reasonable, and in one of those cases it ended up being fast enough (the other was, in retrospect, already doomed before I got there)
Professional software development has plenty of proofs of concept and "iterating on an approach to a" something. I don't understand what confused state of mind leads anyone to assume that any field of engineering, specially software development, doesn't deal with one off quick and dirty tests.
What really boggles the mind is that any data analyst worth his salt knows very well that he needs to track his inputs to be able to analyse the outputs, thus software development practices are far more forgiving than data analysis ones.
It seems to me that the author just wants to make excuses for shoddy work.
I don't necessarily agree, I think it's more to do with market value of output, but a large number of people feel that way. Where we do always agree is that there are Things Which Must Be Done, even when you don't wanna.
The main problem is that "data scientists" don't follow procedures at all in their work. If they applied conventional software engineering so far as it applies they would do a lot better.(e.g. use version controls but don't check the database password in, think carefully what configuration goes with the hardware, the "model" in general, a particular instance of the "model", etc.)
ML projects get you in trouble in other areas.
For one thing, an ML project might involve a step that takes 24 hours or more to train a model. This means you have to schedule in terms of calendar time (e.g. can't start task A until task B is complete) as opposed to punch clock time, story points, or whatever passes for reality in your flavor of agile.
Related to that is automated testing. Normal software "works or doesn't work" but ML software is expected to get it wrong a certain fraction of the time so the question of "Is version x.y.z good to ship?" is fundamentally different in nature with ordinary software.
It irks me particularly that he ends his circle with "harvest insights" when any real game theorist knows it is about taking actions based on information -- e.g. why winners at gambling and trading use the Kelly Criterion.
Seems to me the author decided to write about topics he knows nothing about.
I do tend to agree that it’s made some dramatic improvements regarding utility over the last few years, but...
GP is otherwise right. First figure in the article is a joke. And that makes the article a joke.
Plenty of waterfalls out there. Even more non-waterfalls, aware of iterative, modern process, who will kick you out of the discussion if you try to sell them waterfall.
ML processes need innovation over those of traditional s/w dev. This article is the worst place on the internet to look for that innovation.
with Continuous Improvement you don't want the minimum viable product, you want the Maximum Viable Product.
Seems like a machine learning outfit would really benefit from a leader who is smarter to begin with, a better learner, and already more skilled in the task that you want the machine to do.
The stuff Frank Lloyd Wright didn't know about construction could fill an entire tour. Which I discovered by taking that tour. Boy, did that dude try to ignore physics.
I despise python, but MATLAB is positively satanic by comparison (from a software quality perspective at least)