Machine learning is still too hard for software engineers
nyckel.com
nyckel.com
And then from an actual software engineering perspective, so much ML code is just run-once jupyter notebook stuff... there is a lot we can do. I need to give this more thought, but I think it's acknowledged there is a big opportunity here
If you can’t implement any existing algorithm from a scientific paper it is pretty hard to understand the flexibility that is required designing a training pipeline.
In terms of tech stacks, cloud providers’ sample architectures and code samples are a base for cloud components and API understanding and not actual implementations.
Now, how you enforce this, I don’t know, but this is my pie in the sky dream.
The domains where the quality of code has very high stakes do have very strict processes and have an extremely high bar.
The domains where the stakes are not high, for ex - implementing an API to download user reports, have a very low bar.
And your point about cancelling the license to code if you implement a bad API, that's like cancelling a journalist because they made a typo in their article and it went out.
I would argue that review boards, licensing, etc are only relevant because of the stakes. How to enforce it becomes irrelevant if you can't sell anybody on the idea that it's important in the first place.
Even the strength of your statement "...I don't think it would hurt the industry to have..." doesn't sound like something that any part of government would be willing to get behind.
There's just such a huge leap from "decisions civil engineers make can kill. people" to "decisions software engineers make can annoy people".
Still, you're not wrong that the field might immediately get better in some ways, but I believe we could also go down a lot of paths that would lead to the widespread concept of "stifling innovation so that people are less likely to be annoyed by buggy or hard to use software". I'm sure there are also much stronger arguments in favor of your idea of a better world when you narrow the scope down to specific types of software or use cases.
I'm a big fan of starting simple, which means linear/log regression (normally with lasso, as it does variable selection).
Then, if you can prove business value, it may make sense to start trying to use more unstructured data.
I mean, you can call it optimization with a specially chosen sparse function if you like.
Could the Data/ML team have done something cool and likely complex to show clearer patterns and maybe get something in place to continually refine that data? Probably. Would it have saved more money for the company eventually? Unknown. Would it have cost substantially more money to design and implement? Pretty sure.
As I keep saying to people, sometimes the best model is a well chosen graph.
Author of the blog post here. It was definitely written from my narrow viewpoint and experience. Our goal is to make more solutions accessible to software developers and instinct was the same as yours - a lot of ML can be within the realm of engineering (even small / one-person teams) and that there are accidental complexities standing in the way of wider use. Our solution (AutoML+SaaS) def doesn't work for every situation. I'm curious to hear more of your thoughts on how ML can be made more accessible to Eng (and vice versa).
"Good pipeline and bad model is much better than bad pipeline and good model" (c) someone
I personally mostly talked to friends, Kaggle, and 1-2 Coursera classes.
PS: Kaggle is pretty far from real production ML, but good enough to dive in.
The math you learn when studying electrical engineering plus all the signal processing stuff, gives you a good foundation to pivot to ML.
It's a shame electrical engineering is so poorly paid by default in comparison to SW dev.
I expect the same for ML. The tooling will improve until regular developers can use it without understanding the fundamentals. They will do useful things but some skilled people will be able to do way more advanced stuff.
That's pretty much the path of all technologies. They get simplified to a level where they useful for a lot of people but some experts will be able to get way more out it. And usually the experts look down on the amateurs.
Most of these people claiming expertise do not have a deep grasp of the mathematical fundamentals required for state of the art research in the field. Can you develop a neural network with features that are invariant to permutational and rotational symmetries? If so, how do you efficiently generate the irreproducible representations of the product of the symmetric and special orthogonal groups for use in a fast Fourier transform? What is minimum description length and why is it so fundamental? How do you solve trust-region problems on Riemannian manifolds? Throwing an off-the-shelf PyTorch library at a problem does not make one an expert in machine learning.
Most neural networks that perform well nowadays are highly specialized to a particular problem domain, and I gave an example of an approach that might be used by someone designing a neural network that is invariant to certain types of symmetries on the input data. This isn’t something a typical DS/ML bootcamper would know how to handle, or even how to approach, despite their claims of expertise.
I wrote this as someone who considers himself a half-decent software engineer trying to use ML for a side project and feeling frustrated by all the effort and "accidental complexity" involved. Why focus on software engineers and ML in this post/rant/company? Because "software is eating the world" and having ML be more accessible to software engineers will broaden the range of problems they can solve.
Thanks for all the comments - I acknowledge all/most of the criticisms as valid. A SaaS/AutoML solution won't work for everyone and definitely not for every problem, and it won't be the only answer to making ML more approachable.
Just because both involve coding doesn't mean software engineers should be expected to have the math chops (stats/prob, linalg, calc, etc.) to make machine learning work for them...
Vice versa is a little more complicated, because ML/DS can be done very inefficiently without the proper coding practices, but understanding the math is independent of that so that point still holds for this comparison.
I don't think so. Or more precisely, they might look different academically but you need to have both as a skill to build something useful.
That's not even to mention the really hard part, selecting a good outcome variable and appropriate ways to measure the performance of your system once it hits prod.
You can definitely split this stuff between people, but it gets super-linearly harder as you add more people, so it's really incredible to find people who can do both (and honestly, there aren't that many of them (I'd like to say us, but I'm probably not there yet)).
It's also unnecessary to do so as long as your institutional processes are capable of synthesizing multiple peoples' competencies across multiple disciplines.
How do you think any machine more complicated than a train car was designed? Do my mechanical engineers need to understand the intricacies of avionics?
> it's really incredible to find people who can do both
Absolutely, and I think you'd have a hard time finding someone who disagrees. But you made a very strong assertion about a "need" which requires much stronger arguments to support.
You clearly work at much better run companies than I do ;)
Then we have a team of "Applied ML practitioners", which I am a part of, we productionize the jupyter notebook, by setting up pipelines, services etc. We understand ML algos, stats, probability etc, but not as much as our data scientist team does.
Having both in the same person would be good, but is not necessary.
After about a week of reading through literature and playing with open ai, it became pretty obvious we were still super far away from being able to build something the business could actually find value in.
My problem scales horribly with the ML training angle we have today, because it's the super complex one-off queries we need the most help with, not the simple ones we can anticipate and train against.
What we need is actual intelligence for many of our problems. Things like subjective criteria are important to us. Realizing maybe a recursive query is a fair compromise to reduce a 400 line monster to 30 lines. Assuming the 400 "looks nasty", that is. I guess you could train that bit too, but then your solution space gets even more impossible to target.
ML is kind of like tennis in the sense that you look at Nadal and Federer and all the greats and you can say, "man, I could do that," but you can't, not even close in this life and the next after this one.
In fact, most SWEs who have developed some sort of "intuition" about ML have been quite dangerous, as they tend to be condescending to the real experts, have built things that make no sense at all or fall apart when 1 data point our of 100,000 changes, and when presented with the fact that they have no clue about ML, resort to the "the AI is all hype anyway" comeback. And vice versa (ML practitioners like me who think they can do SWE with one eye closed, "what's the big deal?") is also true.
If you're looking to land a job at FAIR / Deepmind or Google Brain/ Nvidia Research as a researcher or ML scientist the expectations of knowledge are very different than 'data science'. These are research lab groups, that work on pushing the state of the art forward. They are also supported by great engineers, building awesome tools that improve ML research. So transitioning into this sort of role requires more than doing Kaggle competitions, it requires developing an intuition for the respective ML subfield / and trying new things and usually failing. i.e. this is a research role and will require a lot of study and learning
If on the other hand you are looking for datascience / take model and build pipeline to run AI, or perform hyper param sweeps or simply modify some model code, then on I would say that is much more engineering than research ML. This has a much lower barrier to entry coming from engineering and could be a good stepping stone to a transition into pure ML research.
On a more general note to consider when thinking of transitioning to ML is that these systems are probabilistic in nature vs purely deterministic as they are in more general software systems. People (ie humans) are bad at wrapping their heads around distributional processes - you can see this in all fields that deal with them (Quantum vs Classical Physics, Biological Systems etc).
In general I guess what I have seen is when engineers try to dip their toes into ML, what's required is a mindset shift in how to approach problems. Once that happens the depth of that shift determines the type of role with ML you wish to pursue.
Making a successful transition to ML in my opinion, depends a lot on the individual. Without a strong background in calculus, linear algebra and statistics, it's going to be difficult. Training a model is what people tend to focus on, but in my opinion, that's the easy part. Evaluating/validating a model, analyzing and preparing your data, anticipating model performance, understanding what to do to improve your fit, model selection or architecture. Developing custom deep learning architectures at times requires a bit of an abstract mathematical intuition that I think will suit many engineers very well. A lot of engineers are well equipped to be successful in making a transition, but on the other hand, at least as many aren't.
In the future, I think the field will have many varying degrees of expertise, with the barriers to entry becoming lower all the time. We're reaching a point where some common use cases can be solved adequately in a nearly automated fashion. Some "autoML" tools don't really require any real understanding of ML, though I think it's not wise to get in the habit of using them without understanding how to evaluate a fit. These tools will be great for people who want to occasionally use ML to solve some smaller problems, but as a part of their larger job function.
In some middle area, ML engineers and practitioners will be training and operationalizing models, and keeping up with major developments in research. But there will be some significant changes in the next decade. I predict the nebulous of data science, data analysis and machine learning will become formalized into 3 major skills - exploratory data analysis, machine learning and advanced computational statistics.
At the lowest level, researchers will continue developing the field, which like you say, is probably not something you transition directly into.
Who came up with this silly idea that something that is a valid knowledge domain of its own is suddenly going to become "easy"?
So many organisations can't even get the basics of efficiency and custom service right but they still invest heavily into cutting edge tech. I can imagine plenty of companies joining the ML bandwagon and still not even really knowing what they do as a business.
oh shit, are we all supposed to be now?
I want to be able to compose these tools like I would random unix ones: Something like 'Identify album covers in this image | extract the text in said covers | spotify api'.
It seems like there are so many breakthrough models but both due to technical (size/compute) and industrial ($$$) concerns they remain out of reach for random devs, let alone packageable into a `grep` style composable tool.
Assuming you mean coreutils. binutils is for managing/inspecting binary executables.
But to your point: there were two key innovations and criteria of UNIX pipelines: a common and understandable data format, and writing programs to send and receive anonymous data. Crucially, the input and output formats were the same: plain text, separated by newlines.
In contrast neural networks are applied to a variety of data formats. Images, video, audio, text, social networks etc. each with their own encoding into something an NN can work with, with varying dimensions, features, metadata etc. So it doesn't make sense to bundle them as 'neural utils' but rather utils along whatever pipeline already exists, like GraphicsMagick. Which does leave a huge blind spot for the domain transforms like text recognition.
If you stay within the AI ecosystem, you _can_ set up reusable layers for tensorFlow, but typically you cant swap out something in the middle without retraining all layers below it. Which you might treat as a violation of the anonymity criterion, since the behavior / performance of a layer relies on the specific behavior of those above it.
https://blog.acolyer.org/2020/10/19/the-case-for-a-learned-s...
Testing paradigms are either too high level or too specific. Recent work on evolving behavioral tests addresses this but it requires more manual effort and interpretation which kinda defeats to point of automated tests.
This bears more resemblance to traditional manufacturing actually. I think there may be some value in borrowing ideas from statistical process control, rather than trying to force predictions into deterministic cases.
I used to tell my students "software engineering is 70% reading and 30% coding". This remained consistent as I dove into Data Science and now ML with computer vision at Roboflow.
Of course the first time I was exposed to it during a fellowship, I thought I was out of my depths, but this comes with everything new.
To @deepsun point, I've found Kaggle's intro courses quite excellent as well.
Even with a magic API the latency still isnt good enough, so the choice is often entire ML solutions like the OPs product, months of development time, or a really expensive container.
Shameless plug since my Show HN went completely ignored earlier today ;) ...but that's why i built this: https://news.ycombinator.com/item?id=30428664
There's no getting around the complexities of fit, bias and customized models for many ML problems, so my observation above is obviously limited in its applicability.
I just compiled a few resources I found useful on groking computer vision and deep learning.
Short list I’ve enjoyed recently:
- PyImageSearch blog: https://pyimagesearch.com/blog/
- Fastai library and course: https://www.fast.ai/
- Yan LeCun’s 2021 Spring NYU class: https://m.youtube.com/playlist?list=PLLHTzKZzVU9e6xUfG10TkTW...
As an experienced software developer who used to learn a new framework every week, I thought ML was going to be piece of cake. In reality, I went down this rabbit hole 4 years ago, and I'm still in there, I was so innocent back then
Sorry, it had to be done.
I think that is why people find ML so hard. It isn't a single sausage machine to create insight from, it is a set of entire stacks from philosophy all the way down to the electronics.
The article is talking about applications where the value is less-or-equal to the work of learning and applying some library (it mentions days or weeks). What is the actual value of applying ML in these cases?
Also check out Jeremy Howard from fast.ai - also no PhD but amazing teacher and contributes actively to research.
Chris Olah (Google Brain, OpenAI, etc) go to university at all.
PhD definitely not required.
There are a several videos on youtube where the creator reviews a paper and then implements it from scratch. For example this channel is pretty good: https://www.youtube.com/c/AladdinPersson/playlists
Yes, it is going to be laborious to have state-of-the-art deep learning implemented in your infrastructure.
You're absolutely right, sometimes, simple and predictable solutions are much better than AI magic ^^
In general, I feel like ML platforms have the problem outlined by https://xkcd.com/927/ (Standards)
I can't say what we do different from everyone, but a few things that we focus on: * Speed: we train models based on DL in seconds. So you get real-time feedback on your model/data as you annotate and upload more. This is true for a few, but far from all of our competitors. In our benchmarking we find that we still perform on par with the competition (at least in the "low-data" regime https://www.nyckel.com/blog/automl-benchmark-nyckel-google-h...) * Level of abstraction: Many competitors expose some ML knobs for their users thinking it will improve the experience. We found that this induces "ML anxiety" for many. As a result we have zero knobs. Just focus on your data, we do the rest. * API: we have spend a ton of time developing clean API abstractions. Some competitors have great APIs, other don't. * Cost: we are super cheap. Our lowest tier if $50. We don't charge for training or per function/model.
No, actually, you're just being dishonest. Even if you hide TensorFlow and the keras models behind a nice GUI, people still need that mathematics knowledge to succeed. And yes, pre-training is great. But you need a shitload of stochastic analysis to make sure that the pre-trained embedding won't distort your results.
"For a software engineer, the hardest thing about developing Machine Learning functionality should be finding clean and representative ground-truth data, but it often isn’t."
That is (in my opinion) an entirely bogus request. Machine learning is a mathematical / statistical tool for modeling large unknown functions. I feel like this sentence is akin in usefulness to:
"For a nuclear power plant, the hardest thing about building one should be to draw how the finished building will look like in the press release".
Someone "doing" machine learning without the requisite math knowledge is effectively driving blind. And worse than that, they don't even know what they don't see, because they lack the skills to identify their blind spots. That's how you end up with a "tank detection AI" that in reality just classifies the weather into bright vs. dark. [1]
Companies like this who promise advanced mathematical algorithms with no prior skill or knowledge are how we unleash a plague of buggy unverified automatons upon the world.
[1] https://www.lesswrong.com/posts/5o3CxyvZ2XKawRB5w/machine-le...
Clarifai, Amazon Rekognition, Google AutoML Vision, Nykel (the article here), Amazon Comprehend, Google AutoML Natural Language, MonkeyLearn, Lateral, BigML, Azure ML, Lobe, DataRobot, Rapidminer, Dataiku ...
Did I forget anyone?
EDIT: H2O’s Driverless AI, Floyd, AWS SageMaker, Databricks
EDIT2: Pega Platform, MLFlow, Comet.ml
EDIT3: $SNOW SnowFlake, Spell.ml, Cloudera ML, Alibaba Cloud
Personally, if there's ever a downturn, I plan to play Snowflake sales people off each other and get enough credits to last me a lifetime ;)
The Nuclear Power industry is starting to think about stopping doing all designs on paper, maybe in a few decades they will have achieved this, sending a message that good data is the thing they should work on first isn't a bad idea.
Unless you're focusing on German tanks :)
Who cares if it denies bail to minorities or hits a few pedestrians from time to time?
The problem isn’t that ML is too hard, it’s that it’s too easy. Crazy people keep connecting ML to systems that matter- that have real, irreversible impact to humans- and they don’t understand it.
I wish ML were 1000x harder/more expensive to integrate so the economics would drive away frivolity.
I've seen it time and time again: Team has a black box ML/AI solution to a "problem." Team wants to eke out better P/R or deal with some complex edge cases. But team's problem is fundamentally ill-posed and no amount of hacking or kludges will actually produce the success criterion that they need.
The problem is the accessibility to these tools, which in many times has led folks to neglect the subject matter expertise required to effectively apply them in the first place. At least as these tools catch on in popularity in myriad problem domains, there will be a new generation of subject matter / domain experts who are familiar with them, and we'll probably jump over this hurdle.
Because of that, doing deep learning consists of a bunch of cobbled together heuristics for getting good results and probing the model to give a human an intuition for whether it's learning correctly. The tricks and tips for steering that black box have mostly been developed in the last decade: it is not a super deep well.
These tricks and heuristics are like the knowledge needed to be a technician in a nuclear facility, not the knowledge needed to build the nuclear facility in the first place. It's not nothing, to be sure, but unless you're a researcher developing new novel architectures, a very shallow understanding of the statistics will go a very long way.
I mean, even the example given by the OP about the tanks is super well known (apocryphal[0]) and doesn't require math knowledge to avoid. You just have to have heard of this kind of failure mode
Yes exactly. You have to be aware of it, you have to know what it entails and what can cause it and how to diagnose and fix it.
That’s the other half of the domain knowledge, and just “autoML-ing it” or following some set of prescribed steps won’t necessarily get you that solution.
If I hand you a 175B parameter language model, are you really contending we know what's going on in there? At a mechanical level, sure, it's tensor products and activation functions, but that's like saying we know how human brains work because the standard model is very predictive.
Note I'm not saying we perfectly understand everything, or that our understanding is as solid as that involving convex models or linear regression. But "black box" just isn't true anymore and just adding unnecessary mystique. We are somewhere between Newton and the Lord Kelvin and Faraday era of physics, no longer in the ancient alchemist days
A big part of machine learning is looking at weights and outputs to make sure the results are sane and that you have an understanding of what's going on. This is true no matter what algorithm you use to make predictions.
Software developers HAVE to have an understanding of the subject they're developing for. Computers are not brains, and they are not able to understand the objective or context in which they run.
I could spit out their crappy tagline - "the hardest thing about developing __X__ should be __Y__, but it often isn’t." - for almost any topic.
"The hardest thing about developing an inertial navigation system should be getting clean sensor readings, but it often isn't"
"The hardest thing about developing MITM proxies should be getting certs configured, but it often isn't"
"The hardest thing about developing web extensions should be setting up your manifest file, but it often isn't"
"The hardest thing about web development should be handling https requests, but it often isn't"
"The hardest thing about having a baby should be labor, but it often isn't"
"The hardest thing about making a car should be getting high quality steel, but it often isn't"
ON AND ON.
It seems to me the lack of knowledge was not knowing to use a diverse sample set, not some lack of mathematics knowledge.
Also; most ‘black box’ solutions still are not that friendly.
Which is so much bullshit. The hardest thing is validating your hypotheses, which machine learning turns into a black box. When we have coworkers who insist on operating on wishful thinking we try to maneuver them out of a job. Except every 10-15 years when the built up pressure of fads overwhelms reason and we all get stupid for a generation (which in software is about five years).
The things that started as AI that we don’t call AI anymore, and don’t lump in with AI when discussing successes or failures? It’s because they can be explained in plain English and implemented without much or even any special jargon that marks it as anything more than exceptionally clever Logic.
You don't need to know most of the details how your car works in order to drive. You need much more knowledge to build one, yes, but not to drive.
There are different levels of abstractions and depending on your problem you need to understand them only up to a certain level. And different people have different problems to solve.
In most real-world problems today, the difficult part is indeed the data, not the underlying math of the activation function, loss function, or optimizer. Just Google "data-centric AI Andrew Ng" to read more on the topic from one of the most well-known people in ML.
"Even at 200 million frames, there are 10% of games where all algorithms reach less than 10% of human. This final point in particular shows us that all of our recent advances continue to be severely limited on a small subset of the Atari 2600 games."
In short, current AI approaches cannot even reliably win video games from 40 years ago, no matter how much $$$ you burn on GPU power.
How do you expect a non-expert to know if their problem is in the 10% that works well, the 80% that works tolerably, but worse than traditional algorithms, or the 10% where all bets are off?
I work at Nyckel. In fact, I'm the "ml guy" at Nyckel. I have a PhD in ML and did some research at Berkeley, but I mostly consider myself a ML engineer. My most recent job was in the self-driving car industry, leading a ML team there.
Knowing the math/stats is helpful when navigating the vast set of models to choose from when fitting your data. Although I'd argue that some sort of black-magic "intuition" earned by doing this for a long time is more important in practice...
However, when validating a model, there is really only one way: test it on production data. This is what Nyckel does: upload your production data, do some annotations, and see if it works. Nyckel handles model search, cross validation, etc for you which reduces the risk of bugs. In a way we are making the argument that by focusing on your data, you are most likely to do well.
But what about that pesky out-of-domain issue? Like the tank/cats or whatever? Well, our customers are not trying to develop AGI, but solve narrow problems using image and text classification. And they are also doing it for themselves so they have all the incentives to be honest. Consider one example use-case from a health food store we work with: "what type of legume (from the 10 I offer in bulk) is in this picture"? As long as they train and test on production data from the warehouse camera stream, they are in good shape from a statistical perspective. Sure, if they throw in a picture from anywhere else, they are toast, but why would they?
I believe it is a very common mistake for intelligent people to assume that others will behave at least reasonable. But in my experience, when people do AI without understanding it, all bets are off.
"Sure, if they throw in a picture from anywhere else, they are toast, but why would they?" Since you list a Barcodeless Scanner as an example, the manufacturer of strawberries might run a promotion for blueberries on their box. For a non-expert user, it is unimaginable that a model trained on 3D blueberries might be triggered by a 2D photo of blueberries.
Also, I'm going to go with your legume example. As soon as each new truck arrives, the intern runs out and takes photos of the legumes in their boxes for the AI training. He uploads the images to your website and trains a model. TADA! The model is deployed to production and starts causing issues. But the people working alongside the fancy new celebrated machine don't want to lose their job, so they silently fix what's going wrong. You've just reduced productivity by introducing a costly machine.
Turns out, the different suppliers arrive at different times of day, so the lighting is different. And different suppliers use different box types. But without expert domain knowledge, you wouldn't even consider that this might be a problem. Also, why do you assume the customer will verify their model on independently sampled production data? To someone lacking the domain knowledge, using the exact same set of photos for training and for verification seems just fine. Actually, it's a lot less work that way.
That's what I tried to get at with my blind driver analogy. An untrained person will do things that seem absurdly unreasonable to us. But to them, it's the logical choice. They lack the knowledge to properly understand why what they are doing might be problematic.
Based on your description, however, it sounds like you (and your team of experts) are actively working with this customer and giving them feedback on what to do and how to do it. Have you considered making that part of your offering?
"Use Nyckel to integrate state of the art machine learning into your application. Anyone can curate their data set with our ML platform. A quick chat with an experienced AI engineer helps identify the best model and training procedure for your use case. It only takes minutes to finish your first model. Once created, your functions can be invoked in real-time using our API."
I'm pretty sure any serious business user would be happy to spend $100 for a 15 minute chat with someone that checks that their data is OK and their approach is reasonable. And it's also a nice way to segment out those that'll never become paid users anyway.