Machine Learning Is Still Too Hard for Software Engineers
nyckel.com
nyckel.com
For now, ML research and development is too complicated and frustrating for me to dedicate the time and energy to become skilled in it.
This is 96% of how ML is used in practice by companies. Most parameter optimization should be automated by whatever library you’re using, beyond basic sanity checking. The challenging parts are creating high-quality training data and deploying the models efficiently at scale.
I’d be interested to know what the next thing to read or do is if you comfortable with entry level ML.
I think the real blocker is time; modern software devs are expected to be across the whole stack. You can't have one person write your backend, frontend, infra, db admin, build ml architecture, train model, etc. It's just too much.
It's just specialising really, wide and shallow vs narrow but deep domains. Forefront of ML stuff requires researchers specialised in that domain, same as any I guess.
I think making lower level tools (fundamental building blocks) will make things more accessible to software engineers, as opposed to the high level wrappers being written. We can grab a library, read the docs, and put pieces together in an efficient way. We just need some core work horse libraries. Like if llama.cpp was a library at the same maturity as sqlite.
One of the problems I see here is that it seems to me that the math education of many CS grads is woefully lacking. Indeed, deep learning math is basically senior-high-school level calculus. Backpropagation is a straightforwards application of the chain rule. There is really no surprising thing or deep insight. The math for deep learning has been set in place for more than fifty years at this point. The only change was the advent of large enough data and computer chips to process it.
That being said, I don't think understanding how the `transformers` library works or composes is the same as understanding how 'deep learning' works. That's like saying knowing how to use GCC gives you a solid understanding of compilers. That's not a ding against those whose main experience with DL is using those libraries. I use many libraries doing things I don't fully understand. I'm not a graphics expert, but I use graphics programs every day, and graphics libraries regularly. That's fine. But using those libraries doesn't make you an expert in computer graphics.
Actually I find this claim problematic and incorrect. Yes, many parts of ML only require multivariate calculus[0] and those are the parts that most people are exposed to. BUT that doesn't mean there isn't a lot of math hiding around in the background. Even understanding something like activation functions take much higher level math as we got to get into topics such as topology, metric theory, and high dimensional statistics. You don't need this understanding to build models (as you also conclude) and most researchers don't understand much of this either tbh. But I think we need to be clear about these distinctions as model evaluation gets insanely complex. Complexity in evaluation only increases as our performance increases and a major problem we face today is we are marginalizing out important information. I'm sure that anyone that hacks around with LLMs or diffusion will understand how leaderboards are often noisy and how picking top models does not guarantee top performance (and why datasets get continually added). It's because evaluation requires more than boiling down performance to a single number. Here, math skills become extremely important.
[0] btw, many will not find this even offered at their high school. In fact, this is often upper level math for many STEM undergraduates and even an elective in my Uni's undergrad CS program. So let's not be so belittling. Shame the education systems instead of people lacking opportunities.
DL requires a certain degree of mathematical maturity to grasp.
This is my realization too. I think ML for SWE courses should focus on "translation" first. Like by "kernel" they mean this specific thing, not the normal meaning of kernel. This is similar to other fields like finance (which I'm working on). After you learn the language, it's actually not too terrible to understand.
EDIT: Hacker news won't let me respond, but the answers below all seem to be because the original meaning has been lost on everyone.
In English, the word 'kernel' means 'core'. An OS kernel is the core of an operating system. In linear algebra, the kernel of a matrix (or a linear transformation, same thing) is the set of vectors it maps to zero, which is also in a sense the 'core' of the mapping (in so far that zero can be seen to be at the 'core' of the vector space / number line).
So actually, the definition is the same, it's just that the word kernel is rather rare, despite having a well understood meaning. Nevertheless, it is the kernel of many common English idioms such as a 'kernel of truth'.
Going to ML kernels... the etymology is a bit convoluted. I believe they come from operator theory and support vector machines. But nevertheless, the 'normal' meaning of kernel works out because kernels are typically the core fundamental operations supported by a machine learning framework atop which the other operations are built. In that sense, regardless of the etymology, the name actually fits.
Whereas in CUDA programming, a kernel is just the code running on the device. When I first heard kernel in relation to CUDA programming, I expected it to be a) the actual OS kernel, and then b) something akin to a GPU driver.
I can see how NVIDIA got there, but it's not immediately obvious if you're coming from a SWE context.
I think this is still not the machine learning 'kernel' being talked about. While you'd inevitably need to know about this for tooling in the space, looks like kernel is also a mathematical term.
But it was hilarious to see how many other "normal" meanings came up in the comments.
Hell, in math, normal even has multiple meanings. You have the normal distribution and surface normals for example
https://en.wikipedia.org/wiki/Kernel_(linear_algebra)
Of course I'm kidding. It's one of those terms that many fields adopt and give it completely different meaning. There's no "normal" meaning of `kernel`.
Do you mean like PyTorch, TensorFlow, or the others?
But these do exist: plain Jax or Pytorch only give you basic linear algebra, differentiation and some basic layers. And there's a plethora of more or less advanced libraries that add specific functionality, for example torch geometric for graph data and lightning to reduce boilerplate.
Chiming in to say that this precisely our observation. The existing ML/DL libraries are not bad as far as those types of things go. In fact, Pytorch is an amazing library IMO. Especially compared to TensorFlow, Caffe and the stuff that came before that.
But like George points out in the article, unlike "traditional" software, ML requires iteration, data management, monitoring, specific infra reqs, and so on. So our take was that libraries would never be enough, hence the SaaS offering.
It wasn’t until I was “forced” to learn it to solve a problem I was facing, that I realized ML is just like any other engineering topic - whether it’s devops or data engineering. You just need motivation, some patience and ideally a project/problem that you can solve while learning all this stuff.
the math is how it's done, it's statistics all the way down. If you don't have a good grasp on statistics you're not going to fully grok ML.
I don't know if this helps more or make it worse. But I have both and getting the same feelings all the time. But basically you just need good statistics and linear algebra knowledge, and you will be fine (on the math side).
> there’s already so many people much more smarter and advanced for me. Why even bother?
That's the very definition of imposter syndrome put in a very good way. But you can ask the same for everything, not only ML. kudos on getting through it, though.
The core algorithms all build on top of each other. The `algebra` part of linear algebra refers to a `field`, but it might as well also be called arithmetic of tensors.
You're actually describing linear algebra. A core topic is system of equations. You might see a 2-Tensor (matrix) like Ax := [[a,b],[c,d]].[x,y] and you could write it as f = ax + by; g = cx + dy. Often dimensions are implicit so it may not look like this, but it is. But that's a big part of what it is about (there's a whole lot more btw). You're absolutely using linear algebra frequently in graphics. Euler angles are a good example, you're just probably not writing them in matrix/tensor form. You will even get a tiny bit of exposure to {field,group} theory/abstract algebra via quaternions.
In ML I'd say it is very similar. The typical researcher is going to have about the same math skills as the typical person studying graphics (I actually started my PhD in HPC graphics). But, and this holds for both domains, having a deeper math understanding only helps. It makes things easier to debug, gives you a better understanding of what the systems are doing, and gives you a lot of tools to solve many problems. I wouldn't ever use math as a strong barrier to entry, but I feel many get complacent with their skill level and we have been discouraging this myth that math doesn't help. Without a doubt, it does.
FWIW, there are a lot of works that do get in deep to the mathematics of ML and I find these absolutely helpful. Anyone that says theory doesn't help practice hasn't read theory or is operating in bad faith. ML uses A LOT of math, but you just don't need it to create good and/or working models. I think the distinction is important.
ABD (all but PhD dissertation) here with strong math skills. I get the imposter syndrome, but let me absolutely assure you that the community at large does not have strong math skills. I routinely talk to people doing diffusion research that don't know what covariance is or pdf. People from top ranked schools, with high paper counts and high citation counts. Expertise is often more narrow than it appears. That's okay, as long as we're honest about it.
Don't get me wrong, I wish there was more math involved and efforts were more serious. But they aren't. The space is very noisy and little is being done to clean it up (there are some, and I do appreciate those efforts). I'll add that there's one thing more that you need besides motivation and patience: perseverance. ML systems are hard to debug and difficult to evaluate (maybe not for papers, but absolutely for systems that work in the real world). It's okay to not get things perfectly and it is totally okay to not have a model with decent generalization, but context is always important and part of the debugging process is trying to trace these down (which is difficult because you need to do more abstract versions of what is analogous to the silly or random inputs being passed to code). Detecting overfitting is often quite hard and honestly sometimes it is even desirable (GPT being overfit makes it great for information retrieval!).
Also, something I tell my students when I teach ML: you don't need math to train good models, but you do need math to know why your models are wrong. So I highly encourage math, but don't let that stop you from getting started. You can also just have a math heavy person on your team and get many benefits that way.
It doesn’t help that a lot of engineers want to find shortcuts that involve not learning the math. That’s just more engineering thinking. Not all disciplines throw exceptions when the output is bad. Maybe there will be tools that negate this need some day. I have yet to see them.
Software engineers often are missing key skills. They can learn them, but won't automatically get them in their traditional training.
First, measuring success. Actually telling how well a production system is doing is tricky. There's an art to developing metrics that tell you if an ML system is delivering value, and a lot for engineers don't have the metric design skills. Often to productionize an ML system, you need a bunch of proxy metrics and a pretty good backtesting setup. This will often depend on the specific problem, and the skill of it is something you won't get in a standard software setting.
Engineers - and especially designers - also struggle with edge cases when things go off the happy path. It's often easy to make an ML prototype that works in 90% of cases, and get a project started - but a nightmare to solve enough the edge cases for a production grade system. Finding and papering over and designing around all those edge cases effectively can require a deep bag of tricks a pure software engineer won't have.
Finally there's a struggle with tactics and culture. A lot of the bread and butter tactics of high performing software delivery are the opposite of what you need for ML projects. E.g. In high velocity frontend work you want to lock a design early, and your designer can probably do a lot of iteration before engineering starts. In ML projects you want to keep the design floating and low fidelity, as you prototype, and lock it late in the project.
So many development tactics, and cultural patterns, that lead to high performing software teams, in a SaaS setting, say, are anathema to ML projects.
YMMV. Finding and papering over the things that prevent a model from being deployable can also require a deep bag of engineering tricks that an average ML research scientist does not have. In my personal experience I've seen teams burned by this more often than the other way around.
Nevertheless, that entire field depends so much on complex heuristics and subtle optimizations that I completely understand that learning and getting an intuition for those details takes a long time (and is more akin to black magic). The development experience is absolutely horrible. Debugging a model takes such a long time, is rather expensive, and observability is absolutely dismal (at least it was a year ago). It really felt like debugging a program into existence by staring at graphs and retrying infinitely many times.
Machine learning engineers are software engineers, and they exist, so the title is wrong. I suppose it is in Nyckel's interest to claim otherwise.
I've worked with several transformers competitors, and it def wont stay centralized on them
I think it is qualitatively different from programming, because you try to find a reasonable good fit for data you know to guess new future data. Classical programming relies on rules and decisions. ML is closer to numerics, statistics, simulations. Guess the function from the data vs define a function and programm it.
* who are not 23 years old, have real-life responsibilities, and do not have infinite free time to devote to the latest hotness of the week
“Get familiar with one or more ML libraries like PyTorch, Tensorflow, FastAI, or scikit-learn. This is harder than getting familiar with a normal programming library because the concepts and paradigms are very different from what programmers are used to.”
I’d like to see this sort of logic applied to, say, doing an Autocad simulation. I think lots of people use that kind of stuff to do finite element analysis without being experts in the underlying math packages…
In particular treating datasets as "repos" analogous to how code can be versioned seems like a really good way forward.
Learning TensorFlow.js: Powerful Machine Learning in JavaScript Deep Learning with JavaScript: Neural networks in TensorFlow.js
Good luck
Is it just the provocative title?
But now machine learning is another such exception where non-trivial mathematics is important.
https://www.oreilly.com/library/view/hands-on-machine-learni...
https://www.oreilly.com/library/view/deep-learning-for/97814...
But if you really want to understand what's going on I would use a traditional ML textbook. I'm more of a no pain, no gain kind of person.
There's a 0-60 in 3 seconds one. Specifically about deep learning, which is a subset of ML. Deep learning is what is used to build LLMs ala ChatGPT.
(For someone say, who has a CS degree, took a Linear Algebra class a decade ago and doesn't remember much.)
And often the result is failure, or something close to it, as the output isn't very good.
It would have no more value than web development if the average web developer could use it well.
When I got to college, nobody talked about neural networks. Machine learning as a whole was considered a scurrilous science, wasting people and computer time. "There's not enough data. And the algorithms we have don't work! And even if we solved those, computers are too slow".
Fortunately, I managed to fail to get a job in the bio dept and as a consolation, was pointed at a nascent computational biology group in the CS department. I met David Haussler, then one of the few people in CS doing ML. Absolutely genius, he showed me a few papers and I tried to read them/understand them. The math was all over my head. It involved finding analytic derivatives of complicated functions. Fortunately, I was paired up with a grad student and given a reasonable project, where I downloaded all the gene sequence data for E.Coli (which wasn't even finished at the time) and managed to build a simple model of E.Coli genes and write an undergrad thesis that I only partly understood. To me, the magical part was watching gradient descent take those derivatives and update the weights.
When getting ready for my next phase of life, I was terrified that I wouldn't get a job as a programmer in Silicon Valley, because those folks all had CS, not bio degrees, and they knew how hash tables worked, and other complicated stuff that I couldn't wrap my head around. I decided, there was no chance I could afford to live in the valley in 1995 so I applied to grad school and got in; goal was to understand gradient descent.
My PhD was the most exhilirating and exhausting time of my life. Suddenly I was surrounded by people who could solve hard physics questions, understood quantum mechanics, and could come up with interesting experiments that got published in top journals. I felt like an imposter the entire time. But I fell in with a good group that encouraged me to explore things at my own pace, and I spent the next 7 years learning a ton of things, the capstone of which was understanding/converting a molecular dynamics loss function (including those painfully learned analytic derivatives) from FORTRAN to C++, and writing a gradient descent routine straight out of Numerical Recipes. I was thrilled that I could understand the magic of gradient descent but also depressed because that method doesn't solve the hard problems of biology (such as predicting protein structures de novo), and people were still saying that ML didn't work, there wasn't enough data, the algorithms sucked, and computers weren't fast enough.
That didn't sound right to me, because I knew that genetic data was exploding, and computers (especially cheap linux clusters) were changing access to computation quickly. The algorithms (circa 2001) were still mostly garbage, especially in biology. Neural networks to predict protein secondary structure had hit a wall at about 80% accuracy and nobody was doing ML to predict protein tertiary structure.
So I went back to school for a few more years because I still didn't know how to get a job in Silicon Valley. I did a 3 year postdoc with little to no machine learning, just domain-specific biology stuff that I didn't find interesting, and finally managed to get a job as a computer scientist at a national lab. It was a good pivot- I was a principle investigator, meaning I could apply for my own grants, write papers, etc, but didn't have to teach classes. I was MISERABLE! I loved the engineering, but the papers/grants/conference parts were just terrible.
But I still didn't "get" machine learning and wanted to work somewhere that did ML. I tried to get a job as a SWE at google- went through the ringer of all the hard questions, and ultimately got turned down at the last step (thanks, Larry Page) and went to work for a biotech for a year before I finally managed to get hired at Google during the "post-IPO, Google-classic" era, around 2007. My pay started rising faster than the average for the Bay Area which was a nice detail.
When I got to Google I quickly looked through all the projects doing ML and found that other than ads, there really wasn't a lot. There was rephil, and SETI, and SmartASS, none of which seemed even remotely like the ML I was interested in (deep neural networks). So I went and focused on other stuff- learning the distributed technology beneath Borg and Colossus, and mastering the google3 stack and production environments, mainly from an SRE perspective. But my job wasn't very demanding and I spent all my time writing proposals for Google to get involved in biology, because Google had distributed tech that was perfect for doing biology research.
Eventually, some senior engineer found my proposals and introduced me to the right people and I spent the next few years writing and running a large-scale distributed idle-cycle harvester that ran protein folding, protein design, drug discoveyr, and telescope design codes at large scale, while also learning large-scale data processing, because in speeding up the simulations, we were inundated with data to process. I got a few great publications out of this, and ended up being part of the cool kids club (coffee with Jeff Dean and Sanjay Ghemawat, etc) and parlayed this into a job building a new biology-specific platform vertical in Google Cloud so Google could make money off of, and improve the process of, biology research.
At some point I managed to tick off some senior person so I coudlnt' work on Research at Google, but finally- for the first time in my career- managed to land a job working full-time on machine learning- a system at google that almost nobody knows of called Sibyl. Sibyl was an innovative system that used an obscure ML concept- boosting- combining it with mapreduce- to run large-scale ML experiments that were directly part of the serving loop for Youtube, Google Play Ads, and other rapidly growing parts of the company. The profits from sibyl were enough to pay for all of google's research for several years and helped google grow tremendously. All that time I'd spent on machine learning and computer infrastructure... went to writing systems that loaded 80GB hash tables into memory just so a mapper could compute a tiny part of some gradient for some variable.
Unfortunately sibyl was actually a terrible system and I got kicked off the team for telling the leader the right way to do DL was deep neural networks on high performance computing hardware, not mapreduce on cheap linux cluster machines. I hid in a side team for years, playing around with 3d printers and other stuff, not really moving my career forward, but enjoying my job for the first time! At the same time I watched Jeff Dean finally realize that machine learning was an HPC problem (I think vincent vanhoucke managed to speed up voice recognition with 8 GPUs stuffed into a desktop) and he created TensorFlow, which still stumbled around for years before it realized it was an HPC system (see the slow transition to making more and more of the training process be parallel).
Finally, neural networks were vindicated! They solved a wide range of problems and my skills were applicable. We had the data, the algorithms, and the compute, all at once. And even better, you didn't need to be inside google to take advantage of it (except the big data, and that was changing quickly). I understand enough of the math, and the infra to finally be an ML Engineer.
But around that time I also came to a conclusion: most people working in ML are miserable. They are under intense pressure to get results a few percent better than their collaborators, and then once published, pivot to the next-next thing. Thats when I came up with one of my laws: "The very best ML models are distilled from postdoc tears". I saw a few people break down and leave the industry for good just from working on super-stressful projects where they did great work, but only reached parity with a competitor. and so I concluded: I was going to be ML-adjacent. This has been a succesful pivot for me.
What is the moral of this long story? Imposter syndrome drove me to overcome my imposter syndrome, and in doing so, along the way, I learned what I was chasing was not actually what made me happy. I'm far more satisfied puttering about using 5-year-old ML tech like object detectors to improve my microscope's ability to track tardigrades, than I am trying to become a famous researcher who unblocked the hard problems of biology. I guess that's part of the aging process and the stability that comes from having a salary so I don't have to worry if I can make rent.
ML has its own guild-like quality. There is some subgroup of ML people who will always try to move the goalposts, making the math harder, most esoteric, and less practical, while often publishing garbage until you peek under the covers and you realize they just got lucky, and scared away all the competitors with their Big Math. I wish people would stop doing this and instead focus on building relatively simple systems and not trying to chase 1% improvement in performance by making the system 3X more complicated.
The whole thing is just curve fitting. Literally finding some best fit curve across a series of points. This is very very easy for any software engineer to understand. I literally lost interest when I found out that the entire field was just all about messing with the data and the curve to try to get things to fit.
Literally it's just about eyeballing the data and qualitatively picking and training the thing that looks like it's the best fit. But because the data is N-dimensional and in the millions it's impossible to "eye-ball" it with your physical eyes, you have to come up with other techniques equivalent to "eye-balling" it.
Douglas Hofstadter had this whole theory of consciousness and when he found out that an LLM was a simple feed forward network with no feedback loops he went into a crisis. Basically his whole theory in GEB was wrong, according to him.
This stuff is NOT quantum physics. It's startling how simple it is and that's one of the big mysteries about it.
We only understand and build these things at a high level. At the very low level we don't actually understand what's going on. As I stated earlier we understand ML the same way a person understand data from an "eye-ball" perspective so it's impossible to even justify what exactly specifically went on with chatGPT when he answered a specific question correctly.
Did you read his book Godel Escher Bach? It's a good book. CS people love it as it's all about recursion and puzzles.
Moravec's paradox is the observation in artificial intelligence and robotics that, contrary to traditional assumptions, reasoning requires very little computation, but sensorimotor and perception skills require enormous computational resources. The principle was articulated by Hans Moravec, Rodney Brooks, Marvin Minsky and others in the 1980s. Moravec wrote in 1988, "it is comparatively easy to make computers exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility".
The resolution to the paradox is so simple I must be missing something. The amount of data in datasets for 'mobility' is basically zero. You would have to manually construct such a dataset. Whereas, humans have for thousands of years been trained to symbolically encode their reasoning processes in a way that has been incredibly accessible to computers (prose).
If I understand correctly, the scaling laws for mobility are the same for language and reasoning. We need more data.
Basically it's hard to make a machine use and understand how to use it's physical form and in 1988 it was even harder. For chess it was easy. It's easy to get, understand and use chess data.
I tend to agree that I don’t really find ML work all that interesting (much more interested in making it go fast :)), but simple it is not.
Put it this way, it's extremely challenging and not simple at all to walk and balance on a wire. Tight rope walking is not simple at all because very few people can do it.
But tightrope walking is different from something like Quantum physics. I may not be able to tightrope walk but I can understand the concept in it's entirety. For Quantum Physics, many people will never truly understand it.
What I'm saying is this, ML is tightrope walking. Challenging, not simple, but NOT quantum physics. It only seems like quantum physics.
I've always been stronger at discrete type math/programming, which is why I tend to shy away from statistics-based stuff like ML.
One thing to note is that LLMs are indeed feed forward, however the generation of the text (from my understanding) is recursive in that you feed each output token to another forward pass of the neural network.
I think there's a major misconception that ML in the form of deep learning is about statistics. There's no statistics in deep learning models. There are some statistical measurements made of final models, much in the same way a good computer science paper covering implementations of discrete data structures might make statistical statements showing the performance of the author's implementation, but like transformers and traditional neural nets and backprop have nothing to do with statistics.
LLMs are not conscious. The training process for an LLM is not a feed forward network. If we were going to try to fit the idea of consciousness a la humanity (which is really the only fully 'conscious' creature we know of) into LLMs, then 'running' an LLM is identical to cloning a frozen human, thawing it, firing some neurons, reading the result and then destroying the clone.
A better argument for actual consciousness would come from the training process, but that itself is also dubious. It's unlikely consciousness is an emergent phenomenon. Or rather, such a claim is extraordinary and would require an extraordinary amount of proof, which GEB does not provide, sorry.
See here: https://www.nytimes.com/2023/07/13/opinion/ai-chatgpt-consci...
> We only understand and build these things at a high level. At the very low level we don't actually understand what's going on.
Contradictory..
So we create a curve and estimate it. We will never know the true equation. Additionally the curve has hundreds of dimensions and is essentially something that can't be visualized or understood cohesively. We have this neural network that represents the curve but the neural network is a black box.
You're forced to solve problems with inappropriate tools, which is an anathema for a true engineer.
Same as it was with NFTs and where they are?
Same as it is with clouds and k8s;