Machine Learning in Production (CMU Course)
mlip-cmu.github.io
mlip-cmu.github.io
Is it too entry-level? Looking at the labs, a lot of this seems like stuff a mid-level software engineer (or even a motivated beginner) could pick up on their own with tutorials. Git, Flask, container orchestration... all useful, but pretty basic for anyone who's already worked in production environments. The deeper challenges—like optimizing networking for distributed training or managing inference at scale—don’t seem to get as much attention. Maybe it comes up in the group projects?
Also wondering about the long-term relevance of some of the tools they’re using. Jenkins? Sure, it’s everywhere, but wouldn’t it make sense to introduce something more modern like GitHub Actions or ArgoCD for CI/CD? Same with Kubernetes—obviously a must-know, but what about alternatives or supplementary tools for edge deployments or serverless systems? Feels like an opportunity to push into the future a bit more.
That's what I was wondering about too. It seems to me that eventually someone will build a tool that runs any neural network on any hardware, whether local on one machine, or distributed in the cloud.
Relevance? Is there really a huge conceptual difference between Jenkins and the other CI/CD frameworks? If not, if I were them I would just choose a random popular one, and it seems to me that's just what they did.
Does it mean that ML ops is a field that's actually not that hard to approach for a regular developer without a PhD?
Like say you want to generate tons of synthetic data using simulations, you are likely to be more interested in questions say of batching, encoding formats, data loading etc than the actual process of generating unbiased data sets
If you need to collect and sample data from crowd sourcing, you likely need to know less about reservoir sampling than say figure out how to do it, online so it's fast or be efficient with $$$/compute spent on implementing the solution etc.
It's funny to me where statistics ends up sometimes.
For instance, my ML work is almost entirely in the context of engineering simulation regression/surrogate development, where data quality/cleaning is almost no issue at all - all of the work is on the dataset generation side and on the model selection/training/deployment side.
Every job is different!
Looks good to me! Same with the LLM Systems course.
You missed model serving (I think ?). Tough and latency sensitive esp for recommender systems. Prone to latency spikes, traffic spikes. Even with a well-written Python code you can run into limitations quite quickly.
> Nothing fancy.
Well, right now I am seeing lots of low-level innovation for networking/storage along with RoCE, Infiniband, Tesla's ttpoe, the recent addition of devmem-tcp to the linux kernel (https://docs.kernel.org/networking/devmem.html) and wondered if there are approaches on how to plug something like that together on a higher level and what the considerations are. I surely assume EFS or S3 might be too expensive for a (large) training infrastructure, but I can be wrong?
> You missed model serving (I think ?).
I think I have a better grasp on the engineering challenges there and could imagine an architecture to scale that out (I believe!).
So do an end-to-end project where you:
- start from a CSV dataset, with the goal of predicting some output column. A classic example is predicting whether a household's income is >$50K or not from census information.
- transform/clean the data in a jupyter notebook and engineer features for input into a model. Export the features to disk into a format suitable for training.
- train a simple linear model using a chosen framework: a regressor if you're predicting a numerical field, a classifier if its categorical.
- iterate on model evaluation metrics through more feature engineering, scoring the model on unseen data to see its actual performance.
- export the model in such a way it can be loaded or hosted. The format largely depends on the framework.
- construct a docker container that exposes the model over HTTP and a handler for receiving prediction requests and transforming them for input into the model, and a client that sends requests to that model.
That'll basically get an entire end-to-end run the entire MLE lifecycle. Every other part of development is a series of concentric loop between these steps, scaled out to ridiculous scale in several dimensions: number of features, size of dataset, steps in a data/feature processing pipeline to generate training datasets, model architecture and hyperparameters, latency/availability requirements for model servers...
For bonus points:
- track metrics and artifacts using a local mlflow deployment.
- compare performance for different models.
- examine feature importance to remove unnecessary (or net-negative) features.
- use a NN model and train on GPU. Use profiling tools (depends on the framework) and Nvidia NSight to examine performance. Optimize.
- host a big model on GPU. Profile and optimize.
IMO: the biggest missing piece for ML systems/platform engineers is how to feed GPUs. If you can right-size workloads and feed a GPU with MLE workloads you'll get hired. MLE workloads vary wildly (ratio of data volume in vs. compute; size of model; balancing CPU compute for feature processing with GPU compute for model training). We're all working under massive GPU scarcity.
curious: which part of the pipeline does the majority of 'business' value come from?
All the technology challenges are actually on the "cost" side of the equation. Meaning, that the aim wrt business value should be do as little of it as possible (but not less!). For some use cases this can still be quite a lot... But more often on the "all the pieces need to be in place for the whole to work at all" rather than "each piece needs to be super optimized".