We Need DevOps for ML Data
tecton.ai
tecton.ai
When you have engineering team separate than a data science team, you'll inevitably have unproductive conflict & politics. One team might be incentivized for stability and speed (engineering or ops) and the other model accuracy (data science). The end result can be disastrous... An engineering team that wants to bend nothing to help data scientists get their work in production. Or a data science team that only cares about maximizing accuracy, even if it might destroy prod, or be impractical to implement in a performant way.
To hit the sweet spot on accuracy, speed, and stability, you need to have one team that focuses on the end feature. It needs to be cross-functional and accountable for doing a great job at that feature. And the data scientists need to be possibly more focused on measuring and analyzing the feature's success, rather than just building models for models sake.
I'd recommend the book Agile IT Organization Design if you're interested in good team design patterns
The reality is there’s more work to embracing ML than hiring data scientists. Everyone needs to understand ML a little, and it needs to be OK to critically question data science work from product and engineering angles.
We see this in data science and machine learning where people complain about spending their time cleaning data, etc... when their time should be spent "generating insights/etc." We also see that those insights are interesting but not very useful if they aren't actionable, too costly or too impractical to implement.
Ultimate value is related to being able to contribute to and achieve the holistic outcome, but the lens of success is often focused on models or insights instead. This is a cultural and organizational problem, rather than a technological one. It also takes a dose of humility to appreciate the true value of the so-called dirty work.
Spoke to an experienced engineer who used to lead NLP at MSFT and same comment. NLP models are already fantastic and it isn’t very hard to build a smart chatbot. The implementations these days are just very poor because they are not well thought out from a user perspective.
Optimizing a loss function is far far easier than finding the right loss function(s)
Which is the whole idea behind DevOps: to break down the barriers between development and deployment by focusing on rapid iteration to production by continuously integrating changes into that pipeline.
It's ironic that DevOps has become a specialty in and of itself. The idea is to get rid of separate teams, not create a new one!
Getting ML models into production isn't particularly hard... if you put an engineering team on it that know how to write automated release procedures, design architecture that can scale and build robust APIs to surface the data.
But in many companies the engineers with those operations-level skills and the researchers who work on machine learning live completely separate lives. And then the researchers are expected to deploy and scale their models to production!
That's not to say that this organizational problem cannot be solved with technology/entrepreneurship. If a company can afford it it's likely much cheaper to pay an external company to solve your "ML in production" problems than to re-design your organization such that you equip your internal ML teams with the skills they need to go to prod.
It’s actually quite complex, which is why generally speaking very few people do anything like this. I am unaware of any general solution to this problem, either in industry or academia.
Would love to chat if you have further thoughts around the subject - there's a ton of problems we're looking to tackle in the space and would be good to get input.
Building and maintaining ML infrastructure from scratch is a big project. That's why you see FAANG companies hiring for ML infrastructure/platform engineers. Most startups don't have the extra cycles for that big of an undertaking, and so you see a lot of slapped-together, hacky solutions to putting models into production.
I'm biased in that I work on Cortex ( https://github.com/cortexlabs/cortex ), but I think that open source, modular tooling that removes the need to reinvent the wheel is going to have a big impact in terms of making production ML more accessible.
I think surfacing the data is just the first step, often times data scientists need to run some data exploration, the process is generally iterative, and so they need to run several experiments, resume or restart some experiments, scale training with distributed learning using several machines, or run hyper-parameters tuning, which means handling failures, visualize and debug results, before deciding if they should deploy a model. Once a model is deployed the story does not end there, because models become stale and need to be retrained. There are other issues related to compliance that need to be handled as well, and many other problems related to governance, a/b testing, ...
The good news is that there are several open source initiatives to solve several of these problems, at Polyaxon [0] we are trying to solve some of the aspects related to the experimentation phase.
A while back, we published a blog post that discusses how we approached these organizational challenges at Uber: https://eng.uber.com/scaling-michelangelo/. With Michelangelo, we found that the right tooling can both solve technical challenges and help with some organizational challenges. For example: If a standardized and centralized platform is the path of least resistance to get ML into production and solve your business problem, you get the organizational benefits of that centralization (governance/visibility/collaboration) along the way.
https://papers.nips.cc/paper/5656-hidden-technical-debt-in-m...
As someone else says in this comment thread, this is very much an organizational problem, and cannot be viewed as just a technology problem.
The common behavior of individuals and teams is the pursuit of solutions that solve problems for them. Problems here with ML, and as we've seen with "Data Science," along with other magic technologies is that having an appreciation for the domain or context goes a long way. Being familiar with entire process, or "pipeline," is valuable, and role/functional silos often lead the problems people experience.
For some classes of machine learning problems and associated data, sourcing solutions from vendors can work, but as with any tools you can procure, you need the right people to use them appropriately. This also applies to "DevOps" which is used for comparison in the blog post.
--> DevOps example -- the philosophy seems to be about having software developers also share build/release and infrastructure responsibility. But some organizations have made "DevOps" teams to silo build/release and infrastructure work... they ended up renaming what used to be called their Build/Release or SysAdmin teams. Siloing things to be "someone else's" problem doesn't result in the major transformations that are needed.
Now imagine what happens if we substitute MLDevOps for DevOps above.
I'll continue to say "The Role of a Data Engineer on a Team is Complementary and Defined By The Tasks That Others Don’t (Want To) Do (Well)"
Those types of tasks are also often not recognized or rewarded by management, despite being a hugely critical part of the system. I believe the incorrect hiring of scientists who are often strong in terms of core theory or number of papers published but have no clue about building real production ML systems is a huge organizational problem, often causing ML teams to fail to deliver any real value.
Really what we need is version control for data, it's not just an ML data problem. It's a little different though, because you would like to move computation to data, rather than the other way around
It seems to me to be able to time-travel in data you almost need to store the Write-Ahead Log of database transactions and be able to replay that. Debezium captures the CDC information, but it's a infrastructure level tool rather than a version control tool.
In data science, most time-travel issues are worked around using bitemporal data modeling: which is a fancy way of saying "add a separate timestamp column to the table to record when the data was written". Then you can roll things back to any ETL point in a performant fashion. This is particularly useful for debugging recursive algorithms that get retrained every day.
But these are infrastructure level approaches. I'm not sure that it's a problem for a version control tool.
https://www.dolthub.com/blog/2020-04-01-how-dolt-stores-tabl...
Wouldn't discovering what those changes are still entail heavy database queries? Unless Dolt has a hook into most SQL databases' internal data structures? Or WALs?
`SELECT * FROM dolt_diff_$table where from_commit = '230sadfo98' and to_commit = 'sadf9807sdf'`
Right now, Dolt can't be distributed (ie. data must fit on one hard drive) easily so it's not meant for big data, more data that humans interact with, like mapping tables or daily summary tables. But, long term if we can get some traction, we plan on building "big dolt" which would be a distributed version that can scale to as big as you want.
So for most analytic workloads, typically a columnstore db is used due to the need for performance and advanced SQL features (windowing functions) for complex analytic queries -- which I don't expect Dolt to replace. Which means if we wanted to use Dolt's features, we would have to continuously ETL the data into Dolt, which would entail mirroring the entire database (or at least the parts we want to version control).
Dolt essentially becomes a derived database specifically used for versioning. I see how this might work for some use cases.
Until some major data drift happens, but you would notoce it anyways
For these types of recursive model applications, you cannot just fit the model once and forget about it.
Not quite, this is "transaction time". You also need "valid time" to be truly bitemporal. Recovering the database as of some point in time is not enough to answer questions like "when will this fact become false?" or "when did our belief about when it would become false change?", because you didn't preserve assertions about the time range over which the fact was held to be true.
In terms of implementations, ranges are better than double timestamps. They provide their own assertion of monotonicity and can be easily used in exclusion indices.
I found that Snodgrass's textbook was a good introduction to the concepts and it's available for free: https://www2.cs.arizona.edu/~rts/tdbbook.pdf
Thank you for the link to Snodgrass' book. I've not seen a formal book on temporal modeling in SQL before, so this is fascinating.
Some notion of bitemporalism showed up in SQL 2011, but somewhat constrained compared to what Snodgrass describes.
I totally agree that data management problems are not just ML related. But I personally think that there are additional challenges in the space that are not just version control for data.. all the area of data quality management and monitoring for example. I liked the analogy to devops, source version was super critical problem to solve in software development, but it didn't stop there, with things like CI/CD etc. I believe we'll see similar evolution in the data space..
https://www.dolthub.com/blog/2020-03-06-so-you-want-git-for-...
1. To replicate models (needed for regulatory reasons), you need to commit both data and code. If you have only a few models, fine just archive the training data. But, if you have lots of models (dev+prod) and lots of data - you can't use git-based approaches where you commit metadata and make immutable copies of data. It scales (your data!) badly. We are following the ACID datalake approach (Apache Hudi), where you store diffs of your data and can issue queries like "Give me training data for these features as it was on this date".
2. You want one feature pipeline to compute features (not one for training and a different one when serving features). Your feature store should scale to store TBs/PBs of cached features to generate train/test data, but should also return feature vectors in single ms latency for online apps to make predictions. What DB has those characteristics? We say none, and we adopt a dual-DB approach with one DB for low-latency and one for scale-out SQL. We use open-source NDB and Hive on our HopsFS filesystem - where all 2 DBs and the filesystem share the same unified, scale-out metadata layer (a "rm -rf feature_group" on the filesystem also automatically cleans up Hive and feature metadata)
3. You want to be able to catalog/search for features using free-text search and have good exploratory data analysis. The systems challenge here is how to allow search on your production DB with your features. Our solution is that we provide a CDC API to our Feature Store, and automatically sync extended metadata to Elastic with an eventually consistent replication protocol. So when you 'rm -rf ..' on your filesystem, even the extended metadata in Elastic is automatically cleaned up.
4. You need to support reuse of features in different training datasets. Otherwise, what's the point? We do that using Spark as a compute engine to join features from tables containing normalized features.
References:
* https://www.logicalclocks.com/blog/mlops-with-a-feature-stor... * https://ieeexplore.ieee.org/document/8752956 (CDC HopsFS to Elastic) * http://kth.diva-portal.org/smash/get/diva2:1149002/FULLTEXT0... (Hive on HopsFS)
- How can I deliver these features to my model in production?
- How do I make sure the data I'm serving to my model is similar to what is trained on?
- How can I construct my training data with point in time accuracy for every example?
- How can I reuse features that another DS on my team built?
We've found that there's a ton of complexity getting data right for real-time production use cases. These problems can be solved, but require a lot of care and are hard to get right. We're building production-ready feature infrastructure and managed workflows that "just work" for teams that can’t or don’t want to dedicate large engineering teams to these problems.
At the core of Tecton is a managed feature store, feature pipeline automation, and a feature server. We’re building the platform to integrate with existing tools in the ML ecosystem.
We’re going to share more about the platform in the next few months. Happy to answer any questions. I’d also love to hear what challenges folks on this thread have encountered when putting ML into production.
As AWS shows, proprietary all-in-one [platform] is fine as long as it's a-la-carte.
It's easy to do "dev ops" for machine learning. Basically, just automate everything and implement gatekeeping mechanisms along with active monitoring.
It's true, though. I had to cobble together a lot of custom things at the time, but it wasn't that hard to do.
It's a handful of technologies, but they're (generally) mature, battle tested, and well documented.
- managing different data sources
- versioning data
- monitoring how new data affects the model
- testing that certain SLAs are met before new features are deployed
- ability to rollback
- data & model quality monitoring
is technically challenging.
Obviously there are engineers that will quickly hack something together and will falsely think that they have a good enough MLOps solution. I have been part of such teams.
Most companies are not Google, Facebook, or Uber. The large ones very often don't have the know-how to create a robust technical solution around this process, and even if they do it can take them years and the smaller ones lack both the resources and technical expertise.
I'm always looking for new ideas that can become successful business and when I saw the Uber Michelangelo here on HN a few years ago, I was thinking that selling similar tooling to other companies, had great potential. Seems that the right team to create that company was the one the built Michelangelo itself :)
Data pipelines are a real problem though, and I'm very interested in what startups do with this space.
Can you please elaborate more, thanks.
"Static ETL" like running the same database load every day at 1:00am isn't a super challenging problem. Doing it across many tables with complex transformations and multiple steps easily can be. You really have to consider reliability, processing speed, failure methods and other problems that dont really arise until you hit a certain scale.
There's also the issue of what people want out of a Pipeline that's changing. If you want people to be ""data driven"", then that means they need easy access to potentially all of your company's data on an ad hoc basis. So now your boring ETL 1 am pipeline isnt really serving any of these new usecases.
How do you create flexible pipelines that can be created from any dataset on an ad hoc basis? This is where tools like Airflow or Prefect come in. Creating a platform that can create these types of Pipelines is a real problem.
And before you even ask yourself _how_ to process this data, you need to also ask _where_? If you want to do what I outlined above - making your data more accessible and easy to use - then you probably need to rework how you're storing your data. But Data Lakes (and others) are a whole topic in and of itself.
Unless the issue here is data collection in prod to start training your model.
The most common bottleneck is collecting the right data. It can take years, or even a task force just to get the right data before the data scientist can begin.
>I would have thought "a model that works" is much harder
It depends how experienced the data scientist is. Early on into a project a data scientist can do a feasibility assessment. They should identify what is possible, and how possible. Sometimes some data science projects are heavy on the research side where where it can take 2 weeks to 3 months to figure out if something is possible. Sometimes the feasibility assessment ends up being incorrect and a goal is shown to be impossible.
Once research is done it usually takes 4 weeks to 6 months for a data scientist to build a model. The upper bound is rare and happens because of recursive refinement to increase accuracy, trying to get every last drop out of what is possible.
In contrast it can take months to years for the company to begin to collect the right data for a data scientist to be able to begin to do what benefits the company. Sometimes crowd source projects need to be created just to collect the required data. It then takes an average of 3 to 6 months for productionization if there is clever feature engineering in the model. Note: When I say productionization, I mean all the way to the end customer, so setting up and maintaining pipelines, frontend devs updating websites to add the service, and whatever else is necessary. There is more work involved on the production side, but it can be split up to multiple engineers.
IT own the platform, and the software. They should never own the data as well.
I had to search to see it was Machine Learning.
How hard is it to define it the first time you use it?
I can bet lots of people were scratching their heads but didn’t bother to look it up or continue reading...