Who needs MLflow when you have SQLite?
ploomber.io
ploomber.io
- simple logging of (simple) metrics during and after training
- simple logging of all arguments the model was created with
- simple logging of a textual representation of the model
- simple logging of general architecture details (number of parameters, regularisation hyperparameters, learning rate, number of epochs etc.)
- and of course checkpoints
- simple archiving of the model (and relevant data)
and all that without much (coding) overhead and only using a shared filesystem (!) And with an easy notebook integration. MLflow just has way to many unnecessary features and is unreliable and complicated. When it doesn't work it's so frustrating, it's also quite often super slow. But I always end up creating something like MLflow when working on an architecture for a long time.
EDIT: having written this...I fell like trying to write my own simple library after finishing the paper. A few ideas have already accumulated in my notes that would make my life easier.
EDIT2: I actually remember trying to use SQLite to manage my models! But the server I worked on was locked down and going through the process to get somebody to install me SQLite was just not worth it. It's also was not available on the cluster for big experiments, where it would be even more work to get it, so I gave up on the idea of trying SQLite.
It's kind of flaky and slow. It doesn't have namespacing. It's overly opinionated on the workflow (the way that states work, with a model version being in exactly one of dev, staging, prod is super hard to work with).
But beyond that, the biggest problem I have with MLFlow is what I call the "part of this complete breakfast" problem, which the ML/data-science arena is particularly susceptible to these days: the marketing talks a lot about what problems can be solved using the product, but not a lot about what parts of the problem the product actually solves. This is often because an honest answer to the latter question would be "not much". In the case of MLFlow, that would be totally fine, because honestly an opinionated CRUD app is a very useful thing. But it should be a lot more honest about what it does. It's not a system for automatically tracking model metrics, it's a database into which you can write model metrics with a known key structure.
I imagine this is because SQLite is VERY easy to bundle with Python itself. So on some platforms the OS SQLite is used, but on others it gets shipped as part of the Python installation itself.
It even works in WebAssembly via Pyodide!
Thanks for the great open source libraries yall!
Oh yes, I'm glad to see other with similar opinion.
Despite being around it for some time, I’m not sure big data or machine learning needed to be a thing for the vast majority of businesses.
"Let’s now execute the script multiple times, one per set of parameters, and store the results in the experiments.db SQLite database... After finishing executing the experiments, we can initialize our database (experiments.db) and explore the results."
Be warned that issuing queries while DML is in process can result in SQLITE_BUSY, and the default behavior is to abort the transaction, resulting in lost data.
Setting WAL mode for greater concurrency between a writer and reader(s) can lead to corruption if the IPC structures are not visible:
"To accelerate searching the WAL, SQLite creates a WAL index in shared memory. This improves the performance of read transactions, but the use of shared memory requires that all readers must be on the same machine [and OS instance]."
If the database will not be entirely left alone during DML, then the busy handler must be addressed.
When I am working with sqlite I am more likely accessing it from a single machine.
And in this case of ML, most likely from 1 process and by running multiple times in serial.
Then you just suck it up and build one of the totally unnecessary big data systems that have been excreted all over the business world these days. I don't think the problem is that devs are over-engineering.
I wonder what its called, makes me think of tragedy of the commons but probably not quite right.
Edit: Tirole, Jean. "Hierarchies and bureaucracies: On the role of collusion in organizations." JL Econ. & Org. 2 (1986): 181.
At least that was my experience a number of years back.
I think it gets us all the way once you consider the ability to expose domain-specific functions to SQL that are serviced by your application code.
I've always been of the mindset that you can do anything with SQL if you are clever enough.
Claiming that something is "broken" or "trash" when you mean "I don't like it" is a good way to make yourself feel big and smart, but it's not actually constructive.
Yes, I can understand why you comment that. I don't like blind slagging of free software either.
But there are ALSO those whose day job it is, and has been for the last 2 years, to use a badly designed overcomplex horrorshow of a tool that could be replaced easily by something better ... if it wasn't for the lock-in effects and strong marketing.
So I'm ventilating my frustration and at the same time expressing my gratitude to the person who made something fresh, that shows us things can be better.
I can't build the replacement to MLFlow myself, but I can cheer people on who do, and let them know their efforts are sorely needed.
We have another project to cover the orchestration/pipelines aspect: https://github.com/ploomber/ploomber and we have plans to work on the rest of features. For now, we're focusing on those two.
- had a giant pcap
- wrote a perl script to output some of the key value from the dump (e.g. IP and UDP packet lengths) into csv
- loaded the csv into sqlite3 database
- ran several queries to identify microbursts of bandwidth etc
The younger/more junior folks were blown away that you could do this with <100 lines of code and it was pretty fast.
Btw, above was inspired by this: https://adamdrake.com/command-line-tools-can-be-235x-faster-...
Once you have that, you can see the milliseconds with the highest bandwidth. Some extra math can also get you to Gigabits/second in a more network engineer friendly format.
Dropped it in datasette with datasette-vega and got a nice little plot
But, hordes of architects and managers who almost have a clue have been conditioned to want l and expect mlflow. And it's baked into databricks too, so for most purposes you'll be stuck with it.
Props to the author for daring to challenge the status quo.
It sparks joy in my heart whenever I see shade cast against pandas.
Close second is the plotly library.
(I also wrote a Pandas book or two... So there's that)
But for a lot of people who use it infrequently its documentation is a frustrating mess. Simple problems turn into significant time sinks of trying to find which page of the documentation to look at.
A lot of issues are made worse by shit-awful interop between libraries that claim to fully support dayaframes, but often fail in non-obvious ways... meaning back to the documentation mines.
I'd argue that because there's a market for a single author to write two books about it is indicative of documentation problems.
However, I always though the 10 minutes to Pandas page was decent for getting started. I picked up Polars recently and thought it was more difficult than Pandas because there wasn't any quick intro docs. What projects have great introductory docs for you?
Also, I am curious to learn more about the specifics of interop libraries you are referring to.
Learning a new tool is generally a challenge. I think another challenge with a lot of data tools is that non-programmers tend to be the major audience. I make my living teaching "non-programmers" how to use these tools.
That said, I always teach "go to the docstrings and stay in your environment (to not break flow) if you can." The pydata docstrings are better than most, including Python (the language).
However, maybe it makes more sense that it's just a mess that's hard to document.
The user guide material absolutely needs work, and the examples in the reference docs tend to be a little contrived. But I absolutely have seen worse-documented libraries, such as Gunicorn and Pydantic.
No, I want you to force me to provide my data in the right way and raise a noisy exception if I don't.
I agree that the magic type auto-detection is a bit too magical and sloppy, but you have to realize that data analysts and scientists have historically been incredibly sloppy programmers who wanted as much magic as possible. It's only in recent years that researchers have begun to value some amount of discipline in their research code.
x <- 3One look of dplyr code over pandas would of course disabuse anyone of the notion that R is trash and the tragedy is Python will in the current state never have anything like that. That's the advantage of the language being influenced by Lisp vs not.
I agree that it is a trash language and that, outside that many frontier academic ideas are available and some plotting preferences are solidly prescriptive, it should be thrown into the trash bin.
Python, Julia when it gets its druthers for TTFP, Octave, Fortran, C, and eventually Rust. These are the tools I've found in use over and over and over again across business, government, and non-profits.
Everywhere R is used by the org I have seen major gaps in capacity to deliver specifically because R doesn't scale well.
I agree that the standard library is what you might call "a chaotic disorganized mess".
"Trash", despite its connotations of lacking value, is really just a chaotic disorganized mess of something made by artifice with dubious reclaim/reuse/recycle value. Being a subjective assessment, it is natural that one person's trash is a treasure to another.
It's fine that people like it. What's good about it isn't unique, and what's unique about it isn't that great. And there are certainly switching costs for some orgs to consider.
Mainly because those tend to run on Microsoft Azure, which has no decent analytics offering, and are pushing Databricks extremely hard. The CTO or whatever just pushes databricks. On paper it checks all the boxes. Mlops, notebooks, experiment management. It just does all of those things very badly, but the exec doesn't care. They only care about the microsoft credits. Just to avoid using Jupyter so the compliance teams stay happy as well because Microsoft sales people scared them away from from open source.
We pushed back on it very, very, very hard, and finally convinced "IT" to not turn off our big Linux server running JupyterHub. We actually ended up using Databricks (PySpark, Delta Lake, hosted MLFlow) quite a bit for various purposes, and were happy to have it available.
But the thought of forcing us into it as our only computing platform was a spine-chilling nightmare. Something that only a person who has no idea what data analysts and data scientists actually do all day would decide to do.
I ask because normally I tend pretty strongly towards the "NO just let the DSes/analysts work how they want to", which in this case would be running Jupyter locally. However DBr's notebooks seem genuinely useful.
Is your issue "but I don't need Spark" or "i wanna code in a python project, not a notebook?", or something else?
Imo if DBr cut their wedding to Spark and provided a Python-only nb environment they'd have a killer offering on their hands.
Production workloads should be code. In source control. Like everybody else.
Notebooks inevitably degrade into confusing, messy blocks of “maybe applicable, maybe not” text, old results and plots embedded in the file because nobody stripped them before committing and comments like “don’t run cells below here”.
They’re acceptable only as a prototyping and exploration tool. Unfortunately, whole “generation” of data scientists and engineers have been trained to basically only use notebooks.
Talking to a 2000+ person org now that is standardizing data science across the org using... you guessed it
> I found the query feature extremely limiting (if my experiments are stored in a SQL table, why not allow me to query them with SQL).
Far too often, these articles of X is bad, use my homebrew Y instead, without showing comparison to X doesn't help illustrate 'why Y instead'.
You know... <cheeky>For science.</cheeky>
I don't see this scaling to many engineers working in a team, who would want to see each others experiment data, or even store artifacts like checkpoints and such. And lastly, in many cases ACLs are required as well when certain models trained with sensitive data shouldn't be shared with engineers outside of a team/group.
First, the example doesn’t take advantage of sklearn’s built in, super simple parallelization via n_jobs
Then, the entire example could be better wrapped with sklearn’s own cross_validate() which gives you the same functionality: a table of results across experiments.
If you use a different estimator, you can easily concatenate the results into a single df
The rest is the same.
Why you need SQLite for this? (SQLite is great of course for the right use cases)
And if you're doing many orders more experiments (1000s instead of 10s) then that’s probably where MLflow is good (haven’t actually used MLflow)
The company behind DVC is also building a handful of other related tools, e.g. https://iterative.ai/blog/iterative-studio-model-registry
Happy to provide more details on how it's done. It's actually quite interesting technical thing - custom Git namespace https://iterative.ai/blog/experiment-refs
I find the following workflow works well, for example:
1. Define steps depending on a `config.yml`.
2. Run an initial experiment (with an initial config) and commit the results.
3. Update config (preserving the alternate config and using symlinks from `config.yml` to various new configs if necessary), re-run, and commit.
4. Results are then all preserved in your git history.
I don't want to use Git to track all that. I want to use Git to store the final results of running such an experiment in the same commit as the code that implemented it. I just don't like the DVC experiment workflow, but I am more than happy to use DVC for storing the fitted model(s) at the end of the run.
I had researched and spent time with several other tools including DVC, GuildAI and MLFlow but finally settled on ClearML. WandB pricing is too aggressive for my liking (they force an annual subscription of $600 last I checked)
I helped build and use Disdat, which is a simple data versioning tool. It notably doesn't have the metadata capture libraries MLFlow has for different model libs, but it's meant to a lower-layer on which that can be built. Thus you won't see particulars about tracking "models" or "experiments", because models/experiments/features/intermediates are all just data thingies (or bundles in Disdat parlance). For the last 2+ years we've used Disdat to track runs and outputs of a custom distributed planning tool, and used Disdat-Luigi (an integration of Disdat with Luigi to automatically consume/produce versioned data) to manage model training and prediction pipelines (some with 10ks of artifacts). https://disdat.gitbook.io/disdat-documentation
If you work in a DS team where you're the only DS, then it probably suits your needs. Otherwise I can't imagine how you could achieve anything production grade
> There were a few things I didn’t like: it seemed too much to have to start a web server to look at my experiments, and I found the query feature extremely limiting (if my experiments are stored in a SQL table, why not allow me to query them with SQL).
While a relational database (like sqlite) can store hyperparameters and metrics, it cannot scale for the many aspects of experiment tracking for a team/organization, from visual inspection of model performance results to sharing models to lineage tracking from experimentation to production. As noted in the article, you need a GUI on top of a SQL database to make meaningful model experimentation. The MLflow web service allows you to scale across your teams/organizations with interactive visualizations, built-in search & ranking, shareable snapshots, etc. You can run it across a variety of production-grade relational dBs so users can query the data directly through the SQL database or through a UI that makes it easier to search for those not interested in using SQL.
> I also found comparing the experiments limited. I rarely have a project where a single (or a couple of) metric(s) is enough to evaluate a model. It’s mostly a combination of metrics and evaluation plots that I need to look at to assess a model. Furthermore, the numbers/plots themselves have no value in isolation; I need to benchmark them against a base model, and doing model comparisons at this level was pretty slow from the GUI.
The MLflow UI allows you to compare thousands of models from the same page in tabular or graphical format. It renders the performance-related artifacts associated with a model, including feature importance graphs, ROC & precision-recall curves, and any additional information that can be expressed in image, CSV, HTML, or PDF format.
> If you look at the script’s source code, you’ll see that there are no extra imports or calls to log the experiments, it’s a vanilla Python script.
MLflow already provides low-code solutions for MLOps, including autologging. After running a single line of code - mlflow.autolog() - every model you train across the most prominent ML frameworks, including but not limited to scikit-learn, XGBoost, TensorFlow & Keras, PySpark, LightGBM, and statsmodels is automatically tracked with MLflow, including all relevant hyperparameters, performance metrics, model files, software dependencies, etc. All of this information is made immediately available in the MLflow UI.
Addendum: As noted, there is a false equivalence between an end-to-end MLOps lifecycle platform like MLflow and tools for experiment tracking. To succeed with end-to-end MLOps, teams/organizations also need projects to package code for reproducibility on any platform across many different package versions, deploy models in multiple environments, and a registry to store and manage these models - all of which is provided by MLflow.
It is battle-tested with hundreds of developers and thousands of organizations using widely-adopted open source standards. I encourage you to chime in on the MLflow GitHub on any issues and PRs, too!
We'd love to work with the author to make MLflow Tracking an even better experiment tracking tool and immediately benefit thousands of organizations and users on the platform. MLflow is the largest open source MLOps platform with over 500 external contributors actively developing the project and a maintainer group dedicated to making sure your contributions & improvements are merged quickly.