In-Database Machine Learning [pdf]
btw.informatik.uni-rostock.de
btw.informatik.uni-rostock.de
Jesus flippin Chrikey kill it with fire. Academics; where do they come from?
Having scoped out a design for this, seen a design for it in a large corporate research lab, and noting that several large scale machine learning packages .... almost do it .... themselves: you just need to lay out the data in a memory friendly way, and you're freaking done. It's a completely straightforward task which could be done with any number of existing columnar data stores if it hasn't been done already. The only reason it hasn't been outside of (large research lab) is the type of skill needed to lay out the data in the DB properly and the type of skill needed to write efficient machine learning algorithms almost never exist in the same person. Then you get academics .... who think you need to write automatic differentiation in SQL to integrate ML with databases, having never heard of, I dunno, online BFGS or Pegasos, and, like, C.
I'm curious. In your opinion, is there a scenario where there is a possibility and need to do what the authors have done?
The only reason I thought of doing it was skipping the "marshall the data" step for online algorithms; performance basically. If you look at something like Vowpal Wabbit, it stores/consumes data in this ridiculous text format dating from when people invented the Support Vector Machine in the 80s. This annoys the shit out of me, as it involves creating the ridiculous 80s text format. What if it consumed database columns instead of this nonsense, and did the hashing trick and everything you needed to solve the problem? That would be cool. That, in fact, would be something like how the universe is supposed to function instead of the stunted grotesqueries we have today. You could do crap like exploratory analysis right on your database, as a query. You could even get fancy and do wackadoo online matrix decompositions while you're writing the data out in the first place (or at least when nobody's looking), and store it as metadata, meaning you know all kinds of good shit about your data even as you're writing it down. Anyway, because marketing departments keep bellowing about "deep learning" instead of the actual breakthroughs in machine learning and linear algebra of the last 20 years, nobody gave a shit about it. Even (large research group in gigantor corp) couldn't figure out a way of selling the idea. I went on to a productive career in something entirely different, and all I got out of it was the ability to make snarky comments about seemingly clueless academics.
Maybe you want to use an encrypted DB without decrypting it?
This is research, not a vanilla solution.
There are instances where researchers spend their time where the outcome was mostly thought experiment / not directly applicable, and there are instances where the work was useful from an application perspective. Such is life
They care about mathematical background and laying out definitions. Academics care about those pointless things. This is also fine. If you don’t care about these things, don’t read it.
You don’t write a system like this to say: “Here’s a system I wrote, go and use it in practice” as much as “What have we learned from trying this?” It’s the job of an engineer to synthesize useful things from the abstract knowledge.
You failed, first sentence.
You're also the second person asserting some secret knowledge of academic ding dongs who published a paper (as if this is some kind of achievement) without, you know, actually even attempting to make the case.
And I’m not following. I’m not asserting some secret knowledge. This work is knowledge. It’s not practically useful for you. But it doesn’t have to be. And you’re the “ding dong” if you think it has to be valuable to you for someone to get something from it.
The advantage of this approach is that data is never moved outside SQL Server or over the network. The downside I guess is that you need a pretty beefy machine to run the database server.
For uncomplicated ML applications, say a logistic regression over a few columns, this is a relatively easy approach to get results quickly. To me, the actual use cases of in-db ML are limited, but the one case in which I can imagine it being useful is performing live ML on a SQL view that has constantly evolving data -- you save ETL roundtrips to an external ML algorithm.
[1] https://docs.microsoft.com/en-us/sql/machine-learning/sql-se...
Because eventually the smart datastore is just an unperformant dumb one.
VC buddy of mine invested in such a company. So basically the RAM can do small things by itself. Should be very interesting for Databases and HANA. When he and a major Chip producer invested, I even considered asking SAP if they want to be on board.
Data extraction time is insignificant compared to actual training except for trivial models. Databases can read hundreds of thousands of items per second while a model can only process dozens to hundreds.
This is to me a good candidate for an ML workload to run directly in a database:
- Low compute vs storage ratio for the models.
- High number of models.
- Often only want a small subset of data as input to model (a few numbers typically).
- Relational data highly relevant, for contextual data around the entity.
- Simple models with few parameters to store.
- Frequent updates to models.
- Historical models interesting. To implement checking new models against old, running in parallel
Other examples with similar characteristics would be Timeseries Forecasting on many different time-series. Could be sensor data, stock tickers or whatever.
* Better to use a robust analog such as Median Absolute Distance.
But more effective (but still rather simple) models, like using Malahobis Distance, or kNearestNeighbours distance is harder to do efficiently.
One rather mature project for doing ML integrated in SQL databases is Apache Madlib. It has been around since 2015. https://madlib.apache.org/
There is no advequate tooling and all you get is a horrible SQL Developer.
It means that you won't have a debugger, you won't have unit tests, you won't have profilers.
Debugging it is a pain in the ass. Even setting up the environment for work is pain in the ass because setting up Oracle server is pain in the ass.
Ecosystem is non-existant.
If somebody wants to write business logic using it, why not use a normal language like Java or C#?
In general, having business logic inside the database is almost always universally a bad idea.