The reason Spark ecosystem has been so popular is because it enables these types of computations without breaking the model.
The reason Spark ecosystem has been so popular is because it enables these types of computations without breaking the model.
Today’s OSS big data query execution environments are BnL neither secure nor compliant.
This is not a knock. They were simply not designed for those constraints. They were built for academic or single tenant use cases without separation of duties / control.
This also applies to commercial products such as Splunk.
There is tremendous investment going on the last couple years in raising the security and compliance bar. We’ve worked on this with the usual suspects.
But barring a handful of proprietary stacks (the big three CSPs, and a couple enterprise on prem bare metal distros) getting close, we are not there yet.
Trying to land with your security or risk teams or regulators that the CSP pulled it off will likely take you longer than provisioning a compliant laptop you’d keep locked up with two keys.
Unless you have several tens of man years invested in in-depth security wrapping these environments, or can choose a big three CSP w/o answering to anyone, today I’d still recommend the trusted laptop build approach for truly sensitive algorithms and computations.
today I’d still recommend the trusted laptop build approach for truly sensitive algorithms and computations.
You are utterly wrong. Algorithms and computations (especially ML kind) are never sensitive, its the data which is always sensitive. And that ALREADY exists on the cluster.If an Organization already has a Hadoop cluster containing data. You are suggesting that somehow having it downloaded to a secured laptop is better? Than say Spark running on top of the cluster? I think you are deeply mistaken. The cluster instances are already protected (if not you have a bigger problems). Also while an organization might not have Hadoop, they surely have an RDBMS, in which case the algorithms are even more useful.