Docker provides an (IMHO pretty buggy) isolation layer that lies between "keeping things that need to be kept separate in separate folders" and "keeping things that need to be kept separate in separate virtual machines".
I actually don't have the need for the level of isolation below VM and above folder very often. IMHO this level only really makes sense when containing and deploying somewhat badly written applications that have weirdly specific, non-standard system level dependencies (e.g. oracle) that you don't want polluting other applications' dependencies.
I've compiled and installed postgres in separate folders lots of times (super easy) and I've lost count of the number of times people have said "why don't you just dockerize that?" as if that was simpler and/or necessary in some way. That's the effect of "docker hype" talking.
The two primary use cases for Docker is, as far as I can see, is simplifying deployment on varying environments. Variations can happen because of many reasons. Sometimes you have clusters of various sizes in production. Sometimes the environment is a developer laptop. And so on.
Or if you have, y'know, a really simple script.
...the kind which also runs inside most semi-complex Docker containers anyway.
I've spent more of my life and torn out more hair dealing with obscure docker bugs than I have converting scripts from one flavor of linux to another.
> applications that have weirdly specific, non-standard system level dependencies
Spot on. 99% of software the world needs can be written against libraries provided by OSes. And then packaged properly.
I can give my coworker a docker image and it mostly "just work" without failing because she happens to be running a slightly different version of Ubuntu with different system libraries present.
> I can give my coworker a docker image and it mostly "just work" without failing because she happens to be running a slightly different version of Ubuntu with different system libraries present.
"I can give my coworker a VM image and it mostly "just work" without failing because she happens to be running a slightly different version of Ubuntu with different system libraries present."
and also:
"I can give my coworker a full system container image and it mostly "just work" without failing because she happens to be running a slightly different version of Ubuntu with different system libraries present."
A slightly smarter .tar.gz would have solved the problem just as well.
It's called "OS package" ;) and can provide more strict sandboxing using a systemd unit file: unit files provide seccomp, cgroups and more.
Nix tries to solve this, but it isn't there just yet.
Use the same OS and similar hardware for development and production.
Also means developers can work in whatever environment they want, but the result will be reproducible (almost) anywhere.
Yes, systemd unit files are containers, just like Docker.
1) is not a containerisation problem. It’s a team problem. I can jam in a load of npm and pip installs in to a shell install script. Maybe even delete /usr/ for the hell of it. Because the script isn’t isolated from the OS I can cause more damage.
This problem is actually solved by doing code reviews properly and team discussions.
2) errr no. Containers != infrastructure. If you want to deploy on bare metal, you can.
A container is vastly more powerful for running an application than a tar file.
You can often run daemons as different users and set appropriate file permissions. You can add ENV variables to your start up scripts or configuration files. Volumes are mounted by the system (and you set appropriate access rights again). Monitoring and restarting services is managed by your init system (and probably some external monitoring, because sometimes physical hosts go nuts). Depending on your environment you can just produce debs, rpms, or some custom format for packaging/distribution.
Yes, sometimes you still want docker or even a real VM, and there are good reasons for that - I totally agree. But often it is not necessary. I'm often under the impression that some people forget that the currently hyped and cool tech is not always and under every circumstance the right solution to a given type of problem. But that's not an issue with docker alone...
That sounds exactly like creating a Dockerfile. The difference is that your script has to work any number of times on an endless number of system configurations. The Dockerfile has to work once on one system which is a much easier target to hit. The "any number of times on an endless number of system configurations" is a problem taken care of by the Docker team.
before it was just a mess. and it also isn't that much older than docker.
Longer answer here https://thenewstack.io/docker-based-dynamic-tooling-a-freque...
My point was that Docker purports to solve the sandboxing and security problems.
In reality, this is something that 90% of people who use Docker don't give a shit about. For the vast majority Docker is just a nice and easy-to-use packaging format.
The sad part is that
a) Docker failed at security.
b) In trying to solve the security problem Docker ended up with a pretty crufty (from a technical point of view) packaging format.
Maybe we need to start from scratch, listen to the devs this time and build something they actually want.
Says who? The article I linked to you says nothing about security.
>Docker failed at security.
If somebody thinks security is the strong feature of Docker he/she is misinformed.
>For the vast majority Docker is just a nice and easy-to-use packaging format
For the vast majority of who? Developers? Sys admins? PMs?
The big advantage of docker is the self-contained environment for CI builds.
It's great that you compile postgres but I just want to run it in a clean and portable way, along with several other programs, and without learning new workflows for each one. Docker containers give people more options to package and run software in a simple standardized process while offloading the tedious system details that don't matter. That's progress.
As an ops person myself, docket saved me lots of time.. defeated can run their containers locally then hand then over to me to stand up. As we move to hosted services, I don't even need to maintain a server. My role of shifting from spending lots of time on ansible and monitoring servers to helping look at code and spending more time investigating weird bugs outside the developers capacity.
I was an early Hadoop adopter as well... And I agree with people's sentiment here -- it was a tool looking for a problem (outside it's specific use case). I used it for it's intended purpose, and I have with it too make it a web crawler to. It actually kinda worked in that regard, but it's not the right usage. It might be able to expand into new use cases though.
Docker solves (again) a real problem in the industry that had existed for decades... And The problems solution keeps going back and forth. Nowadays we train developers, not systems engineers (I've been trying to hire a systems engineer for almost a year and have nearly no bites... Or developers positions get 3 good candidates worth interviewing in 2 weeks or less). This means we have lots of available developers and not enough ops people. Containers help shift the burden to work in this dynamic to -- it simplifies the process to get the devs application to work in isolation. This means 1 ops guy could support a dozen developers and 30 apps on one server relatively easily compared to before. It shifts the burden of the developers runtime environment to the developer... We can still step in too help, but when file that environment is codified in git.
I've been an ops guy for a decade and unlike my positional colleagues I love Docker, it's let me focus on more important things.
This baseless assertion is patently wrong on so many levels. Building computing clusters on COTS hardware is a very mundane problem. Running processing jobs on data shards is a very mundane problem. Scaling COTS clusters transparently is a very mundane problem.
Many people use/used Hadoop for problems that did not warrant the overhead and complexity that comes with Hadoop. I've seen it countless times with my own eyes that people pre-emptively use tools like Hadoop and Spark because of a chance that they will hit a massive scale in the future.
This happens in both startups and enterprises alike: people like to think they have big problems too often.
A.k.a. resume-driven development. Having Hadoop on your CV looks sexier than awk.
I just sat there thinking I could probably run what they did on my phone.
And just to be clear - it wasn't a PoC or a demo.
Worse still, it didn’t even use HDFS and we eventually got sick of the crappy embedded Zookeeper/Kafka setup.
No, it still remains astonishingly wrong. Even container orchestration platforms are being adapted to provide the same service that Hadoop has been providing for years, and no one in their right mind would claim that running processing jobs on the cloud is a problem that almost no one has.
- Fits on one computer (most of the market)
- Fits on several computers (most of the rest)
- Requires a significant cluster of machines (50+ to store it)
Hadoop only really solves the last one. It has huge overheads in terms of speed and in terms of resources and headcount to run it properly, so it only makes sense at a particular scale. It's like a mainframe – most companies shouldn't buy one.
If you add to this the fact that Hadoop was about batch processing, and its "realtime" capabilities were poor, there really aren't that many potential customers, and many of the potential customers would rather run it in-house, or build their own system.
There's one category above that, which is "Fits in memory" and that is a huge chunk of the market. I've seen first hand people getting way too cute and complicated planning for scale, and then it works out that they don't even have more than a couple of GB of data.
Unless you're storing media, or you are truly "web-scale", your business data will very likely fit in 512GB.
I've heard of a company refusing to purchase an external drive for an employee so they could process a handful of ~50GB datasets on their MacBook Air – instead forcing them to use "the cloud" or constantly download and backup datasets.
I've heard of companies doing extensive work to set up Hadoop to process a few GB.
Roughly I'd suggest that "fits on my laptop" is <1TB, "fits in memory" is < 1TB, "fits on one computer" is < 10TB, "fits on a small cluster" is <100TB, and "might be worth Hadoop" is >100TB. I could be too low on these though.
From Heroku's site, Heroku's 4GB database plan goes for 50$/(instance.month) while Heroku's 8GB plan goes for 200$/(instance.month).
Therefore, it isn't a question of if it makes finantial sense (it does) but how long the startup plans to operate to recover their investment.
> I've heard of a company refusing to purchase an external drive for an employee so they could process a handful of ~50GB datasets on their MacBook Air – instead forcing them to use "the cloud" or constantly download and backup datasets.
I find it rather strange how someone believes that it's a decent idea to conduct a company's data analysis work on what an employee manages to fit on an external HD, as it creates a whole lot of hard problems both legal and technical. I mean, how do you ensure the data's provenance is tracked and other data analysts can access the data? Who in their right mind would put himself on a situation where a minor lapse or misfortune (losing/getting the HD stolen) could put the company at risk?
> Roughly I'd suggest that "fits on my laptop" is <1TB, "fits in memory" is < 1TB, "fits on one computer" is < 10TB, "fits on a small cluster" is <100TB, and "might be worth Hadoop" is >100TB. I could be too low on these though.
That's a rather naive and missinformed take on Hadoop. Hadoop might be conflated with big data but it's actually a distributed system designed to reliably process data shards without having to incur a penalty to move data around. It makes absolutely no sense to base your assertion on data volumes alone. What matters if it the performance increase justifies setting up a hadoop cluster with the resourses available to a company.
Now making (big)money with its ecosystem is another question.
So there should be some money there.
Something that can handle hundreds of terabytes on hundreds of machines and provides useful tools on top of the whole thing (Spark, Hive, etc)?
BTW, "hundreds of terabytes on hundreds of machines" isn't interesting territory any more. Most people's needs are far smaller, so HDFS isn't much help. Those who need more generally need much more, so HDFS isn't much help again. Richer semantics are nice either way. Imagine thousands of machines with dozens of terabytes each and you might start to see the problems with HDFS's design (though you'll still be far short of the domain I work in).
And I am curious about possible alternatives: open source, about 100-200 machines, with good support for analytics and SQlish systems.
Nope. K8s is a cluster operating system that happens to fit very well with the microservices architectural model. Packer is a way to create classic VM imgs (this one is similar-ish to Docker but only if you only care about the 10000 feet non-technical-at-all image. Ansible is an infrastructure as a code tool. You deal with mostly classic infra components and compositions of them as code. Putting all them in the same bag is like saying that all programming languages solve the same problem. It's only true if we cut the conversation down to a level where we consider all digital devices as the exact same thing.
I'm genuinely curious, if you want to search over big data, which should be a pretty common procedure these days, what alternatives are there to a distributed file system? A dfs seems very complex to me. And it is not clear too me what alternative system designs a dfs will outperform. Is a dfs the only solution to big data?
Relational DBs do break down at a certain scale. What system do you turn to next? Nosql? Will that scale infinitely? Will any system scale infinitely?
https://adamdrake.com/command-line-tools-can-be-235x-faster-...
> if you want to search over big data, which should be a pretty common procedure these days
It's not a common procedure by any stretch because the vast majority of datasets aren't really "big".
My guess is that almost all programming jobs are in fields that produce no more data than a gb or two a month.
I'm happy to be proven wrong, but I would guess that there are far more companies making project management software, time tracking apps, invoicing software, etc. than there are facebooks, googles or reddits obsessively logging every user mouse twitch.
And that's data that's much better sitting in a nice, normal, relational database.
Yes, it definitely seems with market leadership comes big data. Seems to me big data is highly relevant. More relevant now than ever.
Even Postgres got this right recently with the introduction of the BRIN index, which is a lot more lightweight.
Look at Netezza, Oracle Exadata, and (disclaimer: I work on this) SQream DB, which can absolutely handle hundreds of terabytes without too much fuss.
First, NLP based search can be executed on top of any engine (APIs are very handy), relational, kv, graph, filesystem .. so that part is totally irrelevant.
Assuming "big data" in this context is still relational data, then any of those systems would suffice, within their own particular tradeoffs and features.
If you're talking about taking some questions and getting graphs from them, ThoughtSpot does a good job.
I.e., if you are generating reports, running aggregations over a large amount of data you definitely need some parallelism and Postgres isn't designed to handle these loads (certainly not petabytes). Even aggregating 100's of GB probably requires (or at least is more cost effective using) multiple machines.
Now Hadoop may not be a particularly efficient solution unless you need 100's of machines. But there is a limit to what a non-parallel single machine database can do. There are other solutions in-between.
And you really don't have to be twitter or google to handle significant amount of data these days. People are recording much more data in the hope of generating new insights and do need tools to process that data.
Aggregating 100s of GB isn't much of a problem for PG these days. Yes, you can be faster - obviously - but it works quite well. And the price for separate systems (duplicated infrastructure, duplicated data, out-of-sync systems, ...) is noticable as well.
But yea, for many petabytes of data you either have to go to an entirely different system, or use something like Citus.
Disclaimer: I work on PG, and I used to work for Citus. So I'm definitely biased.
But I honestly don't think hundreds of users each querying 100s of GBs is all that common.
Then again it's called venture capital for a reason so this isn't exactly unexpected. The question should really be more about the scale and hype that was involved.
Btw. there are many very successful startups in the big data space that understood the limitation of Hadoop and addressed almost every if not all aspects of its shortcomings. A good example would be Snowflake computing.
How are they then querying over big data these days?
We don't know, do we? Or did they open-source their search engine?
By using Hadoop people are trying to not reinvent the big data wheel, partly because it's a motherfucker of a problem to have to solve and party because they want to solve the business problem, not the technical one. I don't see how that is in any way worthy of being frowned upon.
https://www.datacenterknowledge.com/archives/2014/06/25/goog...
https://www.quora.com/Why-did-Google-stop-using-MapReduce-an...
Map -> map, filter, flatmap, etc
Reduce -> reduce, joins, folds, group by, etc
Those other concepts were always expressible as map and reduce, of course, just with a bunch of annoying repetitive work
Here is the abstract: "MapReduce and similar systems significantly ease the task of writing data-parallel code. However, many real-world computations require a pipeline of MapReduces, and programming and managing such pipelines can be difficult. We present FlumeJava, a Java library that makes it easy to develop, test, and run efficient dataparallel pipelines. At the core of the FlumeJava library are a couple of classes that represent immutable parallel collections, each supporting a modest number of operations for processing them in parallel. Parallel collections and their operations present a simple, high-level, uniform abstraction over different data representations and execution strategies. To enable parallel operations to run efficiently, FlumeJava defers their evaluation, instead internally constructing an execution plan dataflow graph. When the final results of the parallel operations are eventually needed, FlumeJava first optimizes the execution plan, and then executes the optimized operations on appropriate underlying primitives (e.g., MapReduces). The combination of high-level abstractions for parallel data and computation, deferred evaluation and optimization, and efficient parallel primitives yields an easy-to-use system that approaches the efficiency of hand-optimized pipelines. FlumeJava is in active use by hundreds of pipeline developers within Google" [0].