Who eats who in open source
platformonomics.com
platformonomics.com
I see this dynamic more and more, companies have figured this out. Sure, buyers prefer open source infrastructure without vendor lock-in, but they don't prefer it enough to spend 2-10x the OpEx, which is often the case in these discussions. The poor optimization of operational costs (except for license fees) is a critical weakness in open source, and it is increasingly being attacked successfully. Open source is a nice idea that most companies love, but they aren't going to spend mountains of extra money in operational costs to get it.
Efforts like https://jortage.com can reduce OpEx quite significantly.
I feel the opposite is true. There are many open source projects that are at least on par with their proprietary counterparts. Postgres is competitive with every SQL DB I've heard of except maybe in some niche use-cases. Redis is open source and incredibly performant. Same with Kafka (if you use it well). Linux sure runs better than Windows in my experience.
There is also a whole class of open-source libraries like React that have taken over. I can't even think of a proprietary frontend framework.
What software are you thinking of when you make these claims?
I believe the point being made was the money you save on licensing fees are insignificant compared to the amount you spend on paying people to maintain your infrastructure.
Postgres is a nice database, but there are enterprise deployment use cases that it doesn't cover, which might be consired a niche, fair enough.
What I don't consider a niche use case, is that its development tooling, meaning graphical DB management, graphical debugging of stored procedures, data modeling tooling, or integration of language runtimes into the database are all a bit behind of what those big DBs are capable of.
First, there is a vast amount of supported operational tooling that is simply missing from open source ecosystems. PostgreSQL, which I still use extensively, is a perfect example of this. The code bases around this tooling are several times the size of the product they support, and it isn't the kind of code that most developers aspire to write despite its high value to the customer. Oracle and SQL Server have architectures that are as obsolete as PostgreSQL, they don't compete on the basis of being modern, efficient designs but on the basis of having dramatically better operational tooling. Cloud RDBMS attack a different aspect of this.
Second, is the operational infrastructure costs i.e. how much hardware is required to run a given workload. This is the larger threat to open source, particularly as the data intensity of business increases. PostgreSQL, Kafka, and Redis (I've used all three operationally) have several times the hardware requirements to deliver a workload in practice than is required with a state-of-the-art architecture. This isn't hypothetical, I've designed engines that replaced them when warranted. An 80% reduction in infrastructure cost is attractive when you are already spending tens or hundreds of millions of dollars per year on it. Some popular proprietary software is just as inefficient but that isn't guaranteed to remain the case (and I've replaced those systems too). Designing efficient software architecture isn't magic (see: ScyllaDB), you simply see very few examples of it in open source because it isn't a priority (see: use of managed languages for data infrastructure despite this being a known limitation).
The reality is that I can guarantee data intensive businesses who make competent use of open source that I can reduce their infrastructure costs by 4x with ease. You can bury a massive amount licensing and engineering costs in those savings.
Also, some companies are explicitly looking at this as a major "green" initiative. The wasteful infrastructure footprint to deliver a workload with open source is viewed as environmentally unfriendly, and this is pushing companies to consider proprietary solutions even if they are not sensitive to operational costs. That isn't a conversation open source is currently prepared to have but it is coming.
I am interested in your perspective and experiences.
> Oracle and SQL Server have architectures that are as obsolete as PostgreSQL, they don't compete on the basis of being modern, efficient designs but on the basis of having dramatically better operational tooling.
What dramatically better tooling does SQL Server has?
> PostgreSQL, Kafka, and Redis (I've used all three operationally) have several times the hardware requirements to deliver a workload in practice than is required with a state-of-the-art architecture.
What should I think of when you talk about state-of-the-art architecture? And what would replace aforementioned tools?
What should the open source offerings do to win you over?
I don't know about Oracle but SQL Server Data Tools (SSDT) and SQL Server Management Studio (SSMS) are vastly better than anything I've seen for open source databases.
SSMS is just a really good visual client with an excellent table/view designer, diagram builder and query builder with GUI management features that let you manage virtually every facet of your server and databases.
SSDT lets you treat your SQL schema, database settings and deployments like code, with versioning. All of your table schemas, views, functions, procedures, settings, seed data, etc. are stored as text in your git repo. When you're creating, it has excellent intellisense/autocomplete. You can use it to diff your schema and data against a given SQL Server database to generate migration scripts. You can also do data comparisons between your seed data and your server and you can diff schemas and data between 2 different servers without even having created an SSDT schema. Another very useful feature is that you can reverse engineer a SQL Server database that was already created in order to start your SSDT project. It also has visual designers for all of the aforementioned things. It used to be a separate tool, but now it's part of the free Visual Studio community edition. Both tools are free, as are various editions of SQL Server.
Also, SQL Server itself has had features for a long time that PG is just starting to get such as real Stored Procedures, ones that can return multiple heterogeneous result-sets, run transactions and do control of flow with SQL-like statements (IF/THEN, DO, WHILE, etc.) It also has table variables, which are like in-memory tables, that you can use as a faster type of temp table for certain sets of data.
Since I moved off of Windows for my workstations, I've been running SQL Server in Docker on Linux for the past few months - it's rock solid.
@jandrewrogers also mentioned a 4-fold performance increase by switching to proprietary solutions from pgsql. I am still wondering what those solutions might be...
---
Have you tried the SQL Server based products at all?
This is the developers perspective. In the time it takes to develop something the licence fees are indeed at tiny part of the cost.
However, over the lifetime of the software itself it's usually a different outcome. Software typically lasts decades. Banks are still running cobol software written decades ago. Perhaps it's possible even over those long lifetimes the licensing costs still win. When it breaks down completely is when you deploy that software multiple times, paying those licence fees each time you do it. If you start paying for 1000 CAL's you have a problem.
The other comment here that is a red flag to me is "all the tools that give you visibility into the database save you time". If it is development time then well and good, but if you are needing DBA's once the thing is deployed then just no. No matter how sophisticated those tools are, they still cost far more than needing no attention whatsoever. This is where open source usually wins, because it's lack of sophistication is usually an case of being pared back to the bare minimum to do the job.
As always, 'commoditize your complement': https://www.gwern.net/Complement
Compare the number of platforms nowadays which (to varying extents) treat Linux as a free set of buggy device drivers.
The author is talking about physical infrastructure as adding value, and they're right, but there's a middle layer that cloud vendors add to their managed tools -- things like auto backup, upgrade and resize for mysql.
These are enhancements to open source DBs etc that cloud vendors keep in-house as secret sauce. This makes the stock version of open source software hard to operate in the way AMZN / GOOG operate it, while still allowing AMZN / GOOG to benefit from community effort without giving much back.
The original thesis would be that brand loyalty would be strong enough to prevent forking like this. "You don't want that third-rate version, get the real thing from the original authors!". It appears that the customer base assumes quality is roughly equal considering its mostly the same code.
It's not possible to say the code is non-free because it's explicitly free, but packaging and placement (such as e.g. default pip installs using a trademarked name) could ensure at least some pain is felt to make a typical client talk to a server that does not have that file. Perhaps the file's license allows to be embedded anywhere so long as it is only used to play the role of a client
Open source is a good business strategy when you're trying to achieve vendor lock-in via price-dumping, a la what Google did with Android (and has been fined for in the EU).
The fact that a company backing an AGPL-licensed work goes under isn't the direct concern of the license. The fact that the code itself remains as Free Software, with a source-release obligation, even when run as an online service, is the principle concern.
Copyleft free software licenses, as with the GPL and AGPL, promote the concept and system of free software, first and foremost.
Under the AGPL, the source-disclosure provision persists.
In both cases, Amazon is doing software development, but in the first, they're not contributing back to the codebase. In the second, they are.
Amazon are maintaining the software. Along with other users.
Note that AGPL applies to few projects at this point. Amazon is very heavily reliant on, e.g., Linux and Xen, both of which are GPL v2.
How do computers earn people money? Speaking in 2019 that's obvious. Cloud providers earn the most.
There used to be other ways people got rich off computers. But all those businesses imploded against the awesome might of '0 marginal cost', 'only good solutions survive', world of open source.
Open source has 'removed more profits'/'provided more value' then the cloud. Cloud is simply the high profit game until competition kicks in, open source tools become the norm, and only the value of the virtual machine is sold at 1% the current price.
> If you squint, open source could be seen as a very generous charitable donation to some of the largest and wealthiest corporations on the planet.
Broadly: yes. Except backwards.
FLOSS can be seen as a public good -- it is non-excludable and non-rivalrous. Excludability is the property that someone can be prevented from using the good (eg, requiring payment to consume a can of soft-drink). Rivalrousness is the property that utility from consumption of a good by one person diminishes utility for another person (if I drink the soft-drink, you get less of it).
Economics predicts that public goods will be underprovided by a pure market. This is because of the free rider problem. Since I can't be excluded from the good, I can consume it without giving something up for it. Since it's non-rivalrous, there's no meaningful back pressure that eventually raises the cost of consumption to an unacceptable level.
A strictly rational agent will always free ride on a public good. And most of the time, most rational agents will not provide a public good, because the cost of paying for everyone else's consumption exceeds the benefits of their own consumption.
This hints at one of the ways that public goods get provided: through subsidy by benefactors, who will capture some but not all of the value created by their benefaction. A wealthy person may so enjoy seeing the opera that they will donate heavily to the local opera house. They don't capture the full benefit -- other folks can watch the same shows -- but they capture enough value that they are satisfied with the arrangement. They might also get value from other factors, such as social approval.
The cloud providers do not sell public goods. Their services are fully excludable. If you refuse to pay, service will end. They are rivalrous at the limit, though, putting them into the category of club goods. These are much more likely to be provisioned in a pure market, since the benefits and costs fall more "correctly" on those who obtain or bear them.
>>A strictly rational agent will always free ride on a public good. And most of the time, most rational agents will not provide a public good, because the cost of paying for everyone else's consumption exceeds the benefits of their own consumption.
Thus the reason Free (As in Freedom) software advocates promote the use of Copy-Left Licensing, and why non-copyleft (like MIT, BSD, and others) are slowly eroding "open source" to less fully formed software and more just the tooling, libraries and dev environments used to create software
And it is already too late to change course back to GPL being widespread.
Every year there is a new project replacing GCC with LLVM, on Android the kernel is the only major GPL piece still standing (with Fuchsia on the horizon), on embedded there is a rise of non-copyleft OSes (including Zephyr, a Linux Foundation project), and so on.
The GPL creates neither excludability nor rivalrousness in the sense I am describing.
True in the general case of course, but there are arrangements that can be used to mitigate this. Crowdfunding with a threshold mechanism, ala Kickstarter, is one of these - reaching the crowdfunding threshold is a stable equilibrium. The intuition is that if every funder's contribution is critical to the threshold being reached, the funders are essentially "matching" each other's contribution, and thus providing a successful incentive even in a strictly rational sense.
If this software works well enough, what would be the competitive advantage of AWS?
What we gained in homogeneity for cloud-native software we're now losing in wide heterogeneity of Operator software.
Would anyone care to give a brief summary? It's the first I've heard of this.
Google didn't make the same mistake with K8S and it killed Pivotal's PCF.
Kubernetes was never really about Pivotal in any universe I can think of. It was about AWS.
Disclosure: I work at Pivotal.
We now await to see if compute will become a commodity or just priced that way to prevent it becoming so. Is any challenger bold enough to bite off more of the pie than they can chew?
I wonder how many companies are successful with open source but paid licence for commercial use.
Think of Office365, SAP Cloud Platform. It's benefical for a closed source vendor to switch to a service model, collect monthly recurrent fees and get free ultimate leverage against the customer.