When AWS, Azure, or GCP Becomes the Competition
gkogan.co
gkogan.co
This in my opinion is the right way to solve the problem, i.e. provide the customers that want a managed offering the means to get a managed version directly from the entity most active behind the project, (e.g. Hashicorp) while getting the most from the cloud provider's infrastructure. I would be willing to trust Confluent or Hashicorp with operating/managing my Kafka or consul cluster but taking another dependency on their respective cloud offerings should be of concern.
I am quite surprised that its Azure that has shown creativity here while AWS and Google have limited themselves to offering marketplace AMIs/images at ridiculous markups.
They have also learned to build platform, unlike MacOS/Android, windows will never pull APIs out form under you.
Supporting historical baggage is why a lot of business people trust Windows.
Eg. You can still run VB 6 applications :p
We now have a situation where a well designed and complete 5 year old application will arbitrarily stop working because someone made an 'improvement' to the OS.
https://cloud.google.com/blog/products/open-source/bringing-...
Also DynamoDB (vs MongoDB)
> We’ve always seen our friends in the open-source community as equal collaborators, and not simply a resource to be mined. With that in mind, we’ll be offering managed services operated by these partners that are tightly integrated into Google Cloud Platform (GCP), providing a seamless user experience across management, billing and support. This makes it easier for our enterprise customers to build on open-source technologies, and it delivers on our commitment to continually support and grow these open-source communities
Are you sure that is accurate for AWS?
https://docs.aws.amazon.com/marketplace/latest/userguide/sof...
This seems to have launched in 2015: https://aws.amazon.com/blogs/apn/saas-partner-program/
Azure Managed Applications launched in 2017: https://azure.microsoft.com/en-us/blog/managed-applications-...
Google's offering seems to have launched this year.
> While AWS has ruffled a bunch of feathers with their Elasticsearch and Kafka managed offerings which can be easily construed as attempts to steam roll the respective open source-first entities, I am actually quite impressed by the mechanism that Azure has employed with Azure Managed Applications:
Here ( https://aws.amazon.com/marketplace/pp/B01N6YCISK?qid=1572333... ) is Elastic Co's ElasticSearch SaaS offering on AWS: > Elasticsearch Service on Elastic Cloud > Sold by:Elasticsearch Inc. > The official hosted Elasticsearch & Kibana offering on AWS. Launch, manage, monitor and secure Elasticsearch and Kibana deployments with the latest versions, and add machine learning and powerful hot-warm architecture with optimized templates.
The other popular product/software to hate on AWS about is MongoDB, which we find here: https://aws.amazon.com/marketplace/pp/B077D557RX?qid=1572333...
> MongoDB Atlas for AWS > Sold by:MongoDB > MongoDB Atlas delivers the world's leading database for modern applications as a fully automated cloud service with operational and security best practices built in. Easily deploy, operate, and scale MongoDB on AWS by letting Atlas take care of time-consuming administration tasks.
I haven't used any of these, but the reviews seem to show that people are running them.
So, your assertion that Azure has shown creativity, when AWS launched the ability for partners to offer SaaS solutions on AWS (with recent improvements such as PrivateLink to allow SaaS partners to offer managed services inside customer VPCs) 2 years before Azure indicates that seem to not be familiar enough with AWS to be making these statements.
I think it is a problem that it is very hard to monetize open source infrastructure libraries. In the past, you could try to do consulting or else offer paid hosting. Now with the cloud providers doing "zero management" hosting of open source libraries, both those avenues have been seriously curtailed.
Honestly, it seems like if you are the primary author of a popular open source library, your best bet is try to leverage that notoriety into getting hired by one of the cloud companies, and try to work on the open source library as a side project. Trying to use that open source library to support yourself any other way, seems destined for penury and misery.
For a single author of a popular open-source product or library, it's still very possible to earn a good living working on the thing that you built.
One author may steer or originate, but over time, Free and Open Source, is the work of many authors, many architectural components, and many network nodes.
Said another way, to understand the motivations of Free and Open Source, one must be able to see beyond one author, to the effects on systems, and systems of systems. Much of the benefits of both Free and Open Source, are not mainly to one author, but to systems, and over time.
This is important, as much of basic computing would not exist as it is now, except for sacrifice and insights by earlier authors and teams.
Of course that will make your project less popular.
Actually, Azure already has ML.NET, but it is about classical machine learning, and their own deep learning product, CNTK, lost the marketing fight with TensorFlow and PyTorch. Which is where I come in :)
Even Sagemaker and Fargate are mostly not useful to me as my company can easily operate its own k8s platform.
Open source tools for model building are great and a small team of engineers can write whatever odds & ends that aren’t covered, or make scalable implementations.
Notebook platforms like Databricks are similarly useless. A notebook is not a frontend to a cluster and notebooks are universally poor development environments even for the very tasks they are advertised.
The main thing that matters for ML development is how quickly & easily you can put an arbitrary container or VM into a deployment environment. You’re always going to need to rapidly change your container or VM, have common ancestors that get specialized to support GPU training or some web server application layer. Making the shortest possible distance between the definition of this environment and deployment of it is the whole game.
The type of tools I’m willing to pay for are specialized database engines, rapid data annotation tools, and specialized hardware environments.
I have no time for some “platform” tool for data provenance, notebook environments, model tracking, model-as-a-service APIs like Rekognition, or anything where I upload data and get back a model.
Any platform that advertises something like you don’t have to spend time defining your training environment, whether it’s Databricks or an out of the box deep learning VM on GCP, is a liability waiting to happen. You always need to define your own training environment, especially because you’ll almost always need specific (likely pinned) versions of all your dependencies, including system dependencies, to manage model training as part of a production life cycle. Very often you also need e.g. custom compiled TensorFlow, custom GPU settings or drivers, etc. It’s very foolish to base that environment on whatever comes out of the box.
Spark also is not universally useful. For training small models many times, like a workload that trains hundreds of small models all day (this was a production use case I had before that my company pilot tested with Databricks), the overhead of Py4J connector is insanely bad. It’s a really terrible paradigm for Python software, meanwhile Scala is a miserable ecosystem for production machine learning models. On top of all this, Spark MLib has huge gaps in functionality and whole classes of problems (e.g. large scale MCMC inference) are not solvable in a way in Spark that is seriously comparable to other tools like STAN & pymc running on not-Spark with simple multiprocessing.
Spark only serves a few special use cases that usually don’t justify its cost, similar to map reduce on Hadoop.
Databricks though (distinct from Spark) serves no use cases.
unnecessarily dismissive, grandstanding.. Bravado of the rejection, not supported at all
Rich-content Notebooks tell a story; well-worn Notebooks are a familiar context for experiment and viz.
They're incredibly non-portable; they don't play nice with version control; I know how they're supposed to be used, but I'm yet to see them actually used in a way that isn't a breeding ground for the worst software engineering I've seen.
Personally I would be happy with just proper version control support.
A notebook frontend however is less than useless, and is actively harmful by propagating poor notebook environments even further into aspects of computing where they cause harm and hurt reproducability and code factoring.
Given this, even if Databricks offered perfectly complete features for all aspects of cluster computing, it would still be inferior to just my own managed EC2 or EMR clusters or equivalent with other providers, where there is no “notebook as control plane” garbage.
But when you add to that the fact that Databricks lacks full features for me to totally own every customized detail of my cluster environment (e.g. how can I run plain Python multiprocessing tasks with zero Spark in Databricks? How can I bring my own custom defined GPU container with custom compiled Tensorflow?) it makes the deal even worse.
Databricks is just another Spark / Hadoop style snake oil seller banking on capturing a bunch of data science teams before people realize that it’s a conceptually junk way to work.
As for the other tooling you mention, I’d almost always say to build it in house. For example, I don’t know of a single A/B testing provider that actually uses frequentist sequential testing to correctly avoid early stopping bias. You actually need real statisticians hired in-house to solve these problems, and the engineering work to set up an extensible A/B test as a service internal tool is not bad. (I’ve built Bayesian A/B test frameworks with teams of 2-3 engineers in 3 different companies). It’s just not cost effective to outsource it on the false hopes of not needing to hire your own in-house statistics experts. Just pony up the dough and hire them.
My anecdotal story is when GCP announced support for running chromium out-of-the-box. I thought that that would be the end of browserless.io. I held back my knee-jerk reaction to shutdown, and after a few months I believe we only lost a single customer. There's probably been a few others that we've lost along the way since the announcement, but it didn't seem to stifle our growth in any significant way.
In any case it really depends on what your product or service is. Just wanted to chime-in and say that the writing isn't necessarily on the wall if this happens to you. Stay the course and offer them the much-needed competition.
Also, If I'm using your service, and a competitor announces a similar service, unless they provide something that's much much cheaper, I will probably not switch.
1. Offer a great user experience. Many cloud providers are just pumping products out the door with little to no concern for UX.
2. Have great personality. That could be through offering outstanding support, or through being the kind of company that people root for. (Honeycomb.io is a good example of this.)
While I think a bunch of nice people work there I have the feeling that the industry sees their opinions as quite controversial.
True UX—not UI, which often gets confused with UX—could matter only if it affects org-wide adoption or ease of use for non-technical users.
That is in fact how FUD got into the vernacular. They would make vague promises that you couldn't quite call outright lies about having a product coming in that market, and a year later they'd put out something tepid (thereby proving they were lying about their progress earlier), but with some bulk licensing deal.
By the time the MS product came out the competitor would already be experiencing reduced sales growth because people would play wait-and-see for the MS product.
There is nothing stopping any of these folks from pulling the same trick. Having a better product will not save you. A brilliant product might, but even that's not a given.
The source is an r/clojure flamewar which I'm not going to link.
Any others?
It's completely a failure of vendors to keep trying to sell software instead of competing on the services. The ones who have rolled out their own offerings are generating plenty of revenue, but the overall breadth, flexibility and quality of is still rather poor. I don't understand why these vendors are moving slowly here. They're the best at running their own software. Writing some blog posts isn't going to fix anything.
https://a16z.com/2019/10/04/commercializing-open-source/
"As software has eaten the world, open source is eating software.
Today, almost every major technology company, from Facebook to Google, is written on the backs of open source software. Increasingly, these companies are building their own open source projects as well – Airbnb, for example, has more than 30 open source projects, and Google more than 2000!
In the future, the virtuous cycle will continue. Technologically, AI, open source data, and block chain are some examples of emerging innovations. The next generation of business models may include ad-supported OSS, as when a large proprietary enterprise supports open source projects; data-driven revenue; and crypto tokens, which monetize blockchain.
I believe Open Source 3.0 will expand how we think of and define open source businesses. Open source will no longer be RedHat, Elastic, Databricks, and Cloudera; it will be – at least in part – Facebook, Airbnb, Google, and any other business that has open source as a key part of its stack. When we look at open source this way, then the renaissance underway may only be in its infancy. The market and possibilities for open source software are far greater than we have yet realized."
If you want no competition -- find a new market (and win it). Otherwise you have to be smarter and faster than existing players.
Decades ago (before Walmart, Target, etc), it was common to have many small specialized stores that did just 1 thing really well (ie. selling just books, selling just shoes, selling just office supplies). Over time, Walmart/Target came along and bundled all of those specialty stores into 1 giant store that sells everything, causing a lot of the small specialty stores to eventually fade away.
AWS/GCP/Azure is sort of like the "Walmart of B2B Internet Services" in that they are in direct competition with companies like MongoDB, Akamai, and the other examples from the article re: offering comparable managed services.
As a market matures, the market consolidates into a small number a big players. To survive, specialty shops need to do things that the big players won't or can't, which I think the article does a good job discussing.
> By the way, I write an article like this every month or so, covering lessons learned from growing B2B software startups. Get an email update when the next one is published
This is the closest I got to subscribing to a newsletter.
Almost got you, huh? :)
I've A/B tested various locations and the inline-embed is the best, by far. This is also why many news sites now include related stories in the middle of the article, not at the end.
Which is stupid because you still can’t reach to your audience if you just buy harvested addresses, so if it’s just people who subscribed to your newsletter, you’re just better off letting them subscribe to your RSS. Perhaps the quantifiability of newsletters vs RSS makes email win. If only RSS had embedded ways to count views and see what users looked at.
Many seem to be doing well: https://a16z.com/2019/10/04/commercializing-open-source/
Some just may not be worth their oversized valuations in the long run but made choices to be valued much larger than their actual worth.