AWS releases Glue Databrew, a visual ETL tool
aws.amazon.com
aws.amazon.com
What AWS has is a culture and institutional knowledge on how to launch new products that take foundational AWS services (S3, Lambda, EC2, DDB, etc.) and glues (!) them together better than what a competing non-AWS company can do. This is a bold claim (since AWS launches some very crappy products), but imagine being able to use AWS infrastructure at cost, having internal knowledge on how to best optimize that infrastructure and access to the engineers that own those services while you build abstractions and better user experiences on top of them.
I don't know how cos that compete in any related space can survive. When AWS is willing to throw whatever against a wall (launching 50+ services a year) to see what sticks, sooner or later they're going to land in your space.
Become more locked into AWS's foundational services -> these abstractions on top of them start to make more sense in engineering complexity / delivery time / possible cost dimensions -> Use more of these -> Become more locked into AWS's foundational services.
This feels very different from Azure or GCP.
Absolutely. I've seen this a handful of times with companies I consult, where they suddenly find themselves competing with AWS. I call it the November surprise because it happens around Re:Invent.
There are several reasons this is a tough thing to compete against, and AWS's vertical integration is just one of them. I've already written about them and also how to come out ahead if you find yourself in this situation: https://www.gkogan.co/blog/big-cloud/
This is true for a subset of products, but not uniformly. To the extent you're building an infrastructure product, you get to choose what axis to compete on. If you're going up against AWS, then trying to compete with them on things like cost and reliability are likely poor choices. But something like user/dev experience isn't. DynamoDB has a mongo compatible API and yet Mongo's Atlas hosted service is responsible for most of the company's growth over the past year. Why? Because it provides a unique offering, not just a 'good-enough' offering, which is what a lot of higher-level AWS services are.
I don't know if this is a strategic difference or an execution / cultural difference (AWS ships products faster, but they're almost barely usable in v1)
Sounds like the widely accepted "minimum viable product" approach.
I claim no special knowledge of AWS, but Azure is full apace, certainly faster than even us global SIs can keep up with in terms of providing support and capabilities.
Their execution is routinely abysmal but it never matters because they have two trump cards:
1. Backdoor through the purchase process bureaucracy
2. Network effects of existing services
Their competitive advantage is their captive customer base, which will much rather pay a premium to use an AWS-managed service than use another vendor.
Note mentioning accessing to road-map, strategic investment, genuine appreciation of product strength and weakness etc.
Former AWS engineer who launched a service here. That, access to source code and being able to setup an hour-long meeting with any engineer are the big points. Not that I think that lacking these is insurmountable, but they're very nice to have.
Which is that companies do not procure individual AWS services but rather AWS itself. Meaning that whenever AWS releases a new tool it is instantly approved and available for use across the company (baring internal processes e.g. security hardening).
Compare this with a startup which has to go through a 6 month long procurement process complete with vendor bake-offs in order to sell their similar tool.
If AWS continues to move into the application space they will surely dominate the enterprise because of this.
-- AWS advertises the tool via the console and preintegrates it from both directions
-- IT+Procurement already approved AWS for projects, so PMs can skip vendor/tool approval+onboarding dances and focus on the budget one
True of not just AWS but Azure + GCP too
Startups can compete, but gets into stuff like deep tech or cross-vendor integrations, where the visibility and integration advantages don't apply as much to the cloud vendors so they rather go after easier targets until they can't. (Folks here posted about UI, but for b2b, I disagree for most cases, unless there's something deeply technical about it that a 20 person team can't copy.)
- Historically, OSS seems to be free product dev for big cloud (... cue AWS's paid PR people to say otherwise ... ). Their integration, advertising, and procurement advantages makes it MUCH easier to win contracts before the OSS devs may even know it is being used and without a bid process. For a fraction of the effort and contribution, they are switching it to a model of monopoly channel owners vs content producers and and driving the sw margins to 0 on the content side. That's why anti-big-cloud LGPL-when-SaaS style licenses are emerging. There are always exceptions, but it's not the axis to compete on unless you do such a license..
- I agree about the community aspect, indirectly. If the software, in addition to being OSS, relies somehow on community and its steward -- not just source code -- and participation in it is somehow what's paying for the OSS dev, yes. For example, maybe the community is also a social network (Slack/Teams across orgs), or generating threat intel -- the software (post-scale) matters less post-scale, so forking is ok.
Just yesterday I selected SNS for a project instead of a local provider because AWS put it through the security audits we care about and we don't win prizes for spending time on these decisions.
Is that true though? IAM isn’t open, and things like service accounts and service chaining (a can only access b through c) are also not.
Yes. Compared to those, newly AWS services are more likely to work with, and integrate with existing services. However, the further you stray from 'Compute' the less likely this is to be the case. More 'esoteric' services tend to be their own microcosm and sometimes feel like they could have come from another company entirely (Quicksight? etc)
This is still light years ahead of Azure (and to a less extent GCP), where even compute services will not necessarily work with one another. You need to make sure the "SKU"s are compatible. Want to use some fancy storage? Oh no you need to use SKUs XYZ and premium this premium that. Whereas if AWS releases a new storage type (such as IO2), you can pretty much assume you can attach that to any of your existing instances (even if some particular types could be recommened).
Not to mention surprising behavior when you try to mix and match features. GCP and AWS, you have instances working perfectly fine, but you have discovered that they provide the ability to create 'internal' load balancers? Cool! Create one, point to the instances, or point to their respective automatically managed groups (ASGs or instance groups). It will be there in case you need it, your workloads are unaffected. Do that on Azure, and now your instances have no internet connectivity whatsoever, as all traffic is now routed through it. There are footguns everywhere.
Technically, GCP tends to be the most advanced of the bunch (their automatic instance migration is brilliant, meanwhile AWS keeps sending emails to us saying that some instance is degraded and it's our problem now). Their networking capabilities are impressive as well (first to have global anycast load balancers, Google's premium network, subnets spanning AZs, etc). However, they do seem to be too opinionated. Want proxy protocol on your NLBs, even though NLBs preserve source IP so in theory you don't need this(but with K8s ingress you might). AWS says sure, we have the feature, enable it, we don't care. Google says: why do you need proxy protocol, the source IP is there. These are not the headers you are looking for. Azure says: proxy protocol wat?
You can't use the SQL Server Virtual Machine extension on an Azure VM to extend the disks if the VM size is one of the AMD EPYC CPU types.
During the support call, the Microsoft tech shared a screenshot of the source code for the SQL VM extension, and it had a switch statement that decides if each feature is "supported" or not.
Let that sink in: Microsoft literally hard-codes their VM-size-to-feature lookups in probably thousands and thousands of places with huge switch statements full of code like this:
case "Standard_M416ms_v2": return false;
case "Standard_M416s_v2": return false;
case "Standard_M64ls": return true;
case "Standard_M64ms": return true;
This is their standard coding practice.So next time you try a new VM size or type, don't be surprised if things randomly don't work or "aren't supported" for mysterious reasons...
And with sales taking up such a huge percentage of a lot of these SAAS companies revenue, Amazon can pass the lack of sales to the customer as cost savings. Skip the sales process and the sales cost. win win.
If not, would you be willing to share where some of the big pain points were?
I eventually got it set up, but for the amount of effort involved, I could have just wrote my own custom ETL solution in less time.
The scheduling jobs and triggers is nice once set up. I do hope AWS makes improvements because Glue has potential, but it doesn't feel like it is ready for prime time yet.
Another example is that some date/time columns got brought in and crawled as a string. That's a bummer because obviously you want to do native operations of these, datediff, datepart, etc. without having to cast all over the place. We manually set them to timestamp and they work perfect (awesome!), even in Athena, so we thought the problem was solved.
However, once we did anything with those columns in Glue ETL, those fields got set as nulls.
The problem can be fixed of course, but these types of issues happen fairly often and they quietly fail (no errors, just values set to null).
Ended up going with an ETL as a service (Alooma and now transitioning to ETLWorks) to extract and load our data into Snowflake.
For this purpose I think it's fantastic. Write a PySpark script and press go (we have a package on pypi called etl_manager to facilitate this). It 'just works' for this use case, and there's a huge amount of value for us in not having to think at all about managing or configuring a Spark cluster.
Our biggest bugbear was slow job startup times and a lack of pip installs, but both of those are fixed with glue 2.0 which was released recently.
We don't use any of the visual/GUI based tools for our jobs, we just write our own Spark code and version control in Github. That's unlikely to change any time soon with products like Databrew. That said, the data profiling tool in Databrew does look like it could be useful as something to refer to when writing code.
(I realise this doesn't help with your specific issue, but i thought it was helpful to offer an example of a good experience)
Super simple incremental loading. This also works when loading data from a relational database, by storing the greatest primary key value.
This is spot on IMO. I use Glue internally (opinions are my own) and still believe that the best course of action is Glue should only run managed Spark. We provide an empty Scala "script" that does nothing, and load a compiled JAR file with the Scala code that actually runs our job as a library and have Glue exec into that.
We can version the ETL in git, run local tests outside of the Glue data plane, prototype in the Spark shell, and much more.
In general, AWS excels at a lot of core features (EC2, S3, the databases) but the higher-level services feel very thrown-together, with awful documentation. They are launching a large amount of half-baked services these days, (intending to capitalize on vendor lock-in?), but it's making the ecosystem start to look like a confusing mess. A lot of these wouldn't survive if they were marketed as standalone products.
But I’m glad the lack of investment is turning around with this, the recently released Glue Studio and the fantastic “glue 2 fast startup” job types.
Given the market power of cloud providers, every infrastructure innovation now a "sustaining innovation" in Christensen's terminology.
The key to success seems to be building a product for a niche the cloud providers think is too small, and then either maximizing your value within that niche so that if your market grows large enough for a cloud provider like AWS to come after you, you can pivot to providing customizations to your highest margin customers. MongoDB is a good example of this.
On the other hand, none of the major cloud providers seem capable of moving up the value chain to the application level, so if I were starting a company today I would focus on leveraging my infrastructure-level innovation to create a vertical opportunity in a high margin market instead of seeking to build a horizontal platform (IoT platform for example).
The benefit to visual ETL is that non-engineers can do a lot of basic data engineering. We tie this into our more complex code-based ETL pipelines. It was a game-changer for us and helps us get a lot more done.
As other posters commented Visual ETL often suffer from source control or limited extension ability but do provide the rapid development environment that users (generally more business oriented) seek. They also tend to trivialize the value of experience/discipline - for example I go to an accountant for my tax because they apply learned-experience relating to tax that I do not have (even though the math is easy) whereas in data engineering seemingly simple tasks such as correctly applying data typing to money or dealing with timezones seems to be glossed over in the pursuit of DIY - and wondering why your money columns don't reconcile or you lose data in failure scenarios.
At the other end of the continuum large teams writing bespoke ETL code for every job does not scale well for many reasons (https://reorchestrate.com/posts/code-doesnt-scale-for-etl/). I think the positive reaction to ideas like Data Mesh comes from the failures of these large, centralized teams which coincided with the Hadoop era.
Our solution has been to develop an open source (MIT) declarative framework (https://arc.tripl.ai/) that allows configuration driven ETL - mostly developed via a Jupyter Notebook environment (to allow rapid development and appeal to a larger audience) - whilst making most of the difficult tasks mentioned above easier. This has been in development for a few years now and continues to evolve. We value your feedback.
this tool requires ready mostly clean-ish data to work with. but the #1 problem in data engineering is lack of such data
And that works great—a great DSL + a visual GUI that edits that DSL. Because like you said you absolutely need VC.
How would you change control something like this?
I would really like a visual editor I can adjust with Regex functions and mappings that I can self host and iterate on.