> You're tying your own software into somebody else's all-encompassing and potentially hostile (to your business goals) framework instead of preferring the library way with clearly defined and exchangeable interface boundaries.
We experienced that recently, moving our ETL jobs away from a Hadoop cluster managed by a sister company that was having some issues (as in "all our data engineers resigned half way through a long overdue cluster upgrade, now nothing works")
Our data science team found Amazon Glue and fell in love with its ease and simplicity...
...and then spent three months trying to figure out why everything kept breaking. TL;DR - Glue's DynamicFrames come with a bunch of caveats and mousetraps, and a severe lack of documentation or source code.
Oh sure, you can read the Python source code for a DynamicFrame... ...source code that promptly delegates computations to a Java class you can't get the source code for.
After a few months of our experienced data engineers (myself included) saying "Just write a Spark job and run it in EMR", they finally did, and now everything's fine.
But the initial simplicity of Glue got them hooked, and they threw good money after bad on it. And as for those goddamned Glue crawlers...
TL;DR - things like Glue are great for early prototyping, but they tie you down to the AWS APIs and infrastructure. At least with EMR the Spark job you're running could be easily switched to another provider's infrastructure or your own infrastructure without any heavy refactoring.