Some time ago I worked with a bunch of people who staked their careers on moving everything to AWS and converting it to serverless architecture. Here is just one example of how fishy it was. They used Dynamo. During presentations to the outgroup, they constatly praised Dynamo as being flexible, scalable and fast. However, internally they constantly struggled with various limitations of the database, from limitations of indexing to not being able to get certain types of metadata.
My point is, the architecture they were designing was heavily bent around Dynamo's strength and weaknesses. Considering how extreme and peculiar those strength and weaknesses are, switching to another database would require heavy re-engineering of the rest of their architecture, much more so than, say, switching between Oracle, MS SQL and Postgress.
I’m no fan of DynamoDB and a multi certified AWS fan, but, I know it’s strengths and weaknesses and so should anyone else.
Here is a 5559 page book of error codes. Imagine having a physical copy of this on your desk.
https://docs.oracle.com/en/database/oracle/oracle-database/1...
This such the perfect description. I was working on a Django feature last year that had to work on all DBs and testing the oracle stuff was like entering a new reality lol.
" CLSRSC-00009: No value passed as OLR locations
Cause: No value was passed as OLR locations.
Action: None "
This happens all the time with databases like Dynamo. All of them feel amazing when you use them in the "intended" way, but they also have extreme limitations and those limitations are different for every product. Why? Because they aren't based on some fundamental data storage model, but rather on an aggregation of specific use cases, which are different across different products.
Here is a mild practical example: https://blog.codebarrel.io/why-we-switched-from-dynamodb-bac...
There are certain (limited) use cases where it is good. Every technology choice takes you down a certain road and limits your implementation choices in certain ways. It’s not the fault of the tool if they choose one that doesn’t meet their needs.
Also, the beauty of something like AWS, is that you can do “polyglot persistence”. You can have different services in your infrastructure use different types of storage depending on the use case.
You could have mapped your use case to this storage model and gotten ridiculously fast queries on vast mountains of data, which is largely the point of NoSQL. But this would have required data duplication, home-made management tools for dealing with said data duplication, and other work that was clearly better spent implementing the much simpler SQL solution.
A benefit of the SQL approach is that if your queries start getting bogged down from growth in data size and/or request volume (not entirely unexpected behavior when you have a couple joins), you can move to a different model where you treat the SQL db as a slow-access source of truth and periodically generate key/value into NoSQL or Redis/ES/etc for actually running queries.
I doubt any real world application comes anywhere close to doing that. Heck, I’d be mildly surprised if any real world enterprise, considering all their apps together, does that for either SQL Server or Oracle.
This seems like a simple case of people choosing the wrong tool for the job, without really understanding its limitations or capabilities.
> say, switching between Oracle, MS SQL and Postgress.
That's because those are all relational DBs that use SQL. If they chose some other NoSQL DB they would likely have encountered similar "peculiar" strengths and weaknesses. And I'm still not sure what this has to do with serverless - they could've have easily implemented a serverless architecture with MySQL or Postgress on RDS instead of Dynamo. But they chose not to.
[1] - https://docs.aws.amazon.com/lambda/latest/dg/running-lambda-...
[2] - https://blog.spotinst.com/2017/11/19/best-practices-serverle...
https://aws.amazon.com/about-aws/whats-new/2018/11/aurora-se...
Possibly because lambda is integrated with dynamodb. You can hook a function to an event like a row being inserted or modified: https://docs.aws.amazon.com/amazondynamodb/latest/developerg...
That lets you skip a lot of explicit queue management code and the change+notify is atomic by default.
The advantage of DynamoDB when using lambda is that it is API based instead of based on database connections and pooling and it doesn’t run inside of a VPC - meaning a network interface doesn’t need to be created inside your VPC every time a new lambda process comes online.
But now Aurora Serverless has the Data API that doesn’t use traditional database connections so there is even less of a need for DynamoDB for lambda performance reasons.
If they chose DynamoDB simply due to this strength (ignoring for the moment that the same thing can be achieved with native functions in Aurora), while ignoring the other things that ultimately turned out to be more important, then this again points to a flawed analysis. It's hard to see how serverless is to blame for this.
I have dealt with all of those, including MySQL. Switching is never easy. Every technology choice, even ones that should be similar, like an RDBMS, have strengths and weaknesses.
As far as being “locked in” because of serverless. With AWS the only thing you are doing code wise to make your code “serverless” is creating one function as an entry point that takes in two parameters - a JSON object and a lambda context. The only thing your handler should be doing is mapping the event to your domain objects and calling your back end code.
I have a c# .Net Core project that can quite easily be deployed as both a lambda that is triggered by an SQS message or as a Windows service. In fact, my CodeBuild step that triggered by a git commit compiles it as both a zip file that is deployed to lambda and a Windows executable - yes you can build a Windows .Net Core executable on Linux and vice versa.
People have a dream of making their implementation vendor neutral but wholesale migrations involving major infrastructure changes rarely happens. The risk of regressions is too high and the benefits too low.
Maintaining a second source is a good business strategy.
A well designed and properly isolated codebase is going to have limited technical debt in switching providers for various services.
But the cost of maintaining vendor neutral code and avoiding making use of the unique product advantages of the provider you've already chosen has significant downside.
It becomes a matter of designing for the lowest common denominator between providers instead of making full use of the ways in which your existing product choice excels.
If you are frequently changing providers, you have other issues afoot. If you rarely change providers, it's doubtful it'd be a net loss to design vendor-specific interactions with those services/products. Just add an intermediate API to isolate vendor specific code, which is going to be a good thing to do anyways for testing.
The vendor lock in argument is one of the few programming boogie-men that really gets under my skin.
The repo abstraction is a fantastic way of declaring your data access implementation and separating its concerns (especially limitations) from the service tier. Once declared - via interfaces - you can substitute the concrete implementation for something that suits the environment and the ever maturing agile use case.
Given how this is a fundamental scale/refactoring primitive, perhaps you haven't considered how others use it. I've used it time and time again at least 3 ways:
Monitor and consolidate data access patterns. If ITopSecret repo requires you pass an IUser to your verbs, you only need to audit/validate this layer to prevent unauthorised access. No one will replicate the query in IExportService and forget to check permissions again.
To scale read-heavy data, with new technologies like predictive search: Slowly pivot away from SQL sprocs/views to Solr/Elasticsearch by swapping out the concrete implementation ISomethingSearch repo with a new one.
For offline 'work' and unit tests: Create new concrete implementation of IBlobRepository that read/writes to your file system or an in-memory chunk.
> As a result, serverless computing today is at best a simple and powerful way to run embarrassingly parallel computations or harness proprietary services. At worst, it can be viewed as a cynical effort to lock users into those services and lock out innovation
Id be delighted to see a universally adopted standard specification for serverless computing including management interfaces and platform compatibility. The sort of specification that competing Cloud providers could implement thus better empowering consumers like myself. I have faith it will come around eventually all of this is relatively new
If the specifications for the serverless platform were "open" then the client could move their stack to different vendor with insignificant planning/risk.
An analogy would be vendors somehow changing machine virtualization such that your code was required to be very hypervisor-aware to function, and porting from vmware to xen imposed significant cost and risk.
The problem for the CNCF and its projects is you can almost always move quicker if subscribing to the walled gardens of the major vendors... even if the experience is less than optimal and definitely not portable.
Hopefully this changes over the next year as better abstractions (think Fn Project/Knative/Rook/someDB) and even some standards (cloudevents) emerge.
That said I suppose data could be traditional DB's (sql/nosql) managed by k8s... not sure about this yet.
1) Much serverless code will be vendor specific anyways. Anecdotally I'd say at least half of all serverless code I've written was to deal/work with vendor services.
2) You can easily write code to be platform independent with a thin wrapper for whichever serverless system you are currently using. This is pretty straightforward unless you are using vendor services, in which case you are locked in anyways due to that dependency.
[1] In fact, for FAAS start-ups (such as my own - PiCloud), the lock-in fear was one of the biggest hurdles.
Edit: to clarify, I’m referring to the assertion at the end of your comment, and I’m genuinely interested.
Everything has a cost. And so being able to move your code around at will is the cost of saving money.