AWS Lambda: a few years of advancement and we are back to stored procedures
it20.info
it20.info
This is just a product of bad design
A better approach would be to decouple all your business logic and provide an interface to it.
e.g.
MyBusinessLogic
doSomeBusiness
MyLambdaRequestHandler
- handleRequest
- call myBusinessLogic.doSomeBusiness
That way if you need to move away from AWS Lambda, you can simply remove MyLambdaRequestHandler and any associated unit tests and your application code is unaffected.If you can do them properly (i.e. source control everything, proper schema update scripts, etc) they can work ok. Just hope you marketing guys don't want to do any A/B testing of various algorithms.
Stored procedures, IMO, are a case where DBAs should be pushing their pain down to the software engineering teams. I think a lot of this delineation will disappear as cross-functional product teams become more integrated across the industry, however.
That's kinda like saying "I personal hate <Language> because it invites bad design. The code for them rarely makes it into source control". I've never worked on a project that used stored procedures where they were not only checked into source control but provided the full necessary context so any "fresh" dev box could grab them and test them. Half the projects I've worked on with stored procedures even had unit tests!
> Stored procedures, IMO, are a case where DBAs should be pushing their pain down to the software engineering teams.
DBAs are just specialized developers or dev-ops. There is no reason stored procedures shouldn't be written by whoever is writing the application itself.
Personally I loved sprocs! I can fine tune my DB access control so it's impossible to do anything but call a set of sprocs, the sprocs has zero dynamic code so no injection of commands are possible, each db user can access different sprocs and I have this nice separation of concerns where my DB layer is simply a dumb, black box to take in input and return output. The same for all of the other layers of the app.
Did you mean to say the opposite of this?
Ideally it would also allow you to restore to a previous point but anything that isn't purely schema related (like an UPDATE OR INSERT) probably needs to be done by hand.
On deploy, all the stored procedures will get updated from that directory. Rolling back would consist of deploying an earlier version of the repository with the procedures in it.
It's more minor / maybe just me, but I also like the fact that I can replace logic in a stored procedure without having to push out a whole new software build unless I change the inputs or outputs.
It's not ORMs vs Stored Procs unless your data access framework is garbage.
Many ORMs will double the DB<->app traffic when a single column of a single row is changed, even if only the primary key of an existing record and the NEW column value are the only things that need to be sent, and the front-end application is the only possible source for BOTH values. Sure, I'd need to check for consistency - is this record the same as when I first queried it? - but a trigger (don't laugh, that'd be a mistake) that locks a row and increments a "row_version" value on update could allow me to optimize the consistency check to a simple "SELECT id, row_version FROM my_table". I'm personally not aware of any ORM that allows for that kind of optimization, though I'm also certainly aware that I'm not an expert in all ORMs and all things related to RDBMSs.
ORMs have their place. Many solve common problems well. Just as many miss some optimizations that may be relevant to one person, but not to you.
But I limit the ways in which I use them. Only to enforce rules intended solely for the database, for example, like shadowing an entry or changeset to another table. I also sometimes use them to optimize the retrieval of some cursors where I know from testing that SQL Server's query plan will end up returning faster than a view or (shiver) an app-side aggregation.
In short, like everything else in our world, fellow hackers, there's a time and a place for everything, and it's (in my opinion) irresponsible to say that you know better than someone else with different domain experience than you.
Edit: to my parent post, I'm agreeing with you, btw. Just happened to be a good place to hang my reply.
Ah yes, SQL injection will be with us for many years yet.
Perhaps, rather, a product of an infelicitous programming language? You don't have to worry about extra function arguments in javascript.
It seems common in C, python, and similarly arity-concerned languages to just accept a pointer to a struct (dict, etc.) when coding a callback for an API like this. That way the callback can examine as much or as little of the struct as it wants, and much less needs to change in future when the callback needs more.
The callback-args struct is common, but has its own set of trade-offs too: it frequently becomes a "data whale", a dumping ground for everything that could happen with no indication in the source code of which parts a subsystem actually uses. Frequently, I've seen these replaced by a dependency injection framework as the software and team gets larger (which has its own set of problems, with it becoming difficult to holistically understand the data flow within the system).
Keyword args are popular in languages that support them, too. They solve the semantic problems of missing positional args, and are semantically identical to the option struct approach but with a bunch of syntactic sugar.
I can't even tell when these things are troll posts and when they are meant to be serious anymore.
Python actually has much better support for varargs than JavaScript used to have. See: https://docs.python.org/dev/tutorial/controlflow.html#keywor...
We're looking to add support for Microsoft's new Azure Functions and Google Cloud Functions, and this will be a matter of creating a single file for each to handle the input.
You should always abstract your dependencies, especially if it's a critical part of your infrastructure.
https://github.com/MitocGroup/deep-framework
From our experience, this is pretty hard, but not impossible!
It's pretty well established that Amazon's usual approach is to provide a "toolbox" of services that can used to build any number of app architecture permutations, without all the typical fluff and polish expected by large business customers. I for one, as an engineer, appreciate this model since it's much more accessible and lightweight.
Its a code evaluation platform, running containerized under the hood, and with a large markup by Amazon. Companies have offered this for years before AWS, but AWS is tooting the horn louder than people have before.
Is it useful? Yeah, sure. Is it revolutionary? Oh come on now.
Just like "the cloud" is just timeshare on someone else's computer, this is simply code management and execution abstracted a few more layers up.
EDIT: Sorry I'm not on the hype train folks.
The innovative aspect of AWS Lambda is the triggers hook into a variety of popular AWS services, e.g. S3, Kinesis etc, with very little setup required from the developer, along with an easy interface and setup process. It abstracts away from containers and operating systems.
Google and Microsoft have recently launched beta versions offering 'Lambda' like functionality. AWS has led on this one, I don't think you can argue against that.
https://azure.microsoft.com/en-us/documentation/articles/fun...
You've been able to do the same thing forever by just writing everything as (non-F/non-WS)CGI scripts, or as isolated PHP servlets. But nobody ever charged for those by tracking CPU time, rather than just imposing a monthly fee + caps.
Very true that a key benefit of AWS Lambda is the ability to hook into the internals, but to the article's point, that's a pretty significant level of lock-in. We recommend to our customers who want a similar level of functionality hook up the internal events to SNS, at which point a job can be triggered on our end.
We operate across any cloud, standardizing through Docker images as the unit of code. It's "serverless" to the developer in that the only configuration is setting the event triggers. Of course there's compute involved, but it's outside of the development lifecycle.
One way to avoid that is to use web frameworks based on AWS Lambda, such as Zappa - https://github.com/Miserlou/Zappa
It seems to me (from reading HN) like most people are avoiding using App Engine because Google tortured and killed their pet, err, I mean cancelled Google Reader three years ago.
(And I guess, the inference is that because Google killed a tiny unprofitable service once in the past, you can not realistically depend on them to continue to provide services they are actually putting real money into because it is of strategical interest. Yeah, that makes total brogrammer sense, let's go with that line of thinking.)
Besides that, when starting last summer to evaluate both services, I didn't find much that was in the benefit of AWS, compared to GC.
(Besides AWS China. GC does not currently operate in China. I really, really want them to fix this.)
Oh, and in case you wanted an explanation: it's called groupthink and echo chamber. In (the more frugal, because less VC) Europe we were all amazed at how much american startups would pay for hosting.
If it's 2010 you're used to the costs of a hosted or coloed system with tens of TB of NetApp storage starting your next project on AWS was compelling in a way discount providers was not.
I think a lot of the popularity advantage AWS has is a result of being first mover in many of these areas,
I remember checking out both in 2007, right when AppEngine first launched. EC2 was just Linux; you had a familiar virtual machine to work with, and you could port your existing code over with virtually no effort. AppEngine required learning a whole new API, which was completely non-portable and locked in to Google, and that was a complete non-starter.
Also, being a PaaS, there was a lot more to get right with AppEngine, and the API changed a lot over the years. I remember looking at it in 2007 and thinking "There is no possible way I can get work done with this." In 2009, it was "Well, it might work, but it's way too much hassle." In 2010, it was "Nope, still too hard." Finally in 2012, I tried it again and it was like "Woah, finally this stuff is usable, and it's actually pretty pleasant." But by then, the outside world had Heroku, it had Beanstalk, it had Parse & Firebase, and there's a huge ecosystem around AWS.
It's a pretty good example of path dependence. If you're just coming to cloud computing now and have never run your own software before, AppEngine is actually a fairly nice alternative. But in 2007-2009, when a lot of us were first checking it out, it was an overly-complicated, vendor-lock-in mess. And vendor lock-in mattered a lot more then; now many developers have just accepted that they're going to be locked into certain platforms, but back then, building on anything that didn't have a publicly-specified API was a huge mistake.
And if you use Flask/WSGI interfaces + Cloud SQL there isn't even a lock-in.
The only big issue I can see: no AppEngine in China. :/
At the moment I'm resigned to use app engine for the world-minus china, and some AWS EC2 inside China, running the same code plus a a lot more admin work.
The one area where App Engine falls a bit short is lack of support for some widely used libraries like Numpy. It would be nice if The Google would add support for those (support for some transcoding libraries would be nice as well). Even better would be an interface to TensorFlow.
Brian - the new "Flexible Environments" (aka Docker aka Managed VMs) give you Numpy, python3, native libraries and more. I took my latest project into production with it, and it's been pretty painless.
https://groups.google.com/forum/#!searchin/google-appengine/...
I worked with Lambda a lot, piping the JSON input into Go programs and I cannot be more happy with something.
I work with Go just like I am used to, testing, compiling, CI and everything and then, I have a shell script that deploys it to Lambda (uploads the zip).
For what I need, lambda is absolutely great!
> In traditional PaaS world the code is the indisputable protagonist (oh, damn, and you also happen to need a persistent data service to store those transactions BTW). > With Lambda the data is the indisputable protagonist (oh and you also happen to attach code to it to build some logic around data). > A few years of advancement and we are back to stored procedures.
I don't agree at all. Lambda is code. It's code that is invoked upon receiving certain kinds of data, but in no way is the Lambda code subservient to the data, or (unlike stored procedures) even colocated with it. Lambda still needs persistent data services to play with persistent data -- there's nothing magic about it.
> There also have been a lot of discussions as of late re the risk of being locked-in by abusing Serverless architectures (like Lambda). > There is some truth to it. This, however, isn’t due (too much) by how you write the code from a syntax perspective: while coding my Python program I noticed that when the function is run in the context of Lambda, the platform expects to pass a couple of parameters to the function (“event” and “context”). I had to tweak my original code to include those two inputs (even though I make no use of them in my program).
The author describes having to write a controller method. This is not surprising, considering he was trying to make a web service.
> So, IMO, the lock-in will not be a function of how different the syntax in your code will be Vs. running it on a platform you control (probably minimally different) but rather in how scattered and interleaved with other services your code will be (at scale).
This is the most cogent point in the article. API Gateway + AWS Lambda can be used to create micro-microservices. Serverless Framework tries to wrangle this potential complexity by allowing users to group related lambdas/endpoints as a whole, but there is still the opportunity to create a real rat's nest of logic if we're not careful.
> P.S. Yes, I know that it’s called “Serverless” but it doesn’t mean “there are no servers involved”. Are we really discussing this?
Yeah, obviously servers are involved. But the fact that I don't have to care about those servers nearly as much (in terms of maintenance or in terms of up-front cost -- both are big wins for most customers) is worth discussing.
Is it the right architecture for everything? Hell no. System architects are in higher demand than ever because there are so many freaking ways to build technology products these days, and it's their job to figure out what tech is right for the job.
But the critical point is that Lambda is distributed and massively scalable, and stored procedures weren't.
Remember that company, Sun? The one that invented Java? When it was still alive, it has an unofficial motto, "the net is the computer". Nobody understood it then; now we know.
All computing science achievements will now be reproduced in distributed environment. OS? Check (AWS/DCOS/Kubernetes). Filesystems? Check (IPFS). IPC? Check (REST/Websocket). Perhaps even "drivers" will be a new thing (for IoT devices).
They give you decent primitives: immutable versions, aliases, easy logging. But everything else you have to build yourself: You have to figure out a development loop, deploys, configuration management. There is nothing built-in to help coordinate lambda deploys across regions.
I expect that you'll see this built out more this year.
https://github.com/serverless/serverless solves (upon my initial read-through) every one of your requirements, including multi-region deploys.
Is this sentence was written in double speak or did I just have a stroke?
Corrected. Thanks.
"Is this sentence was written...", which is even more broken.
There's nothing fundamentally wrong with stored procedures, per se. What was wrong with SQL RDBMS stored procedures was that:
1. each DBMS had its own stored-procedure programming language—and so application frameworks that wanted to provide compatibility with the generic idea of a "relational database" couldn't really use them unless they had devs on their staff familiar with each-and-every DB†;
2. there was—and basically still is—no concept of an RDBMS stored-procedure "view" or "schema"—i.e., API versioning for stored procedures, where a client can request to communicate with the set of stored procedures it was compiled to support, rather than the single version the database is holding onto today;
3. one major RDBMS (MySQL) never supported stored procedures at all, so many devs learned an ossified set of "web development best-practices" without ever being exposed to the idea of stored procedures as an option.
All of these issues are fixable. #3 is just a historical artifact of MySQL's laziness; #1 is likewise an artifact of the proprietary, "enterprise lock-in" nature of the first instances of stored procedures (Microsoft's and Oracle's), evenly fracturing the ecosystem away from adopting either. Neither is likely to repeat.
#2 is more pernicious, and to this day seems ill-addressed.
One place I've seen at least an attempt to resolve it is in the design of Redis's Lua queries, where the "solution" is to refer to the stored procedures by content-hash, with the database having an always-possible error case that requires clients to be able to fall back to inserting the stored procedure again (thus necessitating that all clients track their own copies of any stored procedures they want to call.)
Such a solution could be ported to other RDBMSes; I could imagine Postgres, for example, having a "database view" concept††, where real databases only contain raw tables and indexes, and all the view definitions, triggers, stored procedures, constraints, and even typedefs are held in some record/spec/document that can be both manipulated as data, and connected to as a database. This is sort of equivalent to the CouchDB 'design document' concept.
---
† Sadly, it's really just a syntax problem. If you had several DBMSes that all had the same syntax but different extenional semantics, it'd be very easy to write a single code-generator into your application framework that would spit out appropriate code to take advantage of the extensions available. That's how regular ORM SQL-generation works, after all. But when you have disparate syntaxes, suddenly you need disparate code-generators, which get out of sync and lose features (or an LLVM-like intermediate-representation that you can do the semantic-optimization steps to before finally doing the codegen step, but I can't imagine that'd be cheap enough to slot into a webapp's hot loop.)
†† To go all the way with such a concept, the real 'data' of the database—the tables and the indexes—can become floating objects, not contained "in" anything or defined anywhere, merely existing because of a ref-count from various vDB schemas. You wouldn't explicitly define tables; instead, you'd define your views (relational projections) and then assert identity relationships between some of the columns of those projections, causing one "table" to exist holding the underlying data for both views. This is, AFAIK, what https://en.wikipedia.org/wiki/Dabble_DB was working toward.
- Depending on database, the dev tool environment can be extremely limited
- Depending on database, debugging can be a nightmare
- Difficult, if not impossible, to scale over more than one node
- Another version dependency problem added
- If you want to sell your software, or services depending on an application that uses stored procedures, you need to be very careful how you manage licenses.
If you need to include a bunch of jars, using Maven + the maven shade plugin (or assembly plugin) to generate a 'fat jar' is very simple.
In fact, their official documentation states exactly this http://docs.aws.amazon.com/lambda/latest/dg/java-create-jar-...
- was not clear what was the limit on the total size, I think it was supposed to be 50mb, but it was not exact.
- edit the lambda function code, and wait 10 min till it is uploaded by the eclipse plugin
- finally not being able to attach a debugger
On the plus side, once you had all up and running, I agree it worked nicely. I complain about the developer experience only.
I'll admit that on slow internet connections, uploading a JAR to deploy can be a bit slow though. We have a Jenkins CI build pipeline that automates this process, it builds the JAR (running unit tests etc), then uploads it to S3 and uses AWS Cloudformation to make the Lambda point to the new JAR in S3.
Since my function needs phantomjs, I embedded the binaries into the deployment package and by just doing that I topped up 36 MB of the 50 allowed. Transferring it from my local regular internet was a pain. Now it's nice, I push code to my dev branch, it gets picked up by the CI, tested, built and deployed. I get a notification in the IDE when the whole roundtrip is done, and with the CI in AWS it takes seconds instead of minutes. Without binaries you can still fit jackson, a few AWS clients, groovy runtime, guava and httpclient in a few megs, which is manageable.
I agree about the debugger, but authoring a function is trivial and unit tests can be written and run without considering Lambda. I miss tailing logs, CloudWatch is nice but it's far from realtime.