Run Any JavaScript Function in the Cloud
github.com
github.com
I also remember someone suggested a ±/∓ signs for CoffeeScript at one point, and they justified it as "Well, I can type it using <insert fancy keyboard combination> on my Mac."
I know this is slightly off-topic, but I had to point it out because I feel like it stabbed me in the eye.
Anyway, I do like the idea of Lambdaws, kudos to the author.
Speaking of readability, I actually use a package for my editor that renders word "lambda" as λ. This way, I get to see code that's more pleasant to read and don't have to type in (and save) any greek characters.
There are various code snippets for that floating around, e.g. http://www.emacswiki.org/emacs/pretty-lambdada.el.
Like someone else mentioned you can use the compose-key, or if you use an editor like Emacs, there are commands to insert Unicode symbols. If you're a lisper, you might also have a macro for inserting those :)
https://github.com/mentum/lambdaws/blob/master/lib/LambdaHel...
It seems here is a discussion about the topic: http://perfectionkills.com/state-of-function-decompilation-i...
I'm unsure if you will run into any of the serialization issues mentioned by Kangax, as developers will be using the same ( if not very similar ) versions of v8.
While this might not be a problem for some use cases, I guess at least the documentation could discuss the topic such that the users of the library use the functionality with the restrictions in mind.
suppose:
var i = 0;
var localCaller = remoteBundler(function() {
// this will be run at server-side
i += 1; // oops .. "i" does not exists on server
});
localCaller();
and yes, this is could be easily spotted, but I've seen so many guys really stuck with so naive errors (eg mongodb's reduce function)https://github.com/mbostock/d3/blob/master/src/geo/area.js#L...
It's a pure pain for me to type these letters. How do you write them besides copying/pasting?
[1]: I used Ukelele on Mac but there's also Keyboard Layout Creator for Windows and manually editing xkb files on Linux
There's nothing mystical about this. The source code is 4 files, which I read. It sends your function to Lambda, and that's cool, and the syntax is really elegant. Enough that I could totally see throwing something together using this to solve some complex problems without really thinking about it.
I don't think this is all the way there, but I really like the idea of programming with APIs like this being as easy to use as language libraries.
Especially nice for spinning something up with zero overhead. Maybe not optimal for production apps. Maybe good if you're constrained on server resources but less so on budget. Maybe good if you're still on the Lambda free tier pricing.
Essentially, this can help offset the need for managing extra servers for those kinds of tasks.
I see this being most useful if you have a one-off analytic you need to write against some big data in S3 or RDS. For one-off scripts dealing with the raw AWS APIs is just useless overhead, and the expense of running the script will be negligible.
(a) measure each of the points of your service (b) deploy your code in an automated manner (c) deploy your monitoring in an automated manner (d) make sure your code is under supervision (e) setup alerting on the monitoring (f) scale up / down and within price constraints as needed (g) repeat this for all supporting services (queue, db, etc) (h) write your actual application code
The potential to handle certain classes of problems via SQS/SNS/S3 pipelines is pretty alluring. You still have to do configuration, but the bet is that the configuration necessary for the SQS/SNS/S3/Lambda pipeline is far lower than that necessary to setup random autoscaling Celery, Resque, or random JMS/AMQP system on top of Ubuntu with Chef/Puppet/whatever.
1. I agree that JMS sounds like a hassle but is that really necessary? I would think that you can batch process data on an EC2 instance, then pick it up in your local code directly using AWS APIs... not sure.
2. I am not so familiar with the Lambda system but I'm also not sure how it would scale db as necessary (item "g" in your list) thus overall processing time would still be bottlenecked by other resources (database IO, for example), no? I agree with your points but in all these cloud-compute scenarios I always wonder "Are we trying to reach a theoretical limit of fastest-possible computation, or just reach some reasonable saturation point close to the natural bottlenecks/throttles of our system integrations?".
3. Having been burned a few times now by over-optimizing when considering cloud I would probably now first consider just picking a slightly oversized EC2 instance and throwing some high-performing code onto it (Java, C++). Dynamic languages + auto-scalable resources (though I'm talking about web hosting in particular now) seems to drain clients wallets more than anything. At this point I'd actually recommend anyone with new web infrastructure to just buy a static instance and write optimized Java rather than trying their hand at auto-scaling Ruby/Python/Node. Do you notice a similar issue with your clients regarding code optimization vs. auto-scaling?
I think I remember this concept in Matlab back from when I did some research in grad school -- basically an instance of Matlab can be setup as a compute server, and the parallel processing functions of Matlab code on other computers can portion out work to it. This is the ideal model in my mind.
1. Write high-performance code in any language with some function that should be happening remotely in parallel.
2. Configure AWS to auto-provision the resources necessary.
3. Execute code, have it behave as if it is all running locally.
Really all these things can already be achieved with SOA, RMI, message queues, etc., the trick is just in making it transparent to the programmer so there is no deploy step involved. With the right spec it could even become platform agnostic (change small config file somewhere to target different cloud platform... would be nice to see a JSR about that in the near future!).
And then move it off EC2 onto dedicated hardware and you'll see another ~30x cost savings.
Running a permanent transcode cluster on EC2 would be rather insane. Hetzner rents you i7's for $50 per month, the EC2 equivalent (c3.8xlarge) costs $40 per day.
Yes you can cut EC2 cost with spot-instances, but at least in our case that would still have been significantly more expensive than just renting some scrap metal.
If you need cheap, disposable compute for semi-predictable loads then the Hetzner flea market (yes, they really have one!) is hard to beat on bogomips per dollar.
I was targeting my comment toward a startup that'd likely be building a product like the OP suggested, though, with maybe a dozen devs who are all generalists. Setting up a bunch of EC2 instances to pull videos out of SQS/S3 and run them through ffmpeg is something an ordinary full-stack developer can do in a morning, and scaling it up just involves clicking a button. Running and scaling a dedicated server farm reliably usually needs a dedicated ops person to keep it all working.
Now we're talking. If you can carve your application into small, independent tasks, Lambda can run thousands of them in parallel.
This could be cost-effective if you have a large amount of data stored in small chunks in S3, and you need to query it or transform it sporadically.
So instead of keeping terabytes of logs or transactions in Hadoop or Spark on hundreds of machines, keep the data in 100 MB chunks in S3. Then map-reduce with Lambda.
Set up one web server to track chunks and tasks, have each Lambda instance get tasks from the server. You could effectively spin up thousands of cores for a few minutes and only pay for the minutes you use.
It's suitable for flinging a map-reduce job in response to some event, but I wouldn't try to jam a map-reduce job into those constraints. I mean, sure, yeah, theoretically possible, but really the wrong way to do it. If you're doing a task that even takes a second or two in Lambda you're coming perilously close to being less than an order of magnitude from a hard cutoff, which isn't a great plan in the public cloud. You really ought to be targeting getting in and out of Lambda much faster than that, and anything that needs to be longer being a process triggered in another more permanent instance.
I can stream a 100 MB chunk from S3 and map it concurrently as it streams in 10 to 15 seconds. Sixty seconds is more than enough time to process a chunk.
The bigger issue is that during the preview, Lambda is limited to 25 concurrent functions.
If Amazon delivers a product where "the same code that works for one request a day also works for a thousand requests a second[1]," then you might be able to analyze hundreds of gigabytes of data in a few seconds, spin up no servers, and only pay for the few seconds that you use.
500gb = 5000 chunks of 100mb each.
1000 concurrent tasks each running 10 seconds could process 500gb in 50 seconds.
You would use 5000 Lambda requests out of your free monthly allotment of 1,000,000. You'd also consume 5000 * 0.1gb * 10 seconds = 5000 gb-sec of your free monthly allotment of 400,000.
S3 transfer is free within the same region, and S3 requests cost $0.004 per 10000 GETs, or $0.002 for this query.
Even after you exhaust the free Lambda allotment, processing 500gb would cost $0.000000208 * 100 * 5000 or about 10 cents.
Scaling this up, querying 10 terabytes would take about 20 minutes to execute, cost $2 for the query, and about $300 per month for storage.
For sporadic workloads it might be more responsive and much cheaper than spinning up a fleet of machines for Hadoop or Spark.
[1] http://www.allthingsdistributed.com/2014/11/aws-lambda.html
If anyone is interested in an alternate open-source microservice platform to AWS Lamba, try checking out http://hook.io http://github.com/bigcompany/hook.io
No hook.io users have requested this novel functionality that Lambdaws performs ( yet ), but I suspect we'll add it to hook.io in the future.
* https://medium.com/@AdamRNeary/developing-and-testing-amazon... * https://medium.com/@AdamRNeary/a-gulp-workflow-for-amazon-la...
https://github.com/mentum/lambdaws/blob/master/lib/LambdaHel...
Documentation on this option:
http://docs.aws.amazon.com/lambda/latest/dg/walkthrough-cust...
You can use things like app armor profiles for runtime sanboxing or google's nacl for statically verified binaries. Heck, even java bytecode would be more flexible than javascript.