Distributed Cloud Builds for Everyone
blog.nelhage.com
blog.nelhage.com
While you do want your build to be as fast as possible, the constraint is often throughput and not latency.
If the author is interested in standardizing the same, I'd suggest implementing the REAPI protocol (https://github.com/bazelbuild/remote-apis). It should be amenable to implementing on a Lambda-esque back-end, and is already standard amongst most tools doing Remote Execution (including Bazel! Bazel+llama could be fun). And equally, it's totally usable by a distcc-esque distribution tool (recc[1] is one example) - that's also what Android is doing before they finish migrating to Bazel ([2], sadly not yet oss'd).
The main interesting challenge I expect this project to hit is going to be worker-local caching: for compilation actions it's not too bad to skip assuming the compiler is built into the container environment, but if branching out into either hermetic toolchains or data-heavy action types (like linking), fetching all bytes to the ephemeral worker anew each time may prove to be prohibitive. On the other hand, that might be a nice transition point to switch to persistent workers: use a lambda backed solution for the scale-to-0 case, and switch execution stacks under the hood to something based on reused VMs when hitting sufficient scale that persistent executors start to win out.
(Disclaimer: I TL'd the creation of this API, and Google implementation of the same).
[1] https://gitlab.com/BuildGrid/recc
[2] https://opensource.googleblog.com/2020/11/welcome-android-op...
The difference is with Bazel you also gain a suite of solutions for large software projects like container creation, target visibility control, isolation, dependency fetching, and build graph analysis. Also with the larger community behind Bazel adopting it for a large project is getting easier over time.
The unfortunate reality is that open source developers have decided that they don't want the state of the art because "it's written in Java and is like a 50mb binary". Yes, I know there are certain things about Bazel that don't fit certain open source projects, like it may not be easy to bootstrap a build environment for OSes that want everything to be built from scratch, or want to support 10 different architectures, and I agree that Bazel may not make sense there. At the same time, if you see comments from both HN and on the internet, "it's in Java" is often how FOSS projects are making these decisions :(
I'll always be bitter about this. In particular I feel like Cargo was new enough that Bazel existed and if they had at least built a Bazel compatible tool in Rust (so use Starlark etc.) it would have been very much in line with the Rust ethos of safety and speed.
Fargate Spot pricing is $0.012144 (vCPU Hour) + $0.0013335 (GB hour) + $0.000111 (Ephermeral Storage hour). Pricing is per second with a 1-minute minimum. So for one hour of compiling on 1 vCPU, 1GB ram, 1GB storage, that's $0.0135885, to run any old Docker container.
Lambda cannot execute for longer than 900000 milliseconds, or 15 minutes. For 15 minutes at 1GB for 1 request, that's $0.02. For 4 requests (1 hour) that's $0.06.
This is incorrect. 48 cores, 96 threads.
vCPUs mostly refer to threads, not cores. There are some exceptions. More here...
https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/instance...
...and here (which is inexplicably missing the c5.24xlarge):
https://github.com/nelhage/llama
My question is how much upload bandwidth this needs. The article linked here mentions it in passing, but that's always been the pain point for things like this for me.
IMHO Llama looks so good that it makes me wonder if AWS is undercharging to attract developers.
"Recently, I blogged about building LLVM in 90 seconds [1] using AWS Lambda and my Llama project."
How much work do we have to do to get this "free" thing? What privacy do we give up? What if we're working on something that's part of a product, or otherwise proprietary?
The idea that everyone is doing it, so we should stop being sticks in the mud and just jump on the bandwagon, is harmful and ignorant.
"My data has never been stolen from the cloud" is about as ridiculous as the fortune quote, "As far as we know, our computer has never had an undetected error." How would you know until it's way too late?
Or, even better, you find out the data you've used in the cloud has been leaked on the Internet. Was it AWS? Was it crappy software on your laptop? Or was it a breach on your network? You've just made finding out that much harder, particularly because the cloud doesn't have logs that you can audit, nor real, understanding humans with whom you can correspond.
People who don't understand security can't really be faulted for eschewing the idea of using the cloud for everything, but people who know better really shouldn't be pushing these unhealthy ideas.
I know this is a straw man because I can't know what sort of infra you're imagining, but I suspect a lot of people would be more comfortable with their builds stored on S3 under the scrutiny of their security team rather than on their engineers' laptops.