Managing Apt Repos in S3 Using Lambda
webscale.plumbing
webscale.plumbing
---
For the Apt transport, I recommend apt-boto-s3 (https://www.lucidchart.com/techblog/2016/06/13/apt-transport...) that I authored. Unlike apt-transport-s3, it
* Works with AWS v4 signatures (required in some regions)
* Supports If-Modified-Since caching.
* Uses pipelining.
* Uses standard AWS credential resolution, like ~/.aws/credentials or IAM roles
* Allows credentials to be specified per-repo (in the URL)
* Supports both path and virtual-host styles for S3 URLs
And it works with proxies if you use HTTPS_PROXY/HTTP_PROXY environment variables (though I haven't tested this).That does look really cool though. For proxies it would be nice if it supported the standard apt-get proxy declaration. Apt-transport-s3 basically does that and sets the proxy env vars.
You could probably get yours into the offical ubuntu repos if you wanted, though.
I'd like to point out a tool I'm using now for building and managing our Apt repos- Aptly https://www.aptly.info/
It's really nice and Just Works™. It even has S3 publishing built in. Oh, and you can build Apt repos on Debian, CentOS, MacOS X & FreeBSD (thanks Go!).
To solve for Yum & Apt I wrote a tool in Go that keeps all the metadata in a JSON file in S3. It's just a map of S3 buckets that are repos to a list S3 URLs for each RPM / Deb. Whenever that file gets updated, just like OP, I send an event from S3 to Lambda. Unlike OP, tho, I just have Lambda launch a task in ECS and then my repo builds (& S3 syncing) takes place in containers (1 Ubuntu container for Apt repos & 1 CentOS container for Yum).
(And in general I'm not a fan of trying to do a lot in lambda. I like it a lot more as a glue layer / event handler).
The post mentioned deb-s3 and dismissed it as more complicated to setup than this solution. While the lambda solution is neat, I'm not sure I would describe it as simple. For now I think I'll stick with deb-s3.
But I have to admit I'm a little confused. Is this mostly for large repositories? I ship some software on multiple platforms and part of my build process builds an apt repo (available over HTTP) as follows:
run_gpg_command dpkg-sig -g " --yes" \
--sign builder *.deb && \
rm -rf repo && \
mkdir repo && \
cp *.deb repo && \
cd repo && \
(dpkg-scanpackages . /dev/null > Packages) && \
(bzip2 -kf Packages) && \
(apt-ftparchive release . > Release) && \
(run_gpg_command gpg --yes -abs -o Release.gpg Release) && \
(gpg --export --armour > archive.key)
Then I sync it to S3 using aws s3 sync.What are the problems doing it this way?
It does however not address the issue of signing in this case. Would be interesting to see if someone can extend it using an additional bucket to store credentials with the necessary policies to only allow the Lambda function to get and use them.
I may be about to show my apt inexperience here, but I don't see any explicit callout of a signing step. Does apt handle signing only at the package level, such that you'd just sign when you build a single package and then upload that?
I've written a tool for managing Archlinux repos in S3 ( https://github.com/amylum/s3repo ), which I run in a container myself so that I can use my GPG key to sign the packages and the full repo metadata. I'd love to move to something like this, so I'd be able to let AWS handle running the lambda on my behalf, but I've yet to figure out a good way to do that with the signing structure necessary
I also run my Yum & Apt repo builds in containers and I have lambda launch those containers when S3 is updated...
In order to facilitate signing I have my GPG key on the container host machine and mount that directory as a volume in the container so the container has access to the key as well.
For initially provisioning the container server I deploy the GPG key as an encrypted blob and decrypt it on the host.
That said, it would mean fixing up my original key strategy. At the moment, I'm using a signing subkey of my personal GPG key, since it's just a personal repo and I control the VM/container where the stuff happens. By comparison, if I were dropping it into a lambda/s3, I'd probably be more motivated to split onto a fully separate key for the repo.
Now combining this tool with Jenkins, which triggers a build whenever I change the recipe would be all I need for now.