GitHub Actions as a time-sharing supercomputer
blog.alexellis.io
blog.alexellis.io
For github actions caching was fairly easy to add.
They also cache the Docker images without you having to do anything but it really only works if you use their base images. With custom images the cache rate doesn't look great in my experience but I have no stats to back that up.
When I make an edit and rebuild on my computer, it takes a few seconds to rebuild because most things are cached.
In CI it takes 20 minutes on the MR branches of our repo. 30 minutes on master. Because a whole bunch of crap is downloaded from scratch and rebuilt from scratch.
Our infrastructure guys did set up something that is able to cache stuff that some of the repos use. But with the limited access I have I have not been able to figure out how to make use of it for the repo I work on. And they don’t have time to look into it for me either.
That’s my life lol
Still, on my local machine these jobs would take 5 mins so its not perfect. And as the build gets more complicated and more stages are added, the problem compounds since the initial download is the slowest part.
https://youtu.be/9qljpi5jiMQ (13:58 time stamp for the part that is especially relevant here)
And Microsoft has effectively zero incentive to fix this because it still technically works and they get to charge my employer for all that extra build agent time. Here's a few examples of people reporting similar issues and getting zero help:
https://developercommunity.visualstudio.com/t/devops-build-p...
https://developercommunity.visualstudio.com/t/nuget-restore-...
https://developercommunity.visualstudio.com/t/slow-performan...
Unless you’re booting build servers on demand on your infra, you aren’t paying more for more minutes. Just waiting :(
(Unlike GitHub, where they do charge per minute.)
For example:
The last run didn't clean up, or changed something global.
New runs don't work because some expected state is missing that had been set up on a previous run.
Blowing the agent away each time reduces the chance that something previously run can impact your current run.
This is at the cost of pulling from networks each time. You can reduce those issues by adding network caches for modules/dependencies.
This idea does leech free compute in an API-agnostic way though, if this could be tied together with a GCP free tier, AWS free instance, etc I wonder if you could cobble together enough free resources to run everything you wanted.
A self-hosted macOS runner will be more economical in the long-run, if you have a spot you can hook it up at; or, if you're fine doing things less than legally, you can use https://github.com/sickcodes/Docker-OSX.
But yeah, just drive-by-downloading MacOS to your Windows box it is probably not quite on the up and up.
So if you have 5 10 second actions that run when a PR is created, that 50 seconds of compute is charged as 5 minutes.
It can be secured, as gitpod does, but as I understand it is a PITA.
I always thought this was a hard limitation, but I deployed some self-hosted GHA runners in Kubernetes this week and to my surprise that setup came with an option to run the full docker daemon inside of a container - so apparently it is possible.
Rootless containers are a lot of work and do not support many scenarios that you're going to need.
MicroVMs are the same experience as GitHub, full system and Kernel, do what you will. Even launch a nested VM.
I haven't tried more than three levels but in theory more should work.
What is Docker in Docker?
Although running Docker inside Docker is generally not recommended, there are some legitimate use cases, such as development of Docker itself.
...If you are still convinced that you need Docker-in-Docker and not just access to a container's host Docker server, then read on.
This makes it pretty clear that it's a different copy of the docker daemon (which eg. allows you to test changes to docker itself) and specifically says it's different from "just access to a container's host Docker server".I remember taking a small university course on Ethereum, and getting an introduction smart contracts, trustless environments, and so on. We then heard about a couple of example projects, and were finally asked for our own ideas.
Now, after learning a bit more, I'm pretty sure none if the ideas presented either by the lecturers, me, or the other students really benefitted from the trustless environment, which is mostly what you'd use Ethereum for, and arguably what you pay (a lot) for during contract execution. Yet there were so many ideas about what could be done using smart contracts which were really cool projects on their own.
I think a big world computer with nodes that can perform calculations, react to user input, be called from other nodes, and exchange tokens and information, is somehow an incredibly natural abstraction that humans can work very well with. So, agreed.
The major thing about the API that I don't like is when you initiate a GHA run via the API, it does not give you an ID that you can use to track it. So if you initiate and it either never runs or fails for some reason prior to any code you put on the image, there is no good way to track that.
So it both justifies the mainframe, but if you look at why we all went mainframe in the end, it was to, ironically, connect many computers together.
There's also the overwhelming issue of power and control. Moving computing back to the mainframe allow totalistic control over computing by the service provider. This is a good way to make money. But is it good for the world? And what would the world look like if we had reliable fast trustable interconnected system, instead of cloud mainframes/data keeps?
I'm forgetting which books but some of the books about early computing talked about protests against computerization, against the mass data ingestion (probably among others What The Doormouse Said?). For a while the personal computer was a friendlier less scary mass-roll-out of computing, but this cloud era has not seen many viable alternatives to staying connected while keeping computing personal. RemoteStorage was early in, and Tim Berners Lee trying a seemingly very reasonable Solid idea seem like very reasonable takes, or going full p2p with data/hyper and that world: none of these have the inertia where others can follow suit. The problem is much harder, but I think it's more path dependence and perverse incentives, that breaking out will be found to be quite workable and good and validated, but there's gross inaction on finding the moral, open ecosystem, protocols & standards based alternatives for connecting ourselves together as we might.
Past that, vps can be had so cheap. If we have good software, the computing footprint ought to be tiny.
One real challenge is scale out. Ideally, for p2p to really work, some multi-tenant systems seem required, so we can effectively co-host. I loved the sandstorm model, but didn't actually use it, and I think there's further refinement possible.
Ideally imo, I could host like 10 apps, but if you want to use one, you spin up your own tenant instance. The lambda engine/serverless/FaaS thing wouldn't actually spin up new runtimes, it'd use the same FaaS instances, but be fed your tenant context when it ran and only be able to access your tenant stuff. That way as a host, I kind of know & can manage what runs, but you can have your own freedom of configuring your own instance to a large degree.
Then we need front ends that let you traffic steer and any cast in fancy ways, so you can host your own but it falls back to me, or you have 10 peers helping you host & you can weight between them.
Operationalizing what we have already & finding efficient wins to scale is kind of cart before the horse, since fediverse &Al are so new, but I think the deployment/management model to let us scale our footprint beyond ourselves is a crucial leap. And I think we are remarkably closer than we might think, that the jump into a bigger more holistic pattern is possible if we leverage the excellent serverless runtimes & operational tools that have recently merged integratively.
> Home computers are not unusual, anymore; the novelty is gone.
---
1. https://epicureandealmaker.blogspot.com/2013/02/table-of-con...
I sent this to some of my coworkers for a good laugh.
Super cool, I love this idea!
Less general than you, but I also had a notion to do kewl shit on GH actions, and I figured out how to turn it into a "personal ephemeral VPN-like" using BrowserBox.
Basically, you:
1. fork/generate the BrowserBox project into your own account
2. enable Actions and issues on the fork, and
3. open a new issue from the "Make VPN" issue template which triggers the action to create your remote browser and gives you a login link.
And voila! In a few minutes you get an up and running remote browser/private VPN that runs for 15 minutes or so (conservation, but you can tweak the value in the actions yaml!),
Even cooler, the action job will post comments guiding you through any setup steps you need, and then post the login link for your BrowserBox instance into the comments section of the issue you opened! Here's the action yaml that I use to do this^0
One hassle: I have still not avoided is the necessity of using ngrok for this. Sure, you can get around it using a tor hidden service which we also support, but that requires the user connecting via a tor browser.
Ngrok is required to create the tunnel from the BrowserBox running inside the GH Actions Runner, to the outside world.*
Seems like a lot of steps, right? So I added a "conversational" set of instructions auto posted by the runner as issue comments to guide you through it.
Check it out la! :)
https://github.com/BrowserBox/BrowserBox?tab=readme-ov-file#...
* Technically, this is probably possible by using mkcert on the IP address of the runner, and posting the rootCA.pem as an attachment to an issue comment. You then need to add it to your trust store and you're good to go, but ngrok is the easier way: sign up to ngrok, get your API key, add it to the Repository Secrets in settings and hit the "Make VPN" issue.
0: https://github.com/BrowserBox/BrowserBox/blob/boss/.github/w...
1. feature pipelines
2. batch inference pipelines
notebooks are certainly interactive and Colab's purpose is to allow often low-powered clients to offload ML training/inference to a GNU-enabled server, so kind-of like time-shared systems of the old?
Speaking of mainframes... I ran across this video on the latest workflow for COBOL / CICS development on your IBM mainframe. You have the old 3270 emulation way yes, but then we move into 2010 technology: a custom version of Eclipse IDE that submits the jobs for you, then we move into modern times with Visual Studio Code integration (using something called Zowe CLI):