Our container platform is in production. It has GPUs. Here's an early look
blog.cloudflare.com
blog.cloudflare.com
This does still expose the host's kernel to a potentially malicious workload, right?
If so, could this be mitigated by (continuously) running a QEMU VM with GPUs passed through via VFIO, and running whatever Workers need within that VM?
The Debian ROCm Team faces similar challenge, we want to do CI [1] for our stack and all our dependent packages, but cannot rule out potentially hostile workloads. We spawn QEMU VMs per test (instead of the model described above) but that's because our tests must also be run against the relevant distribution's kernel and firmwares.
Incidentally, I've been monitoring the Firecracker VFIO GitHub issue linked in the article. Upstream does not have a use case for and thus no resources dedicated to implement this, but there's a community meeting [2] coming up in October to discuss the future of this feature request.
[1]: https://ci.rocm.debian.net
[2]: https://github.com/firecracker-microvm/firecracker/issues/11...
It does seem however that Firecracker + GPU support (or https://github.com/cloud-hypervisor/cloud-hypervisor) is most promising though.
It’s surprising that AWS doesn’t have a need for Lambda but with GPU’s to motivate them to bring GPU’s to firecracker.
The runtime may be memory safe, but I'm thinking of the GPU workloads which nvproxy seems to pass on to the device via the host's kernel. Say I find a security issue in the GPU's driver, and manage to exploit it with some malicious CUDA workload.
This is helpful in explaining why AWS hasn't been excited to ship this use case in firecracker.
Say you find a bug in the GPU driver that let's you execute arbitrary code as root. That still all happens within the VM. To attack the host, you'd still need to break out of the VM, and if the VM is unprivileged (which I assume it is), you'd next need gain privileges on the host.
There are other channels -- perhaps you can get the GPU to do something funky on PCI level, perhaps you can get the GPU to crash the host -- but VM isolation does add a solid layer of protection.
I’ve been thinking about QEMM or firecracker instead of just containers for a more robust solution. I have some time before anyone would ask me about GPU workloads, but do you think firecracker is on track to get there or would I be better off learning QEMM?
QEMU can work -- I say can, because it doesn't work with all GPUs. And with consumer GPUs, VFIO is generally not an officially supported use case. We got it working, but with lots of trial and error, and there are still some problematic corner cases.
EDIT: It looks like some people may have been using ghost.blog.cloudflare.com/rss because we used to use Ghost but the actual URL was/is blog.cloudflare.com/rss. We're setting up a redirect for anyone who was using the ghost. URL.
"Cloudflare serves the entire world — region: earth. Rather than asking developers to provision resources in specific regions, data centers and availability zones, we think “The Network is the Computer”. "
So they just use the term "location" instead of "region".
IBM mocked Sun with: "When they put the dot into dot-com, they forgot how they were going to connect the dots," after sassily rolling out Eclipse just to cast a dark shadow on Java. Badoom psssh!
https://www.itbusiness.ca/news/ibm-brings-on-demand-computin...
There really is a wide gulf between the services provided by the older cloud providers (AWS, Azure) and the newer ones (fly.io, CloudFlare etc).
AWS/Azure provide very leaky abstractions (VMs, VPCs) on top of very old and badly designed protocols/systems (IP, Windows, Linux) . That's fine for people who want to spend all their time janitoring VMs, operating systems, and networks but for developers who just want to write code that provides a service it's much better to be able to say to the cloud provider "Here's my code, you make sure it's running somewhere" and let the cloud provider deal with the headaches. Even the older providers' PaaS services have too many knobs to deal with (I don't want to think about putting a load balancer in front of ECS or whatever)
A lot of developers get frustrated at AWS or Azure because they want to deploy their hobby app on it and realize it’s too difficult dealing with stuff like IAM - it’s like trying to dig a small hole in your garden and someone suggests you go buy a Caterpillar Excavator, when all you needed was a hand trowel. The reason this persists is because AWS doesn’t target the hobby developer - it targets the massive enterprise that does need the customization and power it provides, despite the complexity. There are, thankfully, other companies that have come in to serve up cloud hand trowels.
There is no “one size fits all” cloud. There probably never will be. They’re all going to coexist for the foreseeable future.
10 years ago, no such superficial assessment would appear on first page.
This set of words bear little substance and engineering facts.
> AWS/Azure provide very leaky abstractions (VMs, VPCs) on top of very old and badly designed protocols/systems (IP, Windows, Linux) .
AWS cannot be made parallel, they themselves are 2 gens
AWS gen1
Azure gcp gen 2
Gen1 is on vm, ecs ebs s3, for web2 era
Gen2 is on cluster computing which was enable by vm
The then "leaky abstraction" is the mandated abstraction at the time
And GPUs today is about 70s's CPU
For example, you don't have any form of abstracted runtime on GPU, it's like running dos system
It's more leaky than 00s ' vm
It turns out we don’t need React Server Components after all. In the future we will just run the entire browser on the server.
I really need NVIDIA RTX 4000, 5000, A4000, or A6000 GPUs for their ray tracing capabilities.
Sadly I've been very limited in the cloud providers I can find that support them.
https://www.cloudflare.com/en-gb/press-releases/2023/cloudfl...
Paperspace (now DO), Vultr, Coreweave, Crusoe, should all have something with ray tracing.
We did try on the T4 and A10G but the raytracing failed even though those cards claim to support it.
We ended up on Paperspace for the time being but they depreciated their support for Windows so I've been looking for alternatives. Will check out the provides you mentioned. Thanks again.
- Licensing Windows is really, really annoying (and expensive), and BYOL is something people seem oddly reticent to do - Installing (and updating) NVIDIA GPU drivers on Windows is something that requires GUI access (at least the first time)
Paperspace was going to be my answer, but I guess DO didn't like those problems either! NVIDIA has an RTX Cloud, though I admit I'm struggling to find mention of it on their website, maybe something like: https://www.nvidia.com/en-us/data-center/free-trial-virtual-...
Containers on the edge with low cold starts, scalability, the same reliability as Workers, etc would be super cool. In part to avoid the lock in but also to be able to use other languages like Go (which Workers don't support natively).
I've hit a bunch of issues and limitations with Wrangler and Workers locally over the years.
Eg:
Personally I don't want to keep using JS in the server anymore. As more time passes I feel like TS is a hack compared to the elegance of something like Go.
I want to like CloudFlare over DO/AWS. I like their DevX focus too -- I could see issues if devs can't get into the abstractions though.
Any red flags folks would stake regarding CF? I know they are widely used but not sure where the gotchas are.
For headless browsers, the latency benefits of “container anywhere” seems high. For things like AI inference, running on the edge seems way less beneficial than running on the cheapest location possible which would be larger regional data centers.
The operational excellence required to have every successful Internet company manage deployments to a dozen regions just isn’t there. Most of us struggle with three, my last gig tried to do two, which isn’t economical because you always try to handle one region going dark which means you need at least 200% capacity, where 3 data centers only need 150 + ??%, and 4 need 133 + ??%. It has all of the consistency problems of n > 1 and few if any of the advantages.
We need more help from the CDNs of the world to run compute heavy operations at the edge. And if they choose to send them 10-20ms away to a beefier data center I think that’s probably fine. Just don’t make us have to have the sort of operational discipline that requires.
Good point about at the very least not exposing placement to customers. That is a definite win.
One that bit me was https://developers.cloudflare.com/cloudflare-for-platforms/c... - we found an alternative solution that didn't require upgrading to their enterprise plan (yet), but it was a pretty compelling reason to upgrade and if I was doing it again I'd probably choose upgrading over implementing our solution. On balance I'm not sure we actually saved money in the end, considering opportunity cost
Such as? See: https://www.cloudflare.com/trust-hub/compliance-resources/
"We'll keep you on the edge of your seat."
"Nice parade you got there. It sure would be a shame if somebody were to rain on it."
If I understand correctly, you will be running actual third party compute workloads/containers in hundreds of network interexchange locations.
Is that in line with what the people running these locations have in mind? Can you scale this? Aren't these locations often very power/cooling-constrained?
Is that true for China though?
They really have a great engineering team