Truthfully, I don't think anyone would recommend their acquaintances to join Amazon right now.
That said, Amazon is actually winning the AI war. They're selling shovels (Bedrock) in the gold rush.
Truthfully, I don't think anyone would recommend their acquaintances to join Amazon right now.
That said, Amazon is actually winning the AI war. They're selling shovels (Bedrock) in the gold rush.
For senior in-demand talent you are not desperate, and really only desperate people go to work for AWS as they don’t have any better options at a company which respects their employees.
It seems that they just don't care about the high turnover.
AWS is falling behind even in their most traditional area: renting compute capacity.
For example, I can't easily run models that need GPUs without launching classic EC2 instances. Fargate or Lambda _still_ don't support GPUs. Sagemaker Serverless exists but has some weird limits (like 10GB limit on Docker images).
> For example, I can't easily run models that need GPUs without launching classic EC2 instances.
Yeah okay, but you can run most entreprise-level models via Bedrock.
Fargate and lambda are fundamentally very different from EC2/nitro under the hood, with a very different risk profile in terms of security. The reason you can't run GPU workloads on top of fargate and lambda is because exposing physical 3rd-party hardware to untrusted customer code dramatically increases the startup and shutdown costs (ie: validating that the hardware is still functional, healthy, and hasn't been tampered with in any way). That means scrubbing takes a long time and you can't handle capacity surges as easily as you can with paravirtualized traditional compute workloads.
There are a lot of business-minded non-technical people running AWS, some of which are sure to be loudly complaining about this horrible loss of revenue... which simply lets you know that when push comes to shove, the right voices are still winning inside AWS (eg: the voices that put security above everything else, where it belongs).
How?
> The reason you can't run GPU workloads on top of fargate and lambda is because exposing physical 3rd-party hardware to untrusted customer code dramatically increases the startup and shutdown costs
This is BS. Both NVidia and AMD offer virtualization extensions. And even without that, they can simply power-cycle the GPUs after switching tenants.
Moreover, Fargate is used for long-running tasks, and it definitely can run on a regular Nitro stack. They absolutely can provide GPUs for them, but it likely requires a lot of internal work across teams to make it happen. So it doesn't happen.
I worked at AWS, in a team responsible for EC2 instance launching. So I know how it all works internally :)
No? You can reset GPUs with regular PCI-e commands.
> You can't really enforce limits, either. Even if you're able to tolerate that and sell customers on it, the security side is worse
Welp. AWS is already a totally insecure trash, it seems: https://aws.amazon.com/ec2/instance-types/g6e/ Good to know.
Not having GPUs on Fargate/Lambda is, at this point, just a sign of corporate impotence. They can't marshal internal teams to work together, so all they can do is a wrapper/router for AI models that a student can vibe-code in a month.
We're doing AI models for aerial imagery analysis, so we need to train and host very custom code. Right now, we have to use third-parties for that because AWS is way more expensive than the competition (e.g. https://lambda.ai/pricing ), _and_ it's harder to use. And yes, we spoke with the sales reps about private pricing offers.
"AWS Lambda for model running" would be another nice service.
The things that competitors already provide.
And this is not a weird nonsense requirement. It's something that a lot of serious AI companies now need. And the AWS is totally dropping the ball.
> AWS now has to take responsibility for building an AMI with the latest driver, because the driver must always be newer than whatever toolkit is used inside the container.
They already do that for Bedrock, Sagemaker, and other AI apps.
I'm no expert, but I'm pretty sure this[0] is what RTO 5 is.
[0] https://www.phoenixcontact.com/en-pc/products/bolt-connectio...