What would make AWS even better
yehudacohen.substack.com
yehudacohen.substack.com
we even made an interface in code that runs a given task on a lambda (backed by a docker image) or batch (backed by the same docker image) depending on cpu/mem/time constraints. the tasks themselves also send a message when they're done so you can just subscribe to that vs. long polling.
My wet dream is "bidirectional IaC". Let me make changes using the GUI, commit to repo automatically.
And this model lets me be cloud agnostic for the most part - I run data workloads on gcp, dev/build workloads on linode, I've run bare metal in some places where I needed on-prem stuff. It's all just very much simpler than every cloud's flavor of doing everything slightly differently through different apis and tooling...
that's the only service i see OP did not mention
If you do need Spark, Databricks is likely a better option though :)
Same with Twilio. They do kind of compete with them, but not really.
Their managed airflow is insanely basically unusably expensive, I don't get that.
It would be great if the lambda could handle running long. Id probably even be fine if the duration was punitive in that the longer you run over X time it becomes progressively more expensive. This would create a disincentive for using the service wrongly but would allow for oddball tasks.
For something like that, I would use CodeBuild. In essence all CodeBuild is a method to run a list of bash commands in a Linux Docker container.
Standard disclaimer: I work at AWS in Professional Services.
You can run and rest CodeBuild locally. Just download the Docker image and run the shell script.
https://docs.aws.amazon.com/codebuild/latest/userguide/use-c...
A couple of weeks ago, I tried to deploy a lambda function that created Azure Subnets in python, and the Azure client was 265GB alone. My layer creation api call failed because of this.
Out of curiosity, why didn't you use an Azure docker image to back your lambda function?
https://aws.amazon.com/blogs/aws/new-for-aws-lambda-containe...
Lambda is a complex system, and holding those sockets for long times across many services could cause resource starvation issues. You've got load balancers, data plane, control plane, tenant vms, and a whole bunch of caches, and support services that all need to be ready to roll over the lifetime of the invocation.
And you have to consider the use case for draining and patching lambda pools. If someone is running a two hour function and you need to take down any server that's currently holding a thread or socket for it, you need to wait for the function to complete. You can't start a new load, so you are really inefficiently using resources until the function completes.
It feels things like the S3 API are design by committee. If you use tools like the cli you'll notice how clunky it is
Again AWS is great for small businesses just starting up that can make a lot of use of the free tier, and find the pay per use pricing attractive, and for massive corporations that only need one approved vendor that can service all their needs without going through the purchasing process again.
It's the middle where you will get squeezed and not get a cost effective value without a dedicated AWS guy or two.
some external event triggers the boot lambda.
1 minute schedule triggers the monitor lambda.
I've gone through a bunch of audits, and automated scans, and I constantly have to explain this shit, even to AWS Employees.
How it works with ALBs, which do support security groups:
You want to receive traffic on port :443, and allow it to be accessible to the world. You have EC2 instances, and they are listening on the VPC at port :1234
So, you create:
- ALB my_alb which listens on :443, and forwards traffic to tg_traffic
- Target group tg_traffic, which contains the EC2 instances and targets the EC2 instance with port 1234
- Security Group sg_alb, attached to my_alb with two rules:
- rule 1, inbound, from 0.0.0.0/0:443
- rule 2, outbound, to sg_servers:1234
- Security Group sg_servers, attached to the EC2 instances with one rule: - rule 1, inbound from sg_alb:1234
This makes everyone happy. The rules require that traffic from the internet has to go through the ALB.Now how it works on a NLB, with the same scenario:
You want to receive traffic on port :443, and allow it to be accessible to the world. You have EC2 instances, and they are listening on the VPC at port :1234
However, NLBs, as mentioned, don't support security groups.
So, you create:
- NLB my_nlb which listens on :443, and forwards traffic to tg_traffic
- Target group tg_traffic, which contains the EC2 instances and targets the EC2 instance with port 1234
- Security Group sg_servers, attached to the EC2 instances with one rule:
- rule 1, inbound from 0.0.0.0:1234 (not :443, because the NLB translates the port for you, but not the source ip)
...that's it.However, now every audit/automated scan of the EC2 instance & it's security group is going to see that you're listening on some random port, and allowing traffic from anywhere. This throws errors/alerts all the time. Even AWS's automated scans are throwing these alerts.
When it's an auditor you have to take the time to explain that, no, that's how NLBs work. For automated scans, you have to just ignore the warnings/errors constantly.
If your instance has no public IP associated, then at least only that port is exposed, and traffic does have to go through the NLB.
If for some reason the instance does have a public IP associated, then anyone who can reach the public IP can bypass your NLB.
If you could have a SG attached, then you could force the traffic to go via the NLB and not come direct to the instance.
Edit: in the first example the servers should be in a private subnet. And not have a public ip allocated. They would require ssh hopping via a bastion or a vpn.
If you dont believe me, try setting it up yourself.
As for instances being on a private VPC/non-public IPs thats deployment specific.
In any case, everything then complains about listening on strange ports with 0.0.0.0/0
But adding it to my list of things to learn and understand better.
yep, it just works differently from everything else.
Once you realise that, from a network traffic perspective, AWS wants you to pretend its just a magic straw delivering traffic from your source to your targets, and it's basically invisible to everything, then it started to click more for me.
That's why I want NLBs to support SGs. or at least a NLBv2 that does it.
‘ip’ mode is very different from instance mode and might not be what you want. Though, ‘instance’ mode is subtly bugged if you are using cross zone load balancing.
Only six weeks of paid parental leave?
I would absolutely be willing to pay more for AWS if I knew that amount was going to treating the poor folks who built it all better.
Seattle dev: 1st year -> 2 weeks, 2-4th year -> 3 weeks, 5th+ year -> 4 weeks.
California dev: 1st year -> 3 weeks, 2-4th year -> 4 weeks, 5th+ year -> 5 weeks.
15 PTO days. 6 personal days
Yes I work at AWS. I’m never on call and I haven’t worked for more than 40 hours unless I’m learning something new trying to figure out. I control my own calendar and I manage expectations for my projects.
I do work in ProServe though…
Besides, I purposefully put myself in a position that I wouldn’t have to relocate to a high cost of living area. I knew that Azure, AWS, and GCP had a Professional Services department that may require a lot of travel. But no relocation, no on call, etc.
Then a worldwide pandemic happened that reduced travel…
I used to work in one of the DB services and we used to get 20+ pages (sev2) every day. Due to insane amount of pages every day, we used to have daily on-call rotations.