Amazon S3 Object Lambda
aws.amazon.com
aws.amazon.com
It's never about technology. It's about cost. If AWS ever priced themselves wildly above Azure or GCP or ___, shit would hit the fan.
Outside of that, who cares if you're locked into one cloud or another? Features? Maybe if you're super picky about Power BI or SQL Server? It's never made any sense to me.
S3 Lambdas look like a solid addition to AWS.
AWS is for privledged companies playing with silicon valley monopoly money based on made up "evaluations", most people don't have that benefit.
More companies are moving to the cloud from managed colo than the inverse. For example, companies like Uber and Adobe has been moving from their own datacenters to AWS. Even dinosaurs like Disney have moved to the cloud.
Anyone who saves a significant amount of money running their own DCs vs uses the cloud either has significant security requirements or is the in the business of selling infrastructure (ex. Dropbox)
Everyone is waking up to the "total cost" of software development. Market forces are forcing people to realize that developer time is the most valuable commodity there is.
Companies don't go to the cloud because they have silicon valley money to burn. They go to the cloud because every minute of their developer time that is not spent on development operations and maintainance of code that a cloud provider offers instead, is a minute they can spend on the core UNIQUE business proposition of the company.
You can bemoan the skills lost in our industry, but that always happened at every layer of abstraction in every technical industry.
If you price right, then if you scale to the point you need it for ongoing costs, you'll have money to spend on the one-time cost of building an abstraction layer over any locked-in services and switching to that, enabling both a migration off and, potentially, its own resellable product.
Not every platform needs EC2 instances. Most can build on Lambdas, API Gateway, and S3.
Not sure what evidence supports this notion of billions and privilege. If anything, AWS has democratized entrepreneurial endeavors.
> If anything, AWS has democratized entrepreneurial endeavors.
More likely AWS simply made a lot of money on the race to the bottom (or top) with regards to an ever faster, fatter, fancier and shinier Web. (Now everyone needs a CDN, everyone needs many PoPs, because the market forces demand this. Of course the total costs are enormous, but due to the increased scale and decades of technological changes the unit costs have gone down.)
> Most can build on Lambdas, API Gateway, and S3.
Sure, and it's great if people are aware of what kind of workload to put where. And if they build in rate limits. Otherwise they will be surprised when they run into the classic gotchas of "how we burnt all of our seed money on AWS bills".
A simple colocated server nowadays can do almost a petabyte a month for less than $1000. Lookup how much that is on AWS.
Lambdas and s3 are cool until you get ddos'd and now owe way more money. Or, your software doesn't work with lambda, which is true of the overwhelming majority of services.
If you’re building something that a cloud provider might adopt or build their own, well, isn’t that risky regardless of the cloud provider element?
I know some companies avoid one provider or the other out of competitive concerns. That seems easily managed with contractual obligations.
Twitter has enabled some borderline illegal behavior, but the account holder has been the one at risk. I don’t think Twitter systematically allows illegal behavior like parler did.
https://aws.amazon.com/aup/: You may not use, or encourage, promote, facilitate or instruct others to use, the Services.. for any.. offensive use, or to transmit, store, display, distribute or otherwise make available content that is.. offensive.
What does offensive mean? Who defines it?
Does this sound libelous to you, and do you have evidence supporting this? Even AWS's response didn't allege Parler "systematically allowed illegal behavior".
If someone hosted an MP3 piracy site or torrent tracker or carding forum on AWS, it would likely terminated too, regardless of whether Amazon are ideologically against piracy.
If you're a big fish with powerfu rich users like Twitter, then vendor lock-in isn't a risk. If you're a small, inconsequential platform that can be smashed like a fly on the window, then you really need to care about such things.
Parler outright refused to delete posts calling for the mass killing of politicians, even when those posts were specifically pointed out to Parler by AWS. This is all detailed in court documents. AWS very patiently tried to prod Parler into following AWS's terms of service, but Parler refused. This is a very clear-cut matter and those suggesting otherwise are either trying to deceive or being misled.
Twitter also uses GCP, and Azure I’m sure. Their serving layer is not hosted by AWS though, and I’m reasonably certain AWS could not shut down Twitter immediately simply terminating their AWS account.
The idea was to avoid vendor lock and make it "easy" to change the database engine.
A project was on, perhaps 2 years ago, still followed this pattern, but it ended up also adopting somewhere between 10 - 15 AWS services and we never discussed an abstraction layer on top of it.
There are one or two efforts now to create an independent layer on top of cloud services. I have a hard time understanding how that will work outside of basic services.
I have seen this fail in dozens of ways
I’m designing a system right now and I’m specifically not allowing it to become too ingrained into the AWS ecosystem so that I can easy move it to GCP, Azure, or elsewhere (would that Digital Ocean offered better support) if/when Amazon gets too big and starts to take advantage of their size. That will happen. It’s only a matter of time.
It's not clear to me what the real benefits of Object Lambda provides. Indeed it could move your transformation code closer to its data, but the cost of doing that is the management and complexity of the lambdas. I am curious what a best use case example would be for this.
> Resizing and watermarking images on the fly using caller-specific details, such as the user who requested the object.
Now we need to either do rescale directly after storing or need to use another service (on EC2) to do the resize. With Object Lamda, we could do it on-the-fly as the user calls the image. This is operational a lot easier.
Advantage here is that clients can just continue to use their existing S3 interface yet your lambda is still invoked.
Edit: I guess you meant that cloudfront caches it for you automatically.
I start my data as .csv.gz but the first step is a CTAS to extract columns and convert to compressed parquet. This step basically costs the most but gives a 10x data size reduction to downstream steps.
Athena does not work at all if you perform large numbers of small indexed read queries, definitely use a traditional database for that.
I did never work with it, but there is Athena that allows to query S3 in something like SQL. I think there _might_ be some cases were S3 is a useful database, as one can use pre-signed URLs to let any client write to the database (S3). That‘s something I‘m not sure how to set up with a classic database and without any API in between.
IMHO, I think it‘s all about the same functionality with a simpler architecture.
With this we might simplify the pipeline because we can keep the same S3 interface.
I guess that for some systems that are heavily S3-based it will help adding new features without needing to rewrite the interface or duplicate the functionality.
It's nice to know that AWS achieved such lock-in with some customers.
The lock-in mechanism. Imo this isn’t more functional than other approaches, it’s a way to ensure you don’t switch services.
It’s basically doing the same thing as an SQL view. It just hides and abstracts away some complexity.
So you don't want your server waiting around for 15 minutes on an open request, and most systems won't let you do that anyway (browsers definitely won't). 60 seconds is about as long as you'd want to wait around for your object, and arguably even that is too long. If you need to do some sort of lengthy process to your objects then it should probably be done in some sort of asynchronous queue using a normal Lambda.
The brilliant part is that it works even if you don't use AWS for compute: they can start making money on per-request compute for S3 "only" customers.
This was useful to a lot of teams in the org but the downside was trying to lockdown certain fields like IP address to people who needed access to that data for support/operational reasons.
Amazon did come up with some IAM/grouping policy things but these were a bit clunky from what I recall. I wonder if something like this would have made it a bit more flexible to allow us to redact the IP address on the fly
EDIT: although I wonder about the performance implication of this for the "big data" use cases
Despite being familiar with both of those individually, the combination of the two words side by side switched me into an entirely different context that took several tries to escape from...
When you do some data crunching you can offload some work to lambda. Theoretically, bringing execution closer to data brings more efficiency. Though AWS Lamba was very expensive for CPU intensive steady workload. Few times than using EC2 is some use cases.
For larger objects, the WriteGetObjectResponse API supports chunked transfer encoding to implement a streaming data transfer.
Google Cloud Platform (GCP) is a bit of a joke. I really don't think it is all that competitive. Its trying to stay caught up and they do offer a lot of similar services. But their documentation is all over the place. Their service administration is a disaster. You can tell that each product is built by separate teams that only meet together once a quarter over Zoom. Its very disjointed. GCP is definitely trying to play catch-up. That isn't to say that you couldn't build a powerful platform ontop of GCP, many respectable companies are doing it. Discord is one that comes to mind that is (or at least was...) built entirely on GCP. I don't want to say it is shit. But its not as cohesive as Azure and AWS.
Basically AWS and Azure are the two core competitors. I really think that in practical (non-niche, using core services) they are neck and neck. GCP is in the race, but definitely lagging behind the others.
* BigQuery
* Tighter k8s integration with GKE
* Single message service (PubSub)
Unless BigQuery is your first class citizen, I would avoid GCP.
I would only consider them if I need a massive scale of something they are cheaper on, that can quickly be migrated elsewhere.
If you sign a service contract with Google (which you will if you are spending any amount of serious money), then you will have a precisely defined guarantee that they will continue operating GCloud.
Could also be a way to mess with people -- for example, you could have an image src that gives different results depending on the recipient.