How does this work in practice / where can one learn more about this?
How does this work in practice / where can one learn more about this?
It works like this:
User tells the backend, “I want to upload picture.jpeg!”
Backend tells the user, “Alright you have my permission but ONLY for that filename with that extension. Here’s a token, enjoy.”
User uses that signed token and pushes the file to your S3 bucket.
Here's how you do it in Phoenix. https://sergiotapia.me/phoenix-framework-uploading-to-amazon...
You can argue that the user can upload anything using the original api anyway. But in the original case you can do server-side validation before the upload is proxied. I am thinking stuff that are domain specific like only allowing videos that are 6 seconds long or something.
You can move the validation to the client but the client can be easily modified. An actual user might not do this but someone trying steal your storage space (for serving malware or something) might?
These signed urls also seem to expire based on time so you can potentially save the url and upload again later if you allow generous expiration. (again, not really something I see being a huge problem)
But I guess these aren't really serious issues compared to the cost savings. Am I missing other ways this can be exploited?
I am looking into the GCS version, not S3, if that matters: https://cloud.google.com/storage/docs/access-control/signed-...
Do you have links (or just keywords) to learn more? Will I need to add something like Cloud Pub/Sub to my stack? https://cloud.google.com/solutions/using-cloud-pub-sub-long-...
This is more complicated than I imagined so I am not sure the cost saving will still work out (factoring in development time and extra code maintenance cost).
You can use whatever queue you’re comfortable with so long as you can pipe the upload events from the bucket into it. The pattern I’m outlining is just a physical separation of buckets to make access control much harder to screw up.
Not 100% sure what they mean by _vendored_ here, but I'm guessing they make a request to one of their backends to generate the URL and return it to the client for use.
https://devcenter.heroku.com/articles/s3#file-uploads
One thing to keep in mind, users should be able to upload (to the specific signed URL), they should not be able to download from that location. Don't make the files users can upload publicly downloadable, otherwise you can be used to host malware. After the video/image is uploaded, you need to download and process it[1], then upload it to an S3 bucket that allows download (e.g, via CDN).
[1] Use caution when processing user content. It is best to process media in a sandbox that can protect you against exploits in the media processing libraries.
Multipart signed upload is much harder and requires signing every chunk.
Just google s3 signed upload there are a few tutorials from Amazon.