Using a disk-based Redis clone to reduce AWS S3 bill
wakatime.com
wakatime.com
location / {
proxy_http_version 1.1;
proxy_set_header Connection '';
proxy_set_header .....
proxy_cache s3_zone;
proxy_pass https://s3
}
If you wanted to scale this to multiple proxies, without having each one be a duplicate, do a consistent hash of the URL and proxy it to the owner.We used to do this to cache the first chunk of videos. Encoded with `-movflags faststart` this typically ensured that the moov chunk was cached at an edge, which dramatically decreased the wait time of video playback (supporting arbitrary seeking of large/long video file)
Http proxies with etag seem like a good public-facing cache option, but our caching is internal to our DigitalOcean servers. Here's more info on our infra, which might help with our decision:
https://wakatime.com/blog/46-latency-of-digitalocean-spaces-...
Maybe this is something envoy or the newer breads do better.
(1) http://nginx.org/en/docs/http/ngx_http_proxy_module.html#pro...
In my experience, having to perform cache invalidation is usually a sign of design immaturity. Senior engineers have been through this trial before :-)
For our production needs, we use S3 as a database too, and Cloudflare's CDN as the caching layer. Cache invalidation is handled through versioning.
Your own HTTP reverse-proxy caching scheme, meanwhile, can be made durable, such that the cache is guaranteed to only re-fetch at explicit controlled intervals. In that sense, it can be the “canonical store”, replacing an object store — at least for the type of data that “expires.”
This provides a very nice pipeline: you can write “reporting” code in your backend, exposed on a regular HTTP route, that does some very expensive computations and then just streams them out as an HTTP response; and then you can put your HTTP reverse-proxy cache in front of that route. As long as the cache is durable, and the caching headers are set correctly, you’ll only actually have the reporting endpoint on the backend re-requested when the previous report expires; so you’ll never do a “redundant” re-computation. And yet you don’t need to write a single scrap of rate-limiting code in the backend itself to protect that endpoint from being used to DDoS your system. It’s inherently protected by the caching.
You get essentially the same semantics as if the backend itself was a worker running a scheduler that triggered the expensive computation and then pushed the result into an object store, which is then fronted by a CDN; but your backend doesn’t need to know anything about scheduling, or object stores, or any of that. It can be completely stateless, just doing some database/upstream-API queries in response to an HTTP request, building a response, and streaming it. It can be a Lambda, or a single non-framework PHP file, or whatever-you-like.
> Also, because CDNs try to be nearly-stateless, they don’t tend to be built with an architecture capable of fetching one “primary” copy of your canonical-store data and then mirroring it from there.
True, though some CDNs let you replicate data up to 25MB in size globally; others support bigger sizes depending on your spends with them.
> And yet you don’t need to write a single scrap of rate-limiting code in the backend itself to protect that endpoint from being used to DDoS your system. It’s inherently protected by the caching.
Systems in steady state aren't what cause extended outages. The recovery phase can also DDoS your systems. For example, when a cache goes cold, it may over-power an underscaled datastore [0].
> You get essentially the same semantics as if the backend itself was a worker running a scheduler that triggered the expensive computation and then pushed the result into an object store, which is then fronted by a CDN.
Agree, but running one's own high-availability infrastructure is hard for a small team. In some applications, even if scale isn't important, availability always mostly is.
> It can be a Lambda, or a single non-framework PHP file, or whatever-you-like.
Agree, a serverless function a CDN runs (Lambda@Edge, Workers, StackPath EdgeEngine etc) would accomplish a similar feat; and that's what we do for our data workloads that front S3 through Cloudflare Workers.
[0] https://sre.google/sre-book/addressing-cascading-failures/
One could accomplish a lot with a CDN and edge compute add-ons.
---
Basically:
cache validation: Check if a file matches a cached version
cache invalidation: Preemptively tell a cache to discard a cached file
---
For example:
Client A requests file X. Proxy caches file X.
File X changes in the upstream.
Client B requests file X. Proxy does not know the file changed upstream, so it sends a stale version.
---
In this case either the proxy needs to revalidate each request with upstream (which is expensive) or we need some way to tell the proxy that it should discard it's cache for file X since it has changed.
Nginx plus supports this but the free version does not: https://nginx.org/en/docs/http/ngx_http_proxy_module.html#pr...
https://book.varnish-software.com/4.0/chapters/Cache_Invalid...
I think I'll have to look into this to see if we can get away with document/blob cache in one of our solutions actually. I've resisted add-in Varnish in the mix, but it might be time to reconsider.
With your dedicated server, the latency is consistent, No API/network cost. Extra data can be tiered to S3.
SeaweedFS as a Key-Large-Value store https://github.com/chrislusf/seaweedfs/wiki/Filer-as-a-Key-L...
Cloud Tiering https://github.com/chrislusf/seaweedfs/wiki/Cloud-Tier
EDIT: Not necessarily YOU will have to clean it up at some point, but the next guy most likely will.
SSDB has also been wonderful. It's just as easy and powerful as Redis, without the RAM limitation. I've had zero issues with SSDB regarding maintenance and reliability, as long as you increase your ulimit [1] and run ssdb-cli compact periodically [2].
I've tried dual-writing in production to many databases including Cassandra, RethinkDB, CockroachDB, TimescaleDB, and more. So far, this setup IS the right solution for this problem.
[1] https://wakatime.com/blog/47-maximize-your-concurrent-web-se...
[2] ssdb-cli compact runs garbage collection and needs to be run every day/week/month depending on your writes load, or you'll eventually run out of disk space. Check the blog post for my crontab automating ssdb-cli compact.
For startups, moving fast and getting 80% done in 20% of the time is the correct choice. It's frustrating to look at the result as a software developer, but the engineer is there to serve the business, which pays the salary. As a software engineer, I very much disliked working in engineering teams that were engineering for the sake of engineering. It quickly devolves into meaningless arguments and drama.
S3 is fine though because there's any number of compatible alternatives, including self-hosted ones like MinIO.
What I think too few people understand about s3 is that costs on S3 compound. Every month you not only pay for data you stored in the current month but you also pay for the data stored in every month previously.
A related comment about our costs:
Asking because we use MinIO, and for us (a small online place with a single actually active MinIO server) it's been pretty good + super stable.
For ex: We use SQLAlchemy as the source of truth for our relational database schema, and we use Alembic to manage database schema changes in Python. We even shard our Postgres tables in Python [1]. It fits with my preference to keep this caching logic also in Python, by reading/writing directly to the cache and S3 from the Python app and workers.
[1] https://github.com/sqlalchemy/sqlalchemy/issues/6055#issueco...
edit: Found this https://docs.min.io/docs/minio-gateway-for-s3.html
Awesome! So minio can actually act as a caching proxy for S3. Very cool.
A one man shop with paying customers shouldn't spend time on optimization, unless it's just for fun.
If I'm a customer - I'd rather see more value added functionality, rather than cheaper data storage.
But that is my opinion, hope that your customers appreciate you spending time on this.
In fact, adding another component to optimize the cost is adding another point of failure.
A) you first need to run the service to get any customers
B) this might take long
C) you maybe don't want / can't get VC money at this stage
D) you maybe are not the most advanced dev who can properly utilitize s3 from I/O perspective, getting you to higher costs than possible
E) there might be a time period between introducing the service and getting traction which yields enough feedback, so you can start adding more business features
F) when you are burning your own money, you are more senstive to the cost side - which is not ultimately wrong
Just my 2c
Which ends up being back to - purely cost optimization of the initial cost optimization.
It seems like the whole product is focusing on cost optimization. I suspect that the OP would be better off making money by doing cost optimizations for third parties.
Splitting up the storage and execution between cloud providers was clearly an early cost optimization.
Then this cost optimization was to fix with the result of the previous cost optimization.
Adding more code that has no clear end value to your service* never makes anything more stable. Now OP has to maintain three things, instead of two. It's a classic "look at how smart I am" overengineering.
*- OPs customers definitely give 0 f's if everything is on AWS or split between AWS and DO.
or you could look at it as a tool to help the one-man shop make it through a tight cashflow situation that's coming down the line.
I'm curious if you've tested what it would cost to just host on EC2 (or something potentially even cheaper in AWS like ECS Fargate) with a savings plan. At a glance, it looks like AWS would be cheaper than DO if you can commit to 1-yr reserved instances.
That would seem like an easier (and possibly more effective) way to get around costly AWS outbound data costs compared to running a separate cache and sending data across the internet between cloud providers just to save $200/mo.
We went with SSDB because we already used Redis and SSDB, were very familiar with them, and had already worked out any pitfalls with using them. Wish we had found this sooner, thanks!
Off the top of my head, if I were trying to front AWS S3 from DigitalOcean, I'd have gone with MinIO and their AWS S3 gateway[1]. It appears to be purpose built for this exact kind of problem.
(Of course this is a concern only if you don't read Chinese.)
I always found it odd that we can easily port apps, databases, etc., from one cloud to another, but for file storage/CDN it's always some proprietary solution like S3. AFAIK open source solutions never really took off.
https://wakatime.com/blog/46-latency-of-digitalocean-spaces-...
I guess the point is that someone chose Redis and soon realized that RAM was not going to be the cheapest store for 500Gb of data. But that is not Redis' fault.
The reason we don't use SSDB as the main source of truth is because S3 provides replication and resilience.
Can you elaborate?
We've switched to SSDB, but with that we lost the Redis community.
Who says it's faster? This is supposed to make things cheaper.
Ages ago, I just used fsync() and called it good. After reading Dan Luu's article, I'm pretty sure I'd never figure it out on my own.
"Files are hard" https://danluu.com/file-consistency/
That's a bombshell that deserves the graphs to back it up, way hotter of a take than the article.
How is this possible? Surely reads from RAM are faster than disk?
even I am so surprised. I will try to run the benchmarks myself and see.
but... can anyone really explain how can a disk based solution be faster than RAM
Maybe if u have the technological knowledge, start your own block storage using Rook/Ceph object storage on DO. This will reduce your bill even further, and if u know what u r doing, u can improve the reliability
AWS S3 -> AWS EC2 bandwidth is FREE. To pay money to send the data to DO EC2 and then build a Redis cache to save a few dimes on EC2 that the D.O. Redis cluster now spends is...
just wow.
I'm not sure why name-calling and meanness attract upvotes the way they do—it seems to be an unfortunate bug in how upvoting systems interact with the brain—but HN members need to realize that posts like this exert a strong conditioning effect on the community.
It's always possible to rephrase a comment like this as, for example, a curious question—you just have to remember that maybe you're not 100% aware of every consideration that went into someone's work.