AWS services to avoid
medium.com
medium.com
Of course I disagree with that statement. Many serious Redis use cases at big companies assume Redis is there and will hold your data. What they usually assume is that under certain cases during failovers or other serious failures you many lose some acknowledged writes that were sent immediately before. And this most of the times does not happen too. Note that many high performance applications using SQL DBs with a relaxed fsync configuration will have the same assumption. The suggestion should be more: if Redis for you is just a volatile aid, use a given setup, if it's were you store data, use another setup.
Anyway if you want to see a version of Redis where you can store also bank accounts, transactions or any other of the most critical stuff you can store in a DB make sure to check the news at Redis Conf 2020 (the conf is in streaming and free).
EDIT: But even after that announcement the biggest value for Redis is that it can be fast and provide a "best effort" consistency level that is adequate for a lot of use cases in practice. Many systems were sacrified in the temple of wanting to provide linearizability for all the use cases.
> And this most of the times does not happen too.
This rubs me very very wrong. There is no maybe in data integrity. Redis wasn’t designed to be ACID and shouldn’t be treated as such just because “usually” it doesn’t lose data “under the right circumstances.”
If a sysadmin uses an RDBMS with relaxed fsync, they know what they are getting (and durable writes is not one of them, no usually about it). Same for Redis.
Removing fsync is generally referred to as "running with scissors" mode, although postgres now supports it better with unlogged tables.
I came up with a "mental framework" that allowed me to get beyond these details early when I started architecting systems for companies.
1) Identify what are the possible malfunctions
2) What mechanisms will I use to detect these malfunctions as early as possible
3) Do I have a clear and concise recovery path for said malfunctions.
This frame of thought has taken me far (in terms of being able to accomplish a lot in my career so far) and allowed me to move forward without (a) extensive, time-consuming, research into the specific tools I'm using into the specific situations which may never happen, (b) extensive costs hiring experts in these niche areas.
Sometimes you don't need ACID guarantees, which is part of the reason databases like Redis exist, and is why the compromises you're describing are important. But in this context it sounds like the framework you're describing would lead to rolling your own ACID for Redis instead of just using an ACID database to begin with.
Redis was never designed for strong Synchronous replication of transaction or Strong leader election.
So each time a master crash you have a high chance of some recent acknowledged write being lost.
we live in the age of devOps. there's no such thing as sysadmin, or god forbid, DBAs.
/sarcasm
Favorite datastore to date. The only thing I wish for frequently is native global change-data-capture without doing something manually with streams (I've thought about writing a replication protocol consumer, but the way resync's are handled is a little problematic)
Worker is a DIY stream consumer implementation. The stream -> lambda integration works, but its super limited (and ~300ms slower to get new events on average)
Via the lambda integration you should probably expect ~1000ms p99 latency intra-region, and ~2100ms p99 from, e.g, east-1 to west-2. Shave maybe 200-300ms off of those numbers if you do a DIY stream consumer without the aws provided lambda connector.
> "Remember redis (in AWS) holds ephemeral data in memory"
(because EC2 machines don't persist their disk data between start/stops and a lot of people don't realize that until it is too late - hey remember that Bitcoin exchange that got bitten by this?)
EBS is now the default when creating an instance, and instance types of newer generations generally don't have instance store at all (except disk optimized instance types, and those with "d" in the name.)
Yes if you do an ordinary OS reboot the disk remains.
Bad idea. Bad tool for the job.
He sort of implies here, not for you, but it might create that impression on others, was clarifying for those.
I understand your defensiveness, but please understand that because of your massive experience with Redis, either you have never lost data with Redis, which means that nobody besides you understands how to run your software in production, or you have lost data with Redis, which means that it's somewhat hypocritical of you to insist that Redis is durable.
Hogwash and there are plenty of examples out there that prove otherwise.
> somewhat hypocritical of you to insist that Redis is durable
If Jepsen has demonstrated anything, it's that no database is as durable as it claims to be.
I can't think of a single database that's solved the fsync/O_DIRECT issue on Linux completely. Postgres had to patch it _again_ last year.
They always have bugs somewhere, but there are huge differences between bugs that show up for very specific, niche cases, and normal "I wrote to the db and it dropped it".
"We found five safety issues in version 1.1.1—some known to Dgraph already—including reads observing transient null values, logical state corruption, and the loss of large windows of acknowledged inserts."
Loss of large windows of acknowledged inserts. Durability is hard.
As staticassertion is mentioning, some of the violations that were found were only around tablet moves, which happen only in certain cluster sizes and quite infrequently. Of course, Jepsen triggers those moves left-right-and-center to evoke some of those failure conditions; but that's not how tablet moves are supposed to work in real world conditions. This is different from other edge cases like process crashes, or machine failures, network partitions, clock skews, etc., which can and do happen. In those cases, Jepsen didn't find any violations.
We were planning to look into those tablet move issues and get them fixed up (shouldn't be that hard), but honestly, the chances of our users encountering them is so low that we de-prioritized that work over some of the other launches that we are doing.
But, we'll fix those up in the next few months, once we have more bandwidth.
"All of the issues we found had to do with tablet migrations"
"ndeed, the work Dgraph has undertaken in the last 18 months has dramatically improved safety. In 1.0.2, Jepsen tests routinely observed safety issues even in healthy clusters. In 1.1.1, tests with healthy clusters, clock skew, process kills, and network partitions all passed. Only tablet moves appeared susceptible to safety problems."
No one is here to claim that anyone is getting through any kind of rigorous testing without bugs found. But there is a huge difference between "My extremely common write path + a partition = dropped transactional writes" and "Under very specific circumstances, with worst case testing, multiple partitions, and the db in a specific state, we drop writes".
There is an ocean between, say, mongodb's test results, and Dgraph's.
Read Redis's evaluation, for example: https://aphyr.com/posts/283-call-me-maybe-redis
"If you use Redis as a queue, it can drop enqueued items. However, it can also re-enqueue items which were removed. "
"f you use Redis as a database, be prepared for clients to disagree about the state of the system. Batch operations will still be atomic (I think), but you’ll have no inter-write linearizability, which almost all applications implicitly rely on."
"Because Redis does not have a consensus protocol for writes, it can’t be CP. Because it relies on quorums to promote secondaries, it can’t be AP. What it can be is fast, and that’s an excellent property for a weakly consistent best-effort service, like a cache."
Again, Redis is a very different type of database, so expectations should be aligned. Further, this test is quite old.
But that's a huge difference from DGraph's results.
Basically, saying "Well no one does well on Jepsen" isn't really true. Lots of databases do well, but you have to adjust your definition of "do well".
Redis is an extremely reliable service. I've never "lost" data with Redis.
I don't want to pick apart the entire post, but I will say that the Lambda + API Gateway example is maybe the best example of jumping into the cloud-native world with blinders on. Just about every modern programming language has a toolset right now that will do the work of generating a CF template for you that creates an API gateway that forwards all HTTP routes to a single Lambda, and then that Lambda is responsible for handling the actual routing of that request. Examples include Zappa, ClaudiaJS, and Ruby on Jets just to name a few. I can't imagine providing a feature-rich web application in Lambda without this kind of abstraction.
Not knowing that such tooling exists, or explicitly choosing not to use such tooling, and experiencing pain as a result seems to be common theme in this article. If you dive into architecting a system using AWS's product offerings without understanding the tradeoffs you're making, you will experience greater cost and greater friction -- hands down.
This reminded me of a caveat about using some of these abstractions – which is that they are still subject to the limits and restrictions of the underlying platform.
We discovered this the hard way once when an automatically generated function name or something was over the limit in prod (this issue [1] describes a similar problem). We did not catch this in dev because "dev" is one character under "prod" and our autogenerated name in dev hadn't put us over the limit. That was an interesting exercise in leaky abstractions.
I also felt a similar tinge when the post talks about KCL. Yes, you use it, and yes there are rules about output streams - but they all make sense when deployed in fairly complex environments. I can’t remember when I last used stdout for logging outside of greenfield development. Once it’s on prod everything goes to stderr and is highly structured.
Similar sentiments about Cognito. At some point of complexity you must have a server in between using the Admin* range of commands. If you want their “get going quickly” then yes, there are trade offs like WebViews. That said; I’ve never used Auth0 - so insert a Luddite warning here.
Finally, one hundred percent agree on CF.
There’s also truly no reason one lambda function can’t have a collection of APIs (using the proxy method from API gateway). It’s literally no different from any other micro service setup. You can pick any abstraction for a function to cover.
This is what seems to happen when working with Lambda for me.
From the article:
In Lambda proxy integration, when a client submits an API request, API Gateway passes to the integrated Lambda function the raw request as-is.
...and:
You can set up a Lambda proxy integration for any API method. But a Lambda proxy integration is more potent when it is configured for an API method involving a generic proxy resource. The generic proxy resource can be denoted by a special templated path variable of {proxy+}, the catch-all ANY method placeholder, or both.
Which I think is one of the underlying points of the article.
> Many a startup has fallen prey to ElastiCache. It usually happens when the team is under-staffed and rushing for a deadline and they type in Redis into the AWS console:
I think this is a good article in the sense of "these are things that may or will bite you if you've not put a bit of thought into them".
Even the replies on this story upstream hint at that, with people debating Redis' durability, based on this:
> It’s best to design your system assuming that redis may or may not lose whatever is inside.
completely ignoring the follow up:
> Clustering or HA via Sentinel, can all come later when you know you need it!
It's like the first sentence was read and people leaped to the keyboard to reply passionately.
[0]: https://aws.amazon.com/blogs/compute/announcing-http-apis-fo...
So you say Kinesis is bad for message streams where each message should be handled by only one machine? Well duh, that's how it works. Kinesis is meant for every listener to be assured to get the entire stream. It was never designed for that, and I expect you'll have a bad time if you try to use it that way. SQS is obviously what you want if you're interested in guarantees that each message is handled once and only once.
One Lambda per route in a decently-sized restful webapp is ugly and unmanageable? Well duh. Don't do that. Do pretty much anything else instead. Seriously, anything else at all.
This sounds more like, if you want to deploy services on AWS, do either pay somebody who knows what they're doing, or spend some time checking out the services to make sure they are designed for what you want. Don't just pick a random AWS service and usage pattern from a Google search and start building around it without ever checking if it matches what you're trying to do.
You mean SQS FIFO? SQS classic cannot guarantee exactly-once delivery or ingestion, either due to failure at the client or server. SQS FIFO, however; can but requires both the client and the server to work in-tandem to ensure exactly-once processing.
How’s amazon’a kafka offering going?
It’s also what every lambda framework (serverless etc.) dumps you on by default. Often without any alternative being available.
But with ASP.NET Core it's effectively a Web API as a Lambda. It maps the event to a request and passes it through the asp.net framework as if it's a normal website.
The config for this is like 3 lines of code so if you don't want to host in a Lambda anymore, it's trivial to move to IIS, Windows Service, Linux.
X could be:
List prime numbers, List fibbonaci numbers, Conway GOL, etc
I'd just be afraid we'd have to warn the guy who wrote the article not to put any of them in production.
But CloudFormation does have drift detection: https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGui...
I'd pick Terraform before CloudFormation, but there are things that are better with CF, such as rollbacks (they really do work as expected in almost all cases), and AutoScalingGroup's UpdatePolicy (rolling/replacing deploy of new launch configuration with healthchecks), which is not available outside of CF.
This drives me crazy. I never want to touch CF, but if I want rolling ASG, I have to. Why?!?
For a concrete example of one that bit us one time, database parameter groups don't support drift detection.
In my opinion, tf is terrible choice unless you're going to be using it for more than AWS
This is a big deal with classic CloudFormation because there was no escape hatch. Terraform can run arbitrary code if needed but you had to split CF stacks apart so you could have something else run in between. That’s probably better with CDK, of course.
yes, its a complex poorly documented pile of shite. BUT. It does work as a reasonably secure OAUTH2 thingamebob. However I was told by my AWS account manager that auth0 was the way forward, and I agree.
Cloudformation:
Meh, I have about 35k lines of active CF at the moment. Its much of a muchness. Unless you are using parameters with selectors, you are going to have a bad time. Hard linking templates together (I assume thats what nested stacks are) is terrible. I've only briefly used terraform, so I have no idea if its much better.
CF _could be_ a lot better. Like compile time validation, not just in time. that would stop a lot of anger when you realise you've spelt a CF parameter wrong(or the value fails validation) but only after you've spent ten minutes for it to spin up. Thats frankly unforgivable.
Elasticache:
Yes. Its expensive.
KINESIS:
What a disappointment. Stupid naming conventions, Terrible throttling and throughput. Its just horrific. Whats worse is that they looked at SQS and thought: "this compares favouribly" NATS.io is a great fit for certain usecases (no, kafka is never the answer)
LAmbda:
I don't actually get this myself. I made a REST api exclusively in lambda. It meant that I could build a working prototype really quickly. Once proven we ported it to fastapi in an autoscale group.
The API gateway was heavily integrated into the lambda spinup (controlled in CF) so I really don't see what the issue it. Also it understands swagger, so I struggle to understand the criticism
But you can point each route in API Gateway to a different function in the same lambda.
Additionally, using serverless framework facilitate local testing too.
That solves all the complaints from the author.
What do they mean by that?
auth0 is a company that sells a variety of authentication solutions (on premise and SaaS) and a variety of authentication libraries/plugins.
For example they run the site https://jwt.io/ that anybody who's had to work with JWT tokens would have used at some point.
The difference between cognito and auth0 is that auth0 has documentation, code examples and a decent API guide.
Cognito has a poorly documented API, terrible integration guides and even worse debugging options
Once you have it going, cognito is grand. To get there, its a huge learning curve
I've recently finished a project using AWS CDK, which seems to do a certain amount of this. Just using TypeScript and having AWS resource interfaces be fully typed goes a long ways for finding a template mistake quickly.
Oh! I forgot about that mess. Yeah, it takes minutes of deploy time while real resources are spun up to catch some really basic mistakes.
Terraform plan catches a lot of issues, but I've seen cases where it misses something.
High Latency: Kafka's latency depends on your configuration and infrastructure. It has some of the lowest latency out there when configured correctly.
Curios to know why you find it difficult to look after? It's super stable, and is pretty easy to monitor via JMX.
Agree with this 100%. But what about hosted kafka solutions?
> high latency
Not sure I understand this point.
(Cognito groups seemed made for this, except they have a limit of 10k groups. We ended up storing a comma-separated list of ids in a custom cognito tag, which seemed awkward.)
but the documentation is utterly pitiful
Quite a few times over my career I have had a salesperson (not from Amazon) recommend a competitor over their own product. In every single case my respect for the salesperson shot way up. In at least two cases I can recall this helped them close a sale.
A smart salesperson does not do everything possible to push their company's products... a smart salesperson solves their customers' problems.
Bingo. A bad product fit means a bad customer experience, which means a bad review or reputation.
The smaller the company, the more important referrals are from your customers. Sending a potential customer to a competitor will (potentially) earn goodwill and future referrals. At worst, they might not refer anyone your way, but at least they won't be badmouthing you either.
Unfortunately, large companies typically mean large customers, and the people with the buying power aren't the people who will be using the product... so neither party really cares all that much about how well the product fits. This is the old "nobody gets fired for choosing IBM" mentality.
The worst is when medium companies think they are big companies, and try to do that to small customers. I once saw a salesperson push hard for something that was very obviously too small to be worth our time, and the project management overhead would have lead to blowing our potential customer's budget out of the water. In the end, they walked away without working with us, and a pretty sour taste in their mouths from the pushiness of the sales guy.
If you look at it from their point of view:
We were making an API that take images does stuff on the GPU and pushes back an answer
It needed to be secure, fast and easy to look after. If they had forced cognito down my throat, and it stopped me from shipping on time, they would have missed out in $$$ of GPU time. I trusted that architect more, because they were honest, and actually helped. Making me want to stay inside the expensive walled garden that is AWS, more.
I think it comes down to the “more flies with honey” thing. The customer will go on the defensive and shutdown if you tell them they’re stupid.
Elasticache though? Pricey, but I recently got an alert at 1am that one of my core services was down and it was because Redis was out of memory. With Elasticache I doubled the size of the instance in place from my phone so I could go back to sleep. Fixed the key leak in the am and returned to the original instance size.
Did you so that using the AWS console in your mobile web browser? Pretty nifty way to quickly get back to sleep.
When something inextricably effs up, rebuild the environment and more often than not, the problem dissapears.
- The single container Docker platform (not sure if this is an issue with other platforms) can cause the CloudWatch agent on the environments' EC2 instances to stop streaming logs to CloudWatch. This seems to occur when a Docker container fails, for example if the process it's managing stops (e.g., if a Node.js application triggers an exception that is not caught and exits). A new Docker container will be started, but the new container's log file sometimes does not automatically get attached to/monitored by the CloudWatch agent.
- The default CloudWatch alarms created by the environment can create a "boy who cried wolf" situation. For example, when updating the application version for an environment, EB will transition the environment's state from "OK" to "Info" or even "Warning," depending on the deployment policy. This is a regular operation, but CloudWatch will still send an email to the designated notification email address about the state change. If you monitor those emails for environment issues, this normal operation could cause overload, which might lead to ignoring the emails outright. This could be problematic if the environment state transitions to an actual problem state. You can create email client rules for this, but the structure of the alarm email doesn't make this very easy, at least in Outlook 365.
An annoying example of this is when your EB environment auto-scales up due to, for example, an increase in traffic. When the auto-scaling policy scales down your instances (due to normal operation of the policy), you'll get an email that your environment has transitioned into a "Warning" state because one or more of your environment's EC2 instances are being terminated. This looks scary in the CloudWatch email that is delivered, but you have to learn that it's just the ASG doing its thing, terminating unused instances as it's been configured to do. The emails, however, do not provide good context into what has led to the "Warning" state.
- The way environments handle configuration files stored in your application's .ebextensions/ directory can cause inconsistent application state between version deployments on existing/new EC2 instances. For example, if your auto-scaling policy creates a new EC2 instance, but your recently deployed application version doesn't specify some of the commands/settings applied during a previous update to your .ebextensions/ files that might have been deployed to existing EC2 instances, you run the risk of having inconsistent state across your application's EC2 instances. This can be solved by using the "immutable" deployment type, but that's not the default deployment type. It's an edge case, but it's still something that requires you to SSH into your EC2 instances, and possibly manually terminate older instances when you eventually figure out what's going on.
Having said all of that, I think EB is still a reasonable choice for small/beginner workloads: It gives you a number of things (automated deployment, auto-scaling, load balancing, logs, etc.) that you can get by doing things on your own, but lets you get to production quickly. For mature applications, I think you could be better off managing these individual services yourself (EB is mostly just wiring together a number of AWS services with a a few deployment and monitoring agents running on each EC2 instance). If you're comfortable with the components EB is managing for you and if you have a stable CI/CD pipeline, you can get more flexibility than bending EB against its will.
I struggled to get Elastic Beanstalk working well. The documentation is incomplete.
The Elastic Beanstalk console often shows servers as up when they're really down.
Once, my Elastic Beanstalk deployment stopped writing log files. After wasting many hours debugging, I went to AWS Loft and consulted with the support engineer. He had me log into the backing EC2 instance and debug. I was using Elastic Beanstalk so I would never have to log in to EC2 instances. He concluded that the application logs were not appearing due to a bug in the service. He promised to file a bug report.
The last straw was when Elastic Beanstalk team deployed a broken API on a Saturday. THEY DEPLOYED ON A SATURDAY!!! Their broken API added an invalid entry into the beanstalk config database for my account. All subsequent calls to the Beanstalk API failed with a 500 Server Error. I paid for an AWS Support subscription and filed a ticket. Their support engineer told me to install the AWS CLI and run some obscure commands to remove the invalid entry which their botched API deployment added. I asked them to do it and they refused. So I migrated off of Elastic Beanstalk to Docker. I have since migrated off AWS to Digital Ocean.
For low regularity calls, Lambda suffers badly from the cold start problem. The first call to Lambda must actually create the lambda which can take hundreds of milliseconds. This problem also shows up when the level of requests exceeds the number of lambdas currently active, causing a cold start as the new lambda instance is created. It’s not uncommon to see some services inject fake requests with some regularity to insure that there are enough warm lambda instances to avoid the cold start problem, which is silly.
Secondly, all lambda calls are charged by the millisecond, but the all calls are rounded up to 100ms. So if your typical lambda call is 5ms, you might be paying for 20x more time than you’re actually using.
Both of these issues led my former team to use a regular ECS app rather than lambdas.
But also, if you don't have highly variable traffic, why would you use Lambda in the first place? If you have negligible traffic (enough to sit on the free tier), why not just use a single cheap EC2 instance? Lambda trades start time for lack of a server—it's shared resource utilization taken to the extreme. You're lowering your cost by letting AWS use the bare minimum to keep your service available, and that means turning your code off when it's not running. If you want to keep your code available at a hundred millisecond's notice, just have a server running.
I assume most of the folks running into this just couldn't be bothered to pay the $7/mo for a hobby Heroku dyno or run a dirt cheap EC2 instance. Really interested to hear from folks that find Lambda impractical for serious use cases.
I’m speaking from the context of a large company that’s well beyond the free tier, for whom AWS bills matter. If you’re down in the free tier, setup time dominates all other concerns.
I’m not sure why this is the case. You could host all modules in the same lambda endpoint /api/v1/* Using the same technology you would use with any other backend
Article's comments don't seem well researched. Just used tech. Ran into road block. Blogged. Complained.
But I find it's the dependencies and common code that takes most of the space which often is common with all functions.
CloudFormation: this is an excellent resource when you need to tell teams across our org how you want them to set up resources. Rather than have numerous lengthy meetings where we tell them what to do, we just give them the CF template, super simple, and we guarantee everyone has the same setup.
Kinesis: works great for ingesting the data we had them set up resources for with the above mentioned CF script. I can't speak to the Java dependency the author mentioned. Not an issue for us. YMMV.
Lambda: Also works great with the CF template setup mentioned above. Super cheap to use. Maybe the difference between our implementation and the authors is the frequency and trigger used. Our lambda functions are all time based and run once a day, or maybe a few times a day. Super reliable, super easy, super cheap.
I think the overall thesis here is that the usefulness of AWS services depends on what you are using them for.
It helps if you think of Kinesis as AWS Kafkaesque rather than a message queue, because then shards = partitions and how you work with a Kinesis stream makes a lot of sense.
Multiple concurrent consumers? You're going to need shards/partitions.
Now, question is, as a replacement for homerolled Kafka in EC2, or AWS's Managed Kafka, is it cheaper?
So far, magic 8-ball says "benchmarking for your given use case needed". So far my experiments for our workload and patterns say - maybe, but more investigation is needed. I plan to role out a side-by-side Kafka and Kinesis experiment on a given topic and ascertain the costs.
Although ultimately, like any distributed messaging system, you end up engineering to the foibles of the system - in Kinesis' case, it's the fact that it rounds each sent message (which could be a batch of records, or a single record) up to the nearest 5KB for billing that you have to engineer for.
They wanted to have a Kafka but they also had this silly idea that everything needed to be HTTP so they built a poor clone with a HTTP interface.
Unless forced into it by some other system there is never any reason to consider Kinesis over Kafka or Pulsar.
As to the Kafka vs. Pulsar, I've been running a Kafka cluster for several years now, and was interested in Pulsar as a Kafka++, and have been evaluating it for a client who wants to choose between one of the two, and at the moment, I think Pulsar needs about another year before I'd recommend it.
It's exciting tech, but as is normal of recently open sourced projects, there's a lot of bugs being surfaced, and the documentation has a lot of unanswered questions, like, if I'm consuming a multi-partitioned topic in an exclusive subscription, what does that mean for ordering?
I think it has some great ideas (especially decoupling brokers from storage), but yeah, it feels a bit like Kafka pre 0.8, interesting, but you're taking on a lot of work adopting it at the moment.
That said if what you are trying to do fits squarely in the box of things that do work well it's considerably better than Kafka in a few keys ways.
The first is the obvious separation of storage and compute/broker responsibilities, the benefits of which also led to the same design being used in LogDevice - a similar system designed and built completely independently and at roughly the same time at Facebook.
The second is selective acknowledgement, i.e the ability to acknowledge messages as processed by a consumer out of order rather than merely the offset of the latest message processed. This allows Pulsar to be more easily used as a workqueue without the multitude of hacks and layered infrastructure required to get the same out of Kafka.
Shared subscriptions/partitioning model. Compared to Kafka its more flexible, less punishing and less beholden to the architecture of consumers.
Finally I would say tiering, for some it means nothing but depending on your use case it can be defining. Pulsar can offload historical segements to long term stable storage but still present a unified offset like API to consume historical data.
So I think for some Pulsar provides enough benefit to make up for any shortcomings in integration as long as you have sufficient engineers able to debug/patch issues.
Then ontop of that add Consumer Groups which basically deal with the issues in the OP w.r.t consuming a topic from multiple processes along with providing administrative APIs to reset application offsets, inspect lag etc.
Also a bunch of extra features like transactions etc but if you are comparing on the basis of Kinesis like features they aren't likely to matter as much as the core functionality - which is where Kinesis really gets destroyed anyway.
Normally I wouldn't take such an absolutist position but when it comes to Kinesis let me repeat in no uncertain terms.
Never use Kinesis.
Kinesis on the other hand, just works. Yes, it just works. I don't have to have a Kinesis expert on hand, I don't have to configure clusters myself or write Ansible/Puppet, etc. I have a few basic lines of Terraform to create my Kinesis streams and I push data to them and we got it working in minutes and we've had no issues.
Contract this to my previous job where we literally had to hire multiple Kafka experts at high salaries to maintain our Kafka clusters.
This is why you only use Kafka if you ABSOLUTELY NEED it.
I actually did miss something but it's not complexity, you can easily use hosted Kafka (AWS even provides it with MSK now). It's actually authentication. Kafka does have authentication but it doesn't easily tie into AWS or GCP authentication mechanisms without the help of a tool like Vault.
That said.. I don't think that minor benefit is enough to ever justify Kinesis over Kafka.
No, it is not. Also, an equivalent cluster of EC2 instances running zookeeper/Kafka/etc will outprice Kinesis.
I think for any reasonable workload, i.e 10k/s+ and/or throughput over 100mb/s Kinesis gets dumpstered for price every time.
I'm not wedded to one or the other, just interested, as I said in another comment, I'm looking at Kinesis from a cost engineering basis alone, but if there's something I've overlooked, would be keen to know :)
> Lack of Drift detection or reconciliation. With lack of drift detection comes great uncertainty.
Cloudformation has definitely had Drift Detection since 2018: https://aws.amazon.com/blogs/aws/new-cloudformation-drift-de... It's not everything you might want it to be, but it's not like Terraform will reconcile your drift automatically either, that I know of.
Terraform does indeed reconcile drift 'automatically' across all resources, by which I mean the plan will include changing everything that's drifted back to the specified configuration. That may not always be desirable, which is why building a good plan/apply process with approval is important. (Same goes for CloudFormation, though.)
I’ve heard this a couple times after complaining to AWS engineers about CloudWatch shortcomings.
That said, CloudWatch has gotten tons better than a few years ago.
Also the CloudWatch Logs console is unusable for even simple tasks.
Datadog and its competitors have ingestion rate limits, but they are overall quite expensive. Self-hosted log analysis tools are exceedingly complex to set up and maintain: ElasticSearch + Kibana, Grafana + Loki.
1. Deploy host running dockerd
2. Deploy Grafana server
3. Configure Grafana server admin password, organizations, users, and passwords.
If Grafana would just support file-based configuration then a whole stage could be eliminated.
"Loki does not come with any included authentication layer. Operators are expected to run an authenticating reverse proxy in front of your services, such as NGINX using basic auth or an OAuth2 proxy."
https://github.com/grafana/loki/blob/master/docs/operations/...
Unfortunately, this means that any process that can write logs can read all logs. This violates the principle of least privilege, a core part of system security. Prometheus suffers from this, too. To put this into concrete terms: When someone uploads a malicious image through our app which exploits our image resizing server, they will obtain log-writing credentials. Those credentials should not allow them to read all the logs and steal the user data in them. That would be catastrophic for the users and the company.
ElasticSearch supposedly has ACL support, but it is mostly undocumented and full of foot-guns. For example, their security doc omits a necessary flag which enables password enforcement. After following the guide, I discovered that the passwords I had set up were not being checked. I immediately lost all confidence in ElasticSearch as a tool to safely store user data. I deleted it.
I used InfluxDB since it lets me create write-only user accounts. Unfortunately, Grafana's integration with InfluxDB is problematic.
The lack of useful debug logging in Grafana makes troubleshooting especially difficult.
That said, this point gave me particular pause, because it brought in to question the other assertions:
> The application collected small json records, and stuffed them into Kinesis with the python boto3 api . On the other side, worker process running inside EC2/ECS were pulling these records with boto3 and processing them. We then discovered that retrieving records out of Kinesis Streams when you have multiple worker is non-trivial.
Yeah.. because that's really not what Kinesis is designed for. They even, rightly, point out that SQS is a better fit for that purpose. That hardly makes it a service to avoid.
Terraform is better in almost every way. If you use Cloudformation you'll end up writing a bunch of bash script wrappers or similar around it to make it actually do what you want.
Everyone who's tried both at scale has said the same thing.
I disagree. I spent time in Terraform a few years ago working with a client and Terraform had the ability to create but not tear down resources for some services. I was shocked -- check out the Github issues history. I ended up writing a "bunch of bash script wrappers or similar around it".
Overall, I think Terraform is better if you're deploying a lot of inter-connected resources, but Cloudformation makes a lot of things "just work" in ways that Terraform doesn't. I think of it like ECS vs. EKS/Your Own Kube Whatever. ECS is full of gotchas and limitations, but if you play along with it you get a lot of things "for free".
Terraform makes mostly sense -- because there is good tooling around CloudFormation, and CFN more often than not does the well enough -- when you find out that CloudFormation is like Terraform just a driver for AWS APIs, likely developed by a dedicated team, and as such features of the underlying APIs are often not exposed at release time in CloudFormation. I noticed this especially when experimenting with EKS: There's eksctl, which integrates with CloudFormation (generates some stacks) but is utterly useless for integration because you can't import outputs or exports, so you have to hardcode all your SG IDs, VPC IDs, Subnet IDs if you want to integrate with existing infrastructure (https://github.com/weaveworks/eksctl/issues/1141). Dealbreaker, waste of time, disappointing for an "official" CLI. Next, there's pure CloudFormation -- but no luck for you, to this day AWS doesn't support EKS endpoint access control settings through CloudFormation (https://github.com/aws/containers-roadmap/issues/242 - dealbraker, if you need that and are allergic against public control endpoints in your infrastructure. Looking at Terraform -- it supports integration with existing CloudFormation infrastructure and can access CFN exports, it supports all the EKS settings you'd want, it offers a consistent interface to these features, so you got something to use, short of rolling your own.
Mind you, CloudFormation is extensible using Custom resources, but hacking around CloudFormation is likely not worth your time and something Amazon should do. Anyway, CloudFormation is likely one of the most tested and well-working parts of AWS, so I'd prefer it over third-party state management unless I find a very good reason to do so.
You can reference Parameter Store and Secret Manager entries in CF so you don’t have to hardcore values.
So you just click on stuff without knowing how much it costs? Really?
> Secondly that cache.* prefix means this instance costs $0.216/hr instead of $0.126/hr, a 71.4% premium. Then you might think you need one for dev, qa, and prod
Ok and you pick a cache of the same size for dev/qa/prod? Again, really?
> But, not all data/records/events should go into Kinesis. It is not a general purpose enterprise event bus or queue.
Yes, there's SQS for that. But it seems another case of clicky-go-lucky. "Hey it says queue here so we'll just use this right?!"
> Lambdas are great for the following tasks:
Yes, agreed. I would generalize it as: it is good for specific small tasks, especially when plugging into the rest of the AWS infrastructure (reacting to/from SQS, S3 events, etc)
> Lambda is horrible for: > A replacement for REST API endpoints.
Sigh
General advice: if something feels weirdly hard, then you're probably doing it wrong.
These days I would recommend using CDK if you can’t switch to Terraform but I would never under any circumstances recommend YAML since the odds approach certainty that the magic in the parser will cause a problem. Beyond the usual confusion around typing I’ve seen things like significant whitespace breaking functionality[1] and the shorthand types in a few cases make diffs messier. If you find an example in YAML, it takes a second to run it through cfn—flip and you’ll never have to deal with any of that.
For anyone using CF seriously, I highly recommend the cfn-python-lint tool and associated editor support. It will catch many of these cases before you burn a lengthy update cycle.
1. Fun fact: AWS’ own security alert CF stack fails the Security Hub CIS scanner because it adds extra white space the metric filters.
Only if you use a framework that splits it up that way (which I agree is horrible). But nothing about Lambda _requires_ this architecture. In fact, API Gateway has a proxy mode that allows you to serve all requests from a single Lambda.
I've never used Cognito or ElasticCache, but for the other three (CloudFormation, Kinesis, and Lambda), I would agree with what he's actually saying, which is to use these services judiciously.
CloudFormation has some advantages over Terraform, but having used both, Terraform is much more usable for long-lived static resources. I have yet to try Pulumi, and I don't see any reason why I should to be honest. However, CloudFormation is required for automated build and deployment of Lambda functions--both Serverless Framework and AWS SAM use CloudFormation. I've also used CloudFormation StackSets to build out infrastructure for multi-account governance. In normal use though, Terraform is typically more usable, though both solutions can get pretty hairy.
By "Kinesis" the author seems to be referring to Kinesis Streams. In fairness I haven't built anything where Kinesis Streams would be a better use case than SNS/SQS, but that's just as much a statement about the projects I have happened to work on than it is about Kinesis Streams. I have used Kinesis Firehose, which provides a very scalable mechanism for, "multiple clients are intermittently spitting out a massive volume of data points, all of which need to be faithfully logged into S3 somewhere".
And finally, we come to Lambda. Lambda is a good fit for use cases where either you need to run some code on an intermittent basis or when you want to prototype faster and don't want to mess around with provisioning and deploying to servers. Lambda is a great place for little pieces of glue code that get triggered on a x-minute cron or based on a CloudWatch event or from SNS or SQS. It's good for serving an HTTP API that gets less than one request per second. These use cases are extremely common and I've encountered them a lot more than the author seems to, though again maybe that's just me. But for anything high traffic, where the Lambda container is just going to stay hot, just deploy it to EC2 (perhaps with some container magic in the middle), it's cheaper in the long run.
- defining dependencies between stacks
- taking outputs from stacks and feeding them in as parameters to other stacks (not using that awful Import/Export Value crap they implemented)
- deploying as many stacks in parallel as possible and waiting for them to complete before deploying dependent stacks
- dealing with common failures and rollbacks, including handling known "continue update rollback" steps with predefined resources to skip
- pre/post stack create/update/delete actions to make API calls and perform other actions outside of CloudFormation
We basically built the missing pieces of CloudFormation ourselves and have managed to keep CloudFormation holding the definitive state rather than managing it ourselves (or paying Terraform to do it). We have about 500 stacks for a single deployment of our product.
We also implemented a large number of Transform Macro lambdas to composable templates _much_ easier.
There's a lot you can do with CloudFormation, but it takes some investment.
A few things about it that really drive me nuts, though:
- Lagging support for new features/resources
- Parameter count and parameter size limits
- Certain bugs with some resource types that are slow to get fixed. (Redshift cluster password management and issues with elasticsearch resizing triggering blue/green deployment come to mind.
- No ability to modify timeouts for some actions. For example, the timeout for a CustomResource is fixed and cannot be tuned -- if your Lambda never responds, CloudFormation will hang up to 2 hours. We wrote our own Lambda wrapper just to guarantee a response if an unexpected failure occurs
I will say, though, the the KCL library is a PoS and the few containerize services that are reading from Kinesis exhibit this annoying behavior of causing the iterator age to spike to the TRIM_HORIZON when performing a blue/green deployment for the containers.
You do? Why can't a function be written to handle multiple routes?
Many more than the listed services have serious usage limitations or gotchas and the listed ones have more limitations. Workarounds would also be useful to know.
Cloudformation isn’t perfect but it is well integrated within AWS and you can get support for it around the clock. Hashicorp quoted me $13,000 just for support for Terraform. AWS support isn’t cheap either but covers everything in AWS and their support is fantastic - like IBM in their heyday.
You should use CloudFormation if it's there and you're just trying to stand something up quickly. Shell scripts with AWS CLI works too. For more robust long term stuff, use Terraform, Pulumi, or custom Boto3 code. The point is to just start using the Infrastructure as Code pattern early, not to make it perfect. Terraform is surprisingly painful after a while, but it's easier to standardize your organization on large-scale. There's no great solutions in this space.
You should absolutely throw money at ElastiCache to quickly scale up and down a cache. If you're not a very experienced sysadmin (and fuck the entire tech industry for making that term a dirty word - if you are not experienced at administrating systems) you'll be spending unnecessary time and energy to stand up and maintain it. Your caching should ideally just be in your service and scaling your services themselves should be sufficient, but whatever, Redis wasn't invented for people who write good code.
> But, not all data/records/events should go into Kinesis
Not all of anything should go into anything. Use it if it's convenient.
Lambda is very useful for batch jobs and CloudWatch-triggered maintenance tasks. I wouldn't rely on it for anything more than that; I would rather just deploy the same code to a Fargate Spot Instance and not deal with all of Lambda's bullshit.
Don't use: - Cognito (authentication). Because: on mobile social login it isn't native. Instead: use Auth0, OneLogin, Okta or roll your own
- Cloudformation (programmatically configure AWS). Because: various complexities. Instead: use Terraform.
- ElastiCache (managed Redis). Because: expensive. Instead: run Docker Redis in EC2.
- KINESIS (queue), as a general-purpose data queue. Good for: streaming data such as video processing/uploading. Bad for: generic data queue, because its's difficult to route each event in the queue to one of multiple workers. Kinesis is meant for every listener to be assured to get the entire stream. Instead: SNS/SQS (SQS FIFO) or a queuing framework that sits on top of Redis or traditional databases.
- Lambda (server-less), to implement a REST API. Good for: Serving/Redirecting requests to CloudFront, also reacting to events from SNS or SQS by running small asynchronous tasks. Bad for: replacement for REST API endpoints, because it's too hard to work with a zillion lambdas. Instead: use a regular web framework (or, other commentators say, route all requests to one lambda and then do more routing within than lambda)
I do not agree or disagree, I'm just summarizing; although I did add in some information from other comments here, too.
https://github.com/aws-amplify/amplify-js/issues/3495
While I haven't personally had the opportunity to run into these issues, the feedback there shows a serious lack of ownership that I've never encountered elsewhere with AWS.
You may also enjoy trying an incognito window in the future
Regarding cloudformation, I view manually writing cloudformation json/yaml as an anti-pattern. I would recommend looking at using CDK, which is a framework for writing your cloudformation stacks as actual real code, not markup. It's a great tool, and I find it enormously more productive and expressive than terraform.
CDK allows you to write CDK apps in a variety of languages, but I would recommend just using typescript, since that's what its written in. I tried the java bindings, but it was pretty clunky.
For a non-AWS specific CDK like experience, you could also look at using pulumi, but I haven't really used it much myself.
Trying to define infrastructure purely by declarative markup based languages is such a waste of time.
My employer uses three of these heavily (ElastiCache, Kinesis and Lambda) and we get quite a bit of leverage out of them.
ElastiCache in particular surprised me. At first glance I mistook it for a transparent (and expensive) wrapper around sticking Redis on an EC2 instance, but if your usage is heavy enough to need multi-node clusters (e.g. read replicas or full Redis Cluster), its orchestration features are pretty useful. We can resize instances, fail over to a replica, and reshard clusters, with zero downtime, by clicking a button (or a one-line Terraform change). And never having to install security patches is nice too.
It certainly is expensive, though. (But if you're not willing to pay a premium for managed infra, what are you doing on AWS in the first place?)
See https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGui...
As a rule of thumb, we always try to avoid such proprietary services for potential vendor lock-in, whether AWS or not.
The approach CFN used made sense when there was no alternative, but Terraform's approach is so much better. I wonder how AWS approaches this space in the future. CFN has great vendor lock-in, and needs to remain supported for any customers who are locked in, but its harmful for any new development to use CFN.
I get that AWS likes having infrastructure framework level lock-in, but they need to do better.
Does anyone have a feel for how hard AWS support, account reps, or training push CFN versus Terraform? I've been out of the AWS cloud for a while.
Customer obsession is a big deal for support engineers. So if they can, they'll try to help solve a Terraform issue. But it'll be on a personal, best-effort basis.
For such a mud-slinging article, I'm surprised the author didn't complain about the most pressing issue (IMO) with CloudFormation: There have been a few instances where thing in AWS changed and Terraform was updated more quickly than CloudFormation was. Happens less and less now though.
This is not entirely correct. "AWS Support Business and Enterprise levels include limited support for common operating systems and common application stack components" as documented here: https://aws.amazon.com/premiumsupport/faqs/ under "Third-party software"
Elasticache is expensive and the analysis is generally correct on that, however if you do actually need a high capacity rock solid Redis cluster and you can afford it, it does the job. Paying someone to setup and scale a cluster of that magnitude would probably cost a year's worth of usage. Of course if you don't actually need to scale it yet then just use the suggestions in the article to run it yourself.
You can use Terraform with any cloud. You cannot seamlessly move your AWS configuration to Google Cloud but you can still write Google Cloud configuration with the same paradigm as your AWS configuration.
Why should I use Pulumi over Terraform? (Genuinely wondering)
I want infrastructure as code, not infra-as-yaml. I want types checked by the compiler, as is good and proper. Apparently many people don't feel that way though?
To me, Pulumi is a godsend.
I love the simplicity.
I love the 3rd party integration that comes with it.
I love the fact that I don't have to manage user credentials and everything else by myself.
However,
I hate that the UI is not polished, and there's very little customization allowed. As an example, from the sign in/up page, you can't add a link to go back to the home page.
I hate that if user signs up via Google, and later on tries to reset password via the Google email, Cognito silently drops the request.
I hate the js library. Hard to use and documentation is not great.
you pay what you get.
AWS's service is always more expensive than the service you built, but at the same time more realible too
Very Flask like and includes CLI components to automate deployment. Even includes simple decorator based event mapping for S3, SQS, scheduled tasks, etc.
Unlike the author I'm not afraid of lambda but generally prefer containers for apps
I do hope they will get tiered RIs for ElasticCache like they have for RDS.
I'm not following this. You don't read logs one by one. You filter by source, error code, time range, etc. Its been a while, but I'm also pretty sure CW has log groups.
This. If you have three sets of logs intermixed together, things have been architected incorrectly.
Factually incorrect. It takes as many routes as you want. I mostly ok what he is saying but some of his claims are just factually wrong.
I lack experience in AWS or clouds but I wonder what is the alternative to CF if I should avoid it? I can imagine alternatives to the rest but I cannot come up with one to CF.
Terraform https://www.terraform.io/
Pulumi https://www.pulumi.com/
Hand written Cloudformation is on Hold on thoughtworks technology radar. https://www.thoughtworks.com/radar/tools/handwritten-cloudfo...
I've never heard of this magazine but it like a very useful curated list. Thanks for mentioning this!
cloudformation is awesome and it works as advertised. I have literally not seen it fail catastrophically ever since it came out. it just works.
terraform? or terrafail as it’s known. it’s a disaster. it cannot keep track of the resources it creates. when in trouble, it throws its hands up in the air and you’re on your own. do yourself a favor and don’t learn the hard way - in production - that terraform cannot take your infra from point A to point B (and rollback in case something went wrong).
the fundamental problem is that terraform is not even a half baked tool (what version was it again? 0.12?) and people are betting the farm on it. guess i’ll go build more software while others are twirling their hands with the super HCL language (what’s even more aggravating is that Hashicorp got this part - the language - right with Vagrant but I guess reinventing the wheel and using go was sexier).
It's a super obvious way to use S3 and Lambda, but the docs recommend something insane.
https://aws.amazon.com/premiumsupport/knowledge-center/cloud...
https://stackoverflow.com/questions/38752985/create-a-lambda...
better use some randotool and talk big about Iaac when in reality the only people that I have ran across and loved Terraform either 1) used it on a really small scale or 2) don’t really have a lot of experience with the cloud.
the only reason to not use Cloudformation is that you’re not using AWs. Use whatever your cloud provides.
For everyone that has objections, an appeal to authority: I was building stuff on the cloud before you knew the cloud was a thing. I worked with AWS, GCP, Azure and [sad face] OpenStack extensively.
Off-topic, but this is an antipattern right?
The trouble is that without a time machine, it's very tough to realize that your use case isn't the one DynamoDB is good at. The docs and Best Practices certainly suggest that anything is possible.
Do you need God Tier Scale (and understand enough about DDB to achieve it (and have the $$$))? DynamoDB is awesome! Otherwise... maybe consider a few other options...
I tried it and found the experience to be far inferior to Ansible, especially if you provision actual Linux instances with it.
Terraform was always getting stuck in an inconsistent state, and trying to roll back, which doesn't work well for any type of stateful resource.
I guess Terraform may be ok for provisioning resources which are stateless and interchangeable, say security groups and other simple services.
I can't say I enjoyed using it much, though the rest of Hashicorp's offerings, Nomad, Consul etc are really great.
In fact, Terraform + Ansible is a fairly common and powerful stack.
It's all about the right tool for the job, and managing your linux install with terraform is not that.
If what you are CRUD-ing doesn't map well to the terraform concept of a resource, then don't use it. SO i don't like to use to create anaything inside a VM. Using terraform to create an auto-scaling group is good, though, as it can be represented as a stateless resouce. Or rather any state in it is transient.