Moving The New York Times Games Platform to Google App Engine
open.nytimes.com
open.nytimes.com
* This accomplishment would not have been possible for our three-person team of engineers with out Google Cloud (AWS is too low level, hard to work with and does not scale well).
* We’ve also managed to cut our infrastructure costs in half during this time period (Per minute billing, seamless autoscaling, performance, sustained usage discounts, ... )
GCP may have some advantages over AWS, but the reverse is also true and it's hard to take what you say seriously when you say something like that.
> AWS is too low level
It seems very strange to paint AWS with such a broad brush, considering that AWS has tons of services at various levels of abstraction (including high-level abstractions like Elastic Beanstalk and AWS Lambda).
Talking about the use case from the article, they release the puzzle at 10 and need to have infra ready to serve up all the requests. On AWS, you need to pre warm load balancers, increase the quota of your Dynamo DB, scale up instances so that they can withstand the wall of traffic, ... and then scale down after the traffic. All this takes time, people and money. Adding few other things author mentioned: Monitoring/Alerting, Local Development, Combined Access and App Logging ... will take focus from developing great apps to building out the infrastructure for apps.
Currently, I am working on projects that use both Amazon and Google clouds.
In my experience, AWS requires more planning and administration to handle the full workflow: uploading data; organisation in S3; partitioning data sets; compute loads (EMR-bound vs. Redshift-bound vs. Spark (SQL) bound); establishing and monitoring quotas; cost attribution to different internal profit centres; etc.
GCP is - in a few small ways - less fussy to deal with.
Also, GCP console - itself not very great - is much easier to use and operate than AWS console.
Of course, YMMV!
I don't know if this is actually true, I've never done any serious work in AWS.
Additionally, Big Query is far more expensive than Athena, where you have to pay a huge premium on storage.
The biggest difference is that what amazon provides you in infrastructure, where as google provides you a platform. While app engine is certainly easier to use than elastic bean stalk, you have very little control over what is done in the background once you let google do its thing.
We've used 5 different providers for a global system and GCP has won by both performance and price. We still use Azure and AWS for some missing functionality but the core services are much more solid and easy to deal with on GCP, which is also far more than just app engine.
* Kinesis Streams: Writes limited to 1K/sec and 1 MB / shard, reads limited to 2K/shard. Want a different read/write ratio? Nop, not possible. Proposed solution: use more shards. Does not scale automatically. There is another service called Kinesis Streams that does not offer read access to streaming data.
* EFS: Cold start problems. If you have small amount of data in EFS, reads and writes are throttled. Ran into into some serious issues due to write throttling.
* ECS: Two containers can not use same port on same node. Anti pattern to containers.
AWS services have lots of strings attached and minimums for usage and billing. Building such services (based on fixed quotas) is much easier than building services which are billed purely pay per use. This complexity + cost optimization pressures lead to complexity and require more human resources and time as well. AWS got good lead in Cloud space, but they need to improve their services without letting them rot.
One more to add in the list.
In DynamoDB during peak (or rush hour) you can scale which increases the underlying replica's(or partitions) to keep the reads smooth. However, after the rush hour there is no way to drop those additional resources. May be someone can correct me, if I am wrong.
Perhaps you're thinking of some kind of autoscaling that only works in one direction?
You start each 24hr period with 4 chances to scale down, and after those are depleted you can scale down once every 4 hrs regardless
Could you elaborate for this? I'm not sure I understand, are you saying that 2 containers cannot be mapped to the same host port? Because that would seem normal, you can't bind to a port where there's already something listening. But I guess I must be missing something.
The solution is to use AWS's Application Load Balancers instead, which will allow you dynamically allocate ports for your containers and route traffic into them as ECS Services.
That's solved if you use Application Load Balancer instead of ELB classic for your ECS service.
Maybe put all the users data in a blob with a JSON.stringify type thing and then JSON.parse it when you get it back?
What I've found in aggregate is that GCE is a bit easier to use at first as AWS has a LOT of features and terminology to learn. When it comes down to it though, many GCE services felt really immature, particularly their CloudSQL offering.
One client recently moved from GCE to AWS simply because their CloudSQL (Fully replicated with fail-over setup according to GCE recommendations) kept randomly dying for several minutes at a time. After a LOT of back and forth Google finally admitted that they had updated the replica and the master at the same time, so when it failed over the replica was also down.
There were other instances of unexplained downtime that were never adequately explained, but overall that experience was enough for me (And the client) to totally lose faith in the GCE teams competence. Even getting a serious investigation into intermittent downtime and an explanation took over a month. By that time our migration to AWS was in progress.
GCE never did explain why they would choose to apply updates to replica + master SQL at the same time and as far as I know they are still doing this. I asked if we could at least be notified of update events, was told that's not possible.
There were other issues as well that taken together just made GCE seem amateurish. I'm sure as they mature a bit things will get better, and it is cheaper which is why I wouldn't necessarily recommend against them for startups just getting going today. By the time you are really scaling it's like they'll have more of the kinks worked out.
Stuff is just really immature and you'll get blocked by lots of things that make no sense.
If you have to get vendor locked in your quest to not have to manage servers while learning unintuitively misnamed parts of a computer, choose Amazon.
Here's a few I can remember off the top of my head:
- A (relatively) huge 6ms lag between the website and the DB
- One (random) site will mysteriously max out on memory on app-pool startup and take all the others down
- Their scheduler has no concept of timezones
- Their scheduler uses your local time when setting up the job, but UTC for other parts (this has been an open issue for over a year I think)
- Files will get mysteriously locked in deployments and the deployment process will silently fail
- Deployments will suddenly take an absolute age for no reason
- The entire admin UI will slow to an absolute crawl for hours on end
- Some admin tasks always claim they've failed, even though they've succeeded
- Their API wrapper is just wrong on almost every level
Add on top of that the worse management UI I've ever had to deal with and it makes Azure very painful to use at times. Some-one thought nesting menus in a standardised format was a good idea. It wasn't. Everything is fairly terribly named too. Want to see how your deployments doing? That's under "Deployment Options".
Performance is also dogshit compared to the cost, my 4 year old laptop is faster than their "premium" offerings.
Portal i agree is partly confusing/messy but we set up things once and then do deploys via CI infrastructure. And even the initial setup we try to automate using PS instead (to make it reproducible).
When it comes to pricing I agree, but my laptop does not do multi-datacenter so well.
Im not questioning anything you say of course. Maybe I've gone blind or don't see the issues as critical as you, or maybe I'm just more lucky.
DDB -> NoSQL, No Automatic backups, No support for ad-hoc querying, eventual consistency (though you can set to get consistency with few tradeoffs) Spanner DB -> RDBMS, Automatic backups, Enriched SQL, Strong consistency.
Let me if you still think its fair to compare these 2 databases.
Next, your client faced issues with replication in GCE, thats not good to hear, but we do face issues in our AWS RDS MySQL and Aurora very frequently. RDS MySQL error logs not generated properly. Aurora has weird memory leaks, connection spikes, starting to behave sporadically when the memory crosses 80% and so on. We are working with AWS to figure out the issue still (credits to the AWS Support for trying to help us). So, to conclude whether you are in AWS or GCE this is the trade-off of "cloud". We need to live that, if you are moving to cloud !!
My experience with GCP in the past 4 months has led me to revise my "friends don't let friends use App Engine" motto to "friends don't let friends use Google Cloud", there isn't a single service I touched (except maybe Compute Engine) that didn't have half-baked client libraries, documentation, bugs in the server part, or a complete failure by Google to even have their engs use the competitions tooling before inventing their own shitty clone (DNS)
And I would have been totally fucked over by that Postgres connection limit when we went into production, I'm glad I dodged that bullet! I hadn't bumped into that when playing with dev environments, and I haven't seen that limit mentioned anywhere.
Firewall settings work just fine for our platinum clients with complex network architectures, I don't see why it wouldn't in your case unless something was misconfigured.
Normally I'd think I had configured something wrong, except in this case it was insanely simple. A network label that allows all ports both ingress and egress, to any destination/source, definitely applied to the servers, and yet they had constant connection issues with each other.
It probably was something I did, but the combination of those issues, surprise egress bills, very laggy UI, and various other little niggles just made it not worth my time for now. I'm keen to avoid vendor lock-in anyway, so GCP and AWS don't have that many extra features over smaller providers for me.
Email: tsg@google.com
Disclosure: I work on gcp support. Not paid to be here.
With AppEngine, the beauty is that you can have many custom named microservices under one AppEngine project and each microservices can have many versions. You can even decide how much percentage of traffic should be split between each of these microservices.
What's awesome is, in addition to the standard runtimes (Ruby, Python, Go, Java, etc.) Google also provides something called custom VMs for AppEngine, meaning you can push docker based setups into your AppEngine service, with basically any stack you want. This alone is a HUGE incentive to move to AppEngine because usually custom stack will require you to maintain the server side of things, but with Docker + AppEngine, zero devops. Their network panel is also very intuitive to add/delete rules to keep your app secured.
I've been using AppEngine for over 4 years now and every time I tried a competitive offering (such as AWS Beanstalk, for example) I've only been disappointed.
AppEngine is great for startups. For example, a lesser known feature within AppEngine is their real-time image processing service API. This allows you to scale/crop/resize images in real time and the service is offered free of charge (except for storage).
Works really well for web applications with basic image manipulation requirements.
https://cloud.google.com/appengine/docs/standard/python/imag...
The best part is, you call your image with specific parameters that'll do transformations on the fly. For example, <image url>/image.jpg?s=120 will return a 120px image. Appending -c will give you a cropped version, etc.
I really hope to see AppEngine get more love from startups as it's a brilliant platform, much more performant than it's competitors' offerings. For example, I was previously a huge proponent of Heroku and upon comparing numbers, I realized AppEngine is way more performant (in my use case). I'm so glad we made the switch.
If you're looking/considering to move to AppEngine, let me know here and I'll try my best to answer your questions.
I think Google has better things to do than to pay people to comment on HN, but I do think either this person is trying too hard to sell us on Google Cloud because they like it (which isn't a bad thing per say)
Edit: I thought about it and they probably aren't related to it, probably just really enthusiastic about it (good thing) but they want to sell us on it (eh, not sure how I feel about)
Disclosure: Work for one of the cloud providers, but not on cloud itself.
I doubt it would be much different in this thread.
https://news.ycombinator.com/newsguidelines.html
The reason is that the accusation is false orders of magnitude more often than it is true—because people falsely assume someone else can't possibly be holding an opposing view in good faith—and false accusations damage the community.
Being a long time HN member, I would responsibly disclose if I were somehow affiliated with Google (trust me, I wish I was).
But one thing i like about GCP is that it allows to the limit setting in terms of cost and ensure you wont cross it. In case of AWS it can give alerts but for some reason say all you team in in one location and there is emergency like flood etc and you dont check email then you are done. I stopped using AWS after i learned that there is simply no way to set limit. Waiting for GCP to open their Mumbai region. sigh.
Also AWS is very deceiving with free tier , there is simply no way to understand which products get free and worst case is after free tier you will get charged.
WT... I had to reread this to make sure I didnt misunderstand... why not work on making the current arhictecture elastic?! #cloudPorn
"Google provides an SDK that enables users to run a suite of services along with an admin interface, database and caching layer with a single command."
I really wish AWS had a decent local dev story, rather than relying on 10 separate half-baked OSS solutions
That said, boto3 has thorough docs, but I wouldn't consider them particularly well organized. I can only really navigate AWS sdk docs because I already know what I want to do, and can google the specific terminology
If you look at https://googlecloudplatform.github.io/google-cloud-python/la..., for example, there's not even a top level navigation index that I can read through to guess what function I might need by name.
https://cloud.google.com/storage/docs/how-to
Choose a topic, and select "Python" at the top. It should provide instructions and examples using the Python libraries.
Also, we have a repo of demo projects and examples for nearly every GCP product/service, and then some. Great examples to be found here (some might be out of date though):
Thanks for the link! You're right that this is what I was looking for. Unfortunately, that hadn't shown up in a convenient place while I was googling around. Would be good to add direct links to those from the client lib references, because those pop up for e.g. "google storage python" first. (Unless they're already there and I didn't see them).
Links from the READMEs of the repo might be useful too to help people get there? https://github.com/GoogleCloudPlatform/google-cloud-python
Google's cloud salespeople pitch that they don't require any of that.
The "advance notice" and "over provision" advice is still being given for things that could scale up fairly large. (where fairly large isn't anything that exciting, really)
I do know on ELBs though, pre-warming is essential for high throughput.
AppEngine instances can typically start in 30 seconds or so. So if your spike is because your video went viral on facebook and lots of people are looking it it, that's fine.
If your spike is because you have 10 million clients with an app set to do an HTTP request at exactly 10:00:00pm, and they all arrive within a quarter second, thats a problem.
AWS charges by hour. Google by minute. If your peak traffic only last a few minutes, hourly billing is inflexible, simple as that.
I'd bet they could have seen significant cost savings on AWS by migrating to lambda, and gotten continuous scaling to boot.
Of course, there are lots of other consideration that you should be making when going even partially serverless, and it could be that the NYT chose not to experiment with Lambda for other reasons. For instance, you're effectively tied to CloudWatch for logging and monitoring, which could be a deal-breaker. Much of the processing could be happening in the DB, which would make Lambda moot. It may simply be that their estimated usage of Lambda was too costly.
(Disclosure: I work on GCP)
You remove a bunch of ELB per-request costs doing it this way and you can scale it however you see fit.
I've deployed this solution in 50k req/s environments and not seen a single user be a problem like you mention -- any motivated bad actor could cause problems in either scenario I expect.
It out depends on your application and users. Building a website? Probably not much of an issue. Building a low-latency API? YMMV. Keeping your load evenly balanced across your front end cluster also can keep your cost low, since you are able to distribute load more evenly.
That's another point: if you scale your cluster size up and down frequently to accommodate load, doing that with DNS is a nightmare.
It's also, also the case that DNS gets cached and propagating the removal of a broken server could take ages.
There's no incentive for high ranking HN posts, or any HN posts, actually. If there were, you wouldn't see others continually submit our news here before we do. This was a nice and unprompted post for everyone in GCP to read, as well.
(Disclosure: I work on GCP as a product marketer.)
EDIT: It is a double standard though, HN readers want access and responses from people on the GCP team but at the same time tinfoil hat subliminal marketing etc.
Nonsense, there are many Googlers on HN and all it would take is a colleague poking another with a relevant thread.
</noise>
Disclosure: I work on Google Cloud (as an engineer, not in marketing, despite how I like HN).
Yep, engineering blogs are also marketing, aimed right at HN readers. Attracting good engineering talent isn't easy, so companies have to do marketing on that front as well.
But, as usual, people being marketed to (in this case, HN readers) don't realize they are being marketed to.
From my experience working with NY Times they are certainly a top-notch engineering organization. They should be free to advertise that.
I don't know if these engineers are being forced to release these types of blogs, but the far more likely (and respectful to said engineers) scenario is that they just want to talk about their work. This isn't the first time NY Times has done this [0][1].
(Work at Google Cloud on same team as boulos and worked with NY Times on some of their migration pieces like BigQuery)
[0] https://thenewstack.io/caching-hadoop-new-york-times-embrace...
[1] https://open.nytimes.com/faster-simpler-workflow-analytical-...
PS: I don't work Google
<#org.google.subversion#ref-hn5!impact-8!factor-5!aID636TZ>
I can't imagine that the load would be so high that it wouldn't be possible to do it without GCP with three developers.
It would be way more interesting with performance details. :)
But maybe i had bad luck..
Heroku does my deploy in about 5 mins.
https://groups.google.com/forum/#!topic/google-appengine/hZM...
GCloud is cheaper though.
Also, VMs spin up in GCloud amazingly fast. Like 5 seconds. Feels like somebody at Gcloud just needs to go and fix this. No reason this is so bad...
My Google Load Balancers never move.. It is a single thing that points each node (physical machine) in the cluster, and distributes traffic between then.
Each node knows how to route traffic to each app. So when I deploy that app, the software load balancer at the node level will slowly move traffic over from old app to new app. Entire thing is MAGICAL. And 0 downtime, very very fast deploys.
Edit - But yes this explains iy. Changing the google load balancers is like a 5 minute ordeal. Total pain. Nice that with GKE you only need to touch them when your node count changes, which can be very rare (~monthly for me)
Does anyone know how to speed that up?
In the meantime, I recommend the following mitigation strategies:
1. Try to get into the habit of carefully reviewing and testing new versions locally before deployment. Client libraries should still work if you have a valid default application credential set up. I say this because I have a hard time remembering to do this as well.
2. Static content and templates for your site should be hosted on GCS, not deployed with your app in a "static/" folder or something. Easy to fix typos, HTML/CSS, and JS errors by simply using gsutil to copy the fixed file over, takes only a second.
3. Always keep a stable version of your app available in case you broke something in a new deployment. It's quicker to route traffic to an older version than it is to track down a bug and wait for the fix to finish being deployed.
Not ideal, but again, Googlers have to suffer this too, and are very motivated to find a way to fix this.
Also, realtime multiplayer crosswords are coming! I'll be speaking at GothamGo this year about that exact topic.