Google Cloud Platform – The Good, Bad, and Ugly
deps.co
deps.co
Couple things I wish it had:
1. Ability to `keepalive` certain functions. Cold start is real. We have gotten around this via app engine cron that triggers certain functions.
2. Schedule them on cron. We have gotten around this using app engine.
If curious, this is what we are building off of Firebase (and React native) https://itunes.apple.com/us/app/bunch-group-video-chat-games...
As far as I can tell there is no open-source alternative with feature parity that makes it easy to migrate away from Firebase. I'd be curious to hear about how people have done this.
Firebase Functions === AWS Lambda but it's only one of 17 components that is Firebase.
Also, Authentication component is nice (easy way to implement accounts/login with Google/Facebook/Twitter/email).
Only if you are pleased using a "NoSQL" solution.
Also: https://cloud.google.com/firestore/docs/server-side-encrypti...
"Each Cloud Firestore object's data and metadata is encrypted under the 256-bit Advanced Encryption Standard, and each encryption key is itself encrypted with a regularly rotated set of master keys."
So... yup.
I know that this is required by various security certifications - but is there a reasonable threat model that it actually protects against?
The only one I see is someone physically stealing the hard disks out of the servers, which is impossible if you are using a trustworthy cloud datacenter instead of a server in your bedroom.
If you are using a public cloud data center for private data with regulations around authorized access, there is basically 100% chance that people without access authorization have physical access to the servers and their disks in a manner where there is no direct knowledge of the data owner of what occurs, which makes the threat of “unauthorized person gains physical access to the hard drive and steals data” greater, not less, than “a server in your bedroom” (or, more relevantly for corporate use cases, in a corporate data center for which you control physical security.)
Can confirm, I'm also building my startup off of Firebase.
Super easy to get up and running with. I was talking to a friend at Amazon about how I create serverless HTTP entpoints and he was impressed that getting started with it was a 10 minute affair from "never done backend programming" to "REST endpoint setup." Another half an hour got me full end to end auth working from my app to my endpoints through to Firebase DB. FWIW I haven't investigated what it'd take to do the same within Amazon's eco-system, but from what I've gathered I'd have to understand a few more concepts than the simple "npm run deploy" that is the heart of Firebase's serverless offering.
Firebase's serverless stuff is dead simple. Wish the emulator was better though, I can'd do encrypted endpoints on localhost, which means 90% of my API can't be tested locally. Having to deploy to the actual service seriously adds to my code-deploy-debug loop!
Try using ngrok (https://ngrok.com/). It can forward public HTTPS to localhost HTTP.
Sanity is a headless CMS, but the API itself is like a souped-up Firebase and can be used on its own without adopting the Sanity UI. It's got many of the same features -- web API, change watching via WebSockets -- but adds support for joins, fine-grained patches (CRDT-like but not actually), transaction history APIs, fine-grained document-level permissions, and many other neat things. We are particularly proud of the query language, GROQ [2], which is a very powerful superset of JSON that allows you to express transformation pipelines over structured data.
Sanity is hosted on Google Cloud, too, so performance when used inside GCP should be great.
If you're a startup running something on Elixir (or even Rails), AppEngine's experience is hard to beat.
Not many people do know:
* You can run multiple microservices on AppEngine under one application.
* Each of these can have many versions serving different percentages of traffic.
* There is an on the fly image resizing service, which means, you don't need to run complex imagemagick setups on your machine, instead you can simply call your image with parameters appended Eg. - /my.jpg?size=120x80 and it will be resized on the fly. And it's pretty damn fast, too.
* If you're also using Cloud SQL, you can directly let your app talk to it without whitelisting IPs or even using the proxy, just using sockets.
* You can lock your app under what google calls IAP, a login layer that allows only authorized users to access your app. IF you're building prototype for clients, it's a no brainer and saves you from adding custom auth. (https://cloud.google.com/iap/docs/app-engine-quickstart)
* The development experience is superb, and has gotten better in the last 2 years.
For the record, I tried AWS ElasticBeanstalk recently and it has a lot of bugs, especially with the new interface changes, so I kept coming back to AppEngine. Seriously, if you're doing a startup in 2018, there is no reason to not use AppEngine.
[1] https://docs.aws.amazon.com/AmazonS3/latest/dev/mpuoverview....
It's a few extra steps for something that should be simpler, but it's workable.
* No integration with Standard APIs
* Slow / build deploy times
* Minimum pricing of $40 per machine / month
Overall, I feel like the direction App Engine is headed makes it purely better than other PaaS services but if you're running a low traffic website that isn't standard environment (Java 7/8, Python 2.7, Node 8, PHP 5.5, Go 1.6/7/8/9), Heroku is probably still more economical and better developer ux.
I hope they add Ruby and Elixir to the standard environment. I don't think any other PaaS would compare.
I enabled a task queue to run a longer running task in the background but I didn’t realize there’s a default to retry failed tasks indefinitely. This started slowly increasing my front end instance hours (and my bill). A few hundred bucks later I figured it out and specified a retry attempts limit.
I’m personally taking this feedback and making sure it’s shared with the relevant teams (both positive and negative).
On the DevRel team at Google, we often write friction logs (my colleague just authored a post about friction logs in detail: https://devrel.net/developer-experience/an-introduction-to-f...) to help our product and engineering teams identify and fix any rough or unexpected edges in our offerings. This feedback will be valuable to our teams the same way a friction log is valuable.
Thank you again for taking the time to write this up!
First time using said API. Creating the client was straight forward enough...oh look, there is a section of the API docs on Blobs, cool - create a blob...and there's a size property.
And it has no data. Scan through ALL of the docs on the page...ahh, there's a "reload()" method, and sure enough, calling it retrieves the data.
Though...why doesn't constructing the object fetch the properties? And if it doesn't, why don't you just write some sentences somewhere telling people the general flow they should expect for how to interact with the API?
From my experience with all other GCloud docs, the answer is - because Google couldn't care less if their customers can figure out how to use their products. You write some minimal API docs on functions, MAYBE give one example of doing one specific activity, and then it should be OBVIOUS how to do the rest.
Having used both - would never, ever choose to use GCP over AWS if given the choice - and the main reason, aside from any technical differences - is that one company seems to care about what the experience of using their product is, and the other just doesn't.
https://googlecloudplatform.github.io/google-cloud-python/la...
> The size of the blob or None if the blob’s resource has not been loaded from the server.
That's pretty vague imho, because there is no high level overview or tutorial to tell you that you need to explicitly cause the load from the server. Point in fact, the function you need to use is "reload()", and the documentation of this func is "Reload properties from Cloud Storage." To me, the "re" part makes it a bit of a misnomer - as it implies there would have been some sort of initial load in the first place (there is not.)
Imagine if the interface for going to a web page in chrome was: 1. Type the URL in the location bar. 2. Hit "refresh."
> Google couldn't care less if their customers can figure out how to use their products.
That feels like a pretty strong leap to make from your complaint about using the term "reload" when "load" is better.
I can understand you are providing an example, and perhaps you feel the documentation has many such examples. From personal experience reading extensive documentation, I think you are being slightly unfair in this instance.
If I had to guess, given the speed of development of Google Cloud, it will take some time to get to the level of mature documentation of AWS who has had years of head start.
I never made that leap - you're saying that. I'm saying that GCP docs are all severely lacking in "how to" guides.
API references are not so helpful without an explanation of how you're intended to use them. So, the problem isn't that the function name is "reload()" instead of "load()" - it's that nowhere does it tell you the pattern is to instantiate the object you're interested in (in this case, a Blob), and then you must make a call to retrieve the data from the server before the data you're interested in will show up.
Additionally, my commentary was made after now using many different GCP services - this was just something recently in mind.
I kind of agree with the other poster. Getting up and running with .NET on flexible instances has been challenging to say the least. The datastore documentation does not give a good overview or introduction to how interact with this system.
We do care deeply about our users' experiences with both GCP and the GCP ecosystem. Feel free to send me specific areas where you've experience pain in the past and I'll get them addressed (email is the same as my handle - at google - dot com)
Also to be clear - I don't think this particular example I gave is the end of the world/egregious - I just used it as representative of the fact that the general pattern I've seen of the docs following where details are explained, but the "obvious" bits of simple usage are not discussed.
Shouldn't we be able to make these requests given the built in feedback tools in the documentation, and actually expect some result? It's great that you're willing to address this specific issue, but it's frustrating to basically see plain text validation that the feedback forms are not attended to, and developers should be emailing individual team members problematic documentation pages if we're in the cohort exposed to your contact information.
Having used both GCP and AWS, although the core GCP offerings are fantastic and far superior to AWS, the GCP documentation is a disaster.
It’s totally inexcusable, and really disappointing because it has a huge effect on the perception of the platform among folks that are AWS users and take a quick look at GCP.
I can’t see any reason why the documentation issue shouldn’t be solvable in 3 months with enough money. Take a wheelbarrow down to the bank, fill it up, and put out an open offer to everyone on the AWS documentation and developer relations teams to double their salary and stock.
The core of GCP is so great, that it’s such a shame the final polish seems to be impossible to accomplish.
Also agree that Stackdriver is pretty underwhelming, though it does have the perk of providing zero-configuration log capture in GKE.
Lots of problems though -
* UI is super-slow, yet it drops old records when you scroll, so scrolling up and down by more than 100 lines causes a reload of the logs I was just looking at. (Note that e.g. Kibana will do infinite scrolling to load more log entries, but doesn't drop old ones when you scroll, which is much more usable).
* It's really hard to work with large log traces; I'd love a quick way to "show raw logs for this time-period/search" so that I can either copy into a text editor or use browser search to explore. An example use-case here is "select all logs with this request correlator", which could include thousands of logs from a request, then jump around that log stream to look for interesting events. Re-running the search is painfully slow for exploring results like this.
* I've also had problems with alerts; the algorithm used for uptime checks is pretty naive, it just checks if the endpoint was down for the entire threshold period, which doesn't catch a flapping endpoint (e.g. down for 1s, up for 1s). Grafana can be configured to detect that the average latency spiked over the period, even if it's not fully down. This has led to us not being alerted for a production outage, so I've moved our critical alerting off Stackdriver. (Again, the promise here is great; if you could get alerting working as well as Grafana, it would be super-powerful to be able to alert on any synthetic metric generated from any log stream in the cluster. It's just half-baked right now.)
All that said, I've been happy with GCP in general, in particular GKE has made running a k8s cluster a lot easier.
I'm also going to share your feedback about Stackdriver with our product team as well. Thank you for that!
I think a lot of problems would be solved if there would be a way to setup cloud sql proxy that comes with transaction pooling, especially when using k8s and having a pooling namespace (multiple load balancers for HA) where all other apps connect to.
also cloud sql and ipv6 is not funny. (i.e. a lot of things in gcp is ipv4 only which is sad)
In AWS for a simple app I'd just whitelist 10/8 and have the DB open to all internal instances (with TLS for added security as required).
I'll follow up by email too, happy to provide detailed feedback on this sort of thing.
Thanks for taking the time to engage with the community on this stuff!
This would be handy with tags in gcp.
(1) Some time ago I needed to get a month of logs (filtered by an expression) in order to compare the logged data against a database. The Console does not have any such option. You can set up "exporting", but that only starts replication from the moment you set it up; no historical data. You can use the APIs or "gcloud beta logging read", but the performance is terrible. I got maybe 200-500KB/sec. After keeping the command running continuously for 3 days (!), I still only had 20% of the log data.
These days, we have exporting set up to pipe all the logs into BigQuery and GCS, just to get around the crappiness. That gives some query capabilities, as well as the ability to quickly grab the original data for processing with other tools.
(2) We're on GKE, and none of the container's labels end up in log entries. Surely this is pretty key stuff. There seems to be no way to customize the payloads that GKE's Fluentd sends to StackDriver.
(3) There's no way to tail logs. The gcloud command doesn't have it, and the Console doesn't either. It feels like such a missed opportunity.
I think Google would be better off ditching StackDriver Logging entirely and instead let the user set up rules to send logs to different services that actually work. The default could be BigQuery. I don't understand the point of having a separate, proprietary query language for logs when you have SQL.
I've looked briefly at the other StackDriver stuff (most of which is still on stackdriver.com, always causing a confusing redirect and often re-authentication), and it's been similarly underwhelming. We use Prometheus with Grafana, and StackDriver seems pointless in comparison.
For example:
> I have reported bugs/clarifications against AWS docs and got prompt feedback and even requests for clarification from AWS team members. This has never happened for my comments submitted against Google’s documentation.
I always submit "feedback" (for the last 2 years), but the lack of changes in the docs make them seem to disappear into a black hole, so I stopped doing so.
Would be great if the docs for GCP were on GitHub (like the K8s, and now AWS, documentation) so we can at least see that docs are being worked on and iterated, and share/validate common issues and suggest fixes.
Minor things like this also create a lot of developer confusion in large user orgs:
> The API libraries and tools are spread across several GitHub organisations including GoogleCloudPlatform, Google, and possibly others, which can make it a little difficult sometimes to track down the definition of something.
> It would be easier if I didn’t need to think about this API access
Everything in the "ugly" is spot on.
I also totally agree that these little things matter, even for small organizations. An incorrect or incomplete piece of documentation can cost someone hours or days, not to mention the emotional cost. Rest assured that we are working to make the experience better.
For example, see this 3 year old thread on the same topic that is still pertinent - https://news.ycombinator.com/item?id=9497576
Specifically, support and billing.
We've used Gold support for a while, and our experience wasn't great. Pretty much the same as the OP described with Silver. Perhaps response times were slightly faster I imagine, but I would be surprised if quality was any different.
Regarding billing, we've been going through some kafkaesque cycle there trying to set invoiced billing. I've filled a form on their invoiced billing page, a sales contact talked to me on the phone and I explained everything. She then sent an email asking me to fill another form, which I did, and then got contacted by another billing team who basically sent me to the original page. I explained the situation and they just didn't really seem to care or understand what's going on. They then tried to arrange a call, but missed two schedules that I've set, called the wrong number, wrong country code... Not sure where this phone call would lead. Absolutely mind boggling experience there. Teams don't talk to each other. They send the customer around running in circles. It's amazing that it's happening with one of the most sophisticated companies in the world in the 21st century.
It gives me a somewhat unique perspective where I'm happy to give them lip when they need it and they have at times and they also ask for my input fairly frequently. In my opinion they are all in on GCP and it's definitely going nowhere but up. They've been on massive hiring sprees and it's to the point that when I visited them at the Googleplex a few months ago a bunch of their campus buildings were now labeled as Google Cloud buildings.
Obviously logos can be removed and maybe it will say Android/Allo/whatever in 3 years but the team is incredible and they seem laser focused on growing the platform. You can go through their blog https://cloudplatform.googleblog.com/ and the improvements they're making and the pace at which they're making them is pretty great.
I've used them for ~4 years so I've definitely had frustrations, especially prior to them reorganizing their Support/Account Management structure about 6-12 mo ago (I do k8s so I interact with this side of things as infrequently as possible).
This is a huge money maker for them and I think they're doing a great job albeit they still need to focus on their soft skills (support, billing pains, etc).
I have migrated a few companies from AWS to GCP and they've all been happy. If the day ever comes to leave GCP I have no problem making that call; I'm not handcuffed to them. Hopefully they just continue to improve.
I must stress out however that these are just my thoughts based on information from Internet, because I don't have extensive experience with either, so I might be wrong.
I was impressed.
Just got on the Silver plan with Google and only had one support ticket so far, but they got it resolved. Haven't really needed to push it yet, but we will see what happens.
edit: found some myself: https://www.zdnet.com/article/all-of-amazons-2017-operating-...
And that's before you even get to your legitimate concern about Google's long-term commitment to any of their offerings.
Other commenters asked if they still had jobs after saying that.
Had a problem with kube-api going down. I have alerting setup to detect such a thing. Opened a ticket, "I noticed outages for kube-api" gave them the specific times for the alerts, asked for an RFO and got back a response, "can you send me a screenshot of your monitoring software." Following which the support person set the case to "customer pending".
This kind of thing happens on 100% of the cases I open, usually multiple times from multiple support staff.
1. Stackdriver gets a bad rap. Maybe there are better solutions out there, but it hasn't been our weakest link. It's gotten very expensive, though.
2. Google recently deployed a new HTTP load balancer that is way better than the old one.
3. Instance usage on app engine is a little opaque and very hard to tune. A lot of it is trial and error.
4. Cloud SQL is garbage and the author is being generous if anything. Haven't had a chance to try Spanner, but it seems a lot better. App Engine + SQL has a lot of little hidden gotchas that are impossible to debug. I'd never recommend anyone use it in production.
Google Cloud has come a LONG way, and one of the biggest reasons we continue to invest in it is they keep improving based on feedback. Their product team is very engaged. Their IAM additions, improved load balancing, and additional availability zones have been huge for us.
Public IP access is available but that means going through the public internet, and also means you have to whitelist each individual IP accessing it, which is hard or impossible when trying to use ephemeral or private-ip-only instances.
Familiar.
On the other hand, the article doesn’t mention Stackdriver Profiler, which has an amazing UX. Definitely recommended if you are spending significant money on CPU at GCE.
Are you eventually going to support high cardinality events with the ability to aggregate and then break down (as opposed to say Prometheus which pre-aggregates, so you can’t break down to debug, and doesn’t support high cardinality).
There is a terrific explanation of the difference between Alpha, Beta, and Stable in the Istio project [0].
Istio was started by GCP, so while this isn’t officially Google’s definition I imagine it’s peobably the same.
It basically says, Beta is mostly usable in production, but things might change in the future that require work from you to upgrade.
AWS is to Linux as Google Cloud is to FreeBSD. "Rock solid performance and everything is exactly where you think it should be"
For instance I wanted to create a project progmatically a few months ago - so check out the docs. Ok v1 of the api has it...wait there’s a v2, it’s been completely removed. where has it been moved to...absolutely no clue. so what am I using now? A deprecated api?
Things I would add to the Good side.
- ability to commit to multi year cpu/ram without upfront costs
- they can live migrate your vm to new physical hardware, so no need to have random reboots of your instances
- https load balancers allow traffic around the world to enter google network closest to user, this reduces handshakes and reduces comcast like outages to affect users.
the above greatly decrease my stress.
I run into the same problem while ago. It may solve your problem as well.
Having supported VMWare at multiple locations, and having worked for a smaller cloud host, it shocks me that AWS don't support live migrations. VMWare can live-migrate VMs between datacenters in different geographic regions, while maintaining active sessions!
This is the only sane way to manage variables in terraform right now IMO. It does require that you specify a file when you run it, but that's somewhat of a bonus as it forces you to be explicit about the environment you are running against. You shouldn't really be running Terraform locally anyway so it's all handled by the CI server in most situations.
You totally can run locally, provided you are a single developer in your project. Once you start collaborating, you need this to be in a CI server or similar. Make it in such a way that the code is guaranteed to be properly versioned, taking the latest from your 'production' branch, as you can run into a situation where you are running terraform based on local changes which are not in version control, by accident or on purpose. Which now means that the state no longer conforms to the latest TF scripts, and this will come back to bite you. This also forces you to provide any settings as part of your CI job, not a one-off command in the CLI. It also give you auditing, and many other benefits.
But the most important piece of advice I can give is: no matter what your particular circumstances are, what your project size is, your Terraform skill, or whatever other variable, always use a Terraform remote state, with whatever backend you are most comfortable with. No exceptions. Even if you are running locally, I don't care. Unless you are doing testing while developing the scripts (and I'd argue even then), there is no reason to use a local state, ever.
I can’t speak for HashiCorp’s own examples (I mean, I could, I used to work there), but we have a lot of documentation and examples in the Google bits. If there’s a specific thing you think is missing, I’d be happy to help address it.
It was meant more as a critique of people choosing Terraform over CloudFormation for AWS to prevent "vendor lock in".
Their own github repo issue list said they wouldn't support deployments when it was in beta (hashicorp won't do anything that is in beta, another problem with google cloud and its long, stable betas).. Kuberenetes deployments have been out of Beta for at least 6 months, and no movement on the kubernetes provider
Deployments have been in GA 7 months now, and NO feedback from hashicorp on this one...
The two things that swayed me toward CF was reading that TF is usually slightly behind CF with adding support for new features when they were introduced by AWS and my rule of always choose the platform providers preferred solution.
Rather than Terraform offering a lowest common denominator set of resource definitions, each provider can be designed to work in a way that most naturally maps to their offerings and API.
In theory you could write modules that provide abstractions over multiple cloud providers, but I don't know how good a result that would give.
https://github.com/terraform-providers/terraform-provider-aw...
vars { ROLE = "bastion"
Then append this to the user_data.tpl file
chef-client --audit-mode enabled -R -E ${chef_environment} -r role[${ROLE}]
Allows you to use the same user_data.tpl file for multiple autoscale groups
As far as what CF can’t do with respect to AWS, most if not all of the missing pieces can be remedied with custom lambda backed resources and/or Python scripting with troposphere.
- very often something new is introduced on AWS without any CloudFormation support, so you will either need to write custom resources with lambda functions, or be prepared to wait a long time. From the few times I checked TF is actually faster to support new services / features.
- CloudFormation has extensive documentation, but very often you end up having to also read AWS API documentation and going through a dozen trial & error attempts before you get something working.
- occasionally CloudFormation can get stuck and leave you in a state you can not recover from. Luckily AWS support tends to be very responsive and can help you here, and it hasn't been happening as much as 1-2 years ago
- CloudFormation has very little support for reusing things: no macro support, very limited include support, no support for YAML aliases (unless you use "aws cloudformation package" as a workaround)
- CloudFormation changesets are nice, but do not work if you use sub-stacks (which you should use)
Just like Zope there is a bit of a Z-shaped learning curve: there is a pretty steep learning curve to start, after which a lot of things become easy. But when you get to more complex things suddenly everything becomes frustratingly difficult again. That may come with the territory; I have not used other tools such as TF so I can't tell if that is a problem-space specific thing.
occasionally CloudFormation can get stuck and leave you in a state you can not recover from. Luckily AWS support tends to be very responsive and can help you here, and it hasn't been happening as much as 1-2 years ago
I had just the opposite experience with AWS support. CloudFormation was stuck because I had a syntax error in my Node based custom resource so of course it waited for a response from the lambda that it wasn’t getting and then when I tried to cancel the stack creation, it again called the same broken lambda to delete the resource and hilarity ensued. This was after I corrected the lambda and ran it successful to create another stack.
I did a live chat with AWS support and all they did was quote the user documentation - after trying to explain that I am not trying to create a lambda resource and it’s not trying to delete an ENI.
I finally just gave up and waited 4-5 hours for the rollback to time out.
Just like Zope there is a bit of a Z-shaped learning curve: there is a pretty steep learning curve to start, after which a lot of things become easy.
I’m very much at the second valley on the Z. I spent quite awhile getting the hang of CF just to make my code deployments easier (creating parameters, autoscaling groups, launch configurations, lambdas, etc.) but I haven’t had to do anything especially complicated.
Note: If you work at Google, you might want to pass this message up the chain, thank you.
PS: i've played with all types of admob ads to try to make the system profitable, even the IN-YOUR-FACE interstitial ad that everyone hates (even Google themselves hate it, they are limiting its use)
PS #2: Google Admob has also removed the ability to create native ads.
> ...highest level of availability and performance is ideal for low-latency, high QPS content serving
And the cost using Google Cloud CDN isn't necessarily going to be a major improvement over Google Cloud Storage. Egress cost is similar depending on usage, plus you pay for cache fill and api calls.
Are they trying to get engineers to write the docs rather than hiring professional technical writers?
Truly baffling.
Thank you for the constructive comments regarding GCP and in particular documentation; we are working on addressing the points raised.
GCP has a dedicated tech writing team which is growing (we're hiring!) and the documentation is written by tech writers. We work with our colleagues in developer relations (developer advocates, developer programs engineers), UX and the software engineers who created the product.
All that said, we know our documentation is not perfect and are constantly working to improve. To that end if you see something that is broken/incorrect, please file a bug (the "SEND FEEDBACK" link in the top right of the documentation page is a bug submission form). We love to hear from users and I can assure you docs bugs do not get routed to /dev/null (we have a bug SLO, just like our SWE colleagues).
As a more general tech writing comment, documenting large, fast growing distributed systems such as public clouds is tricky. There's a lot more to documentation than just writing up instructions, such as thoughtful information architecture. It's a challenge that we, and I'm sure our counterparts at AWS et al., are grappling with.
As tech writers, we wince when we see errors in our work and feel pain when users are not able to enjoy the product/service as intended due to issues in documentation. The whole developer relations team (DAs, DPEs and tech writers) proactively try and catch these mistakes but of course, we miss some. It is nice to see users feel so passionately about documentation, so please let us know where we have made mistakes and we will fix them and try harder next time.
AWS is "infrastructure as a service". Hosts, network switches, load balancers - things that historically cost an arm and a leg, but they could virtualize. Add some elasticity and auto scaling, and you've got the foundation to build anything else on top.
The hyper-specificity of SQS vs Kinesis, for example, comes out of the same philosophy: Provide the infrastructure, and let customers figure out what to build on top of it.
Google, meanwhile, started with AppEngine - run your apps in the cloud without concern for the hardware. AWS's equivalent is Elastic Bean Stalk - more than one layer of abstraction higher than the default offerings.
Google now has a line-up of products directly competing with equivalent AWS services: VMs (Compute Engine), queueing (Pub/sub), key/value storage (GCS, which even has S3 API compatibility), SQL databases (CloudSQL, Spanner), Redis, document storage (Google Datastore), distributed file system (Filestore, like EFS), big data store (BigQuery), etc.
Microsoft's bread and butter was always the Fortune 500's and I have the impression that's where they're making most of their Azure sales.
I did a fair bit with AKS. It had a lot of shortcomings that would take a while to go into. They ended up suggesting I use ACS-engine.
There are some products that are neat, that I haven't really used at scale, but seem good. Most of the bigdata stuff is pretty legit based on my limited usage.
The customers of my employer are large enterprises. Azure definitely has penetration there. I only have a small sampling, but many of them end up leaving Azure for AWS.
None of them are really considering GCP.
- when scaling nodes, didn't get the desired count
- creating and deleting clusters takes a while (30 minutes)
- cluster operations get stuck in a loop
- disappearing node did not bring up a new node
- external management setup was difficult (GKE assigns publicly reachable IP by default)
- changing SSH user or key through azcli or GUI didn't work
- default storage class wasn't set
- admission control disallows deploying a registry in kube-system
- nodes randomly go into notready state
- no ability to add new node pools or modify pool without destroying cluster
This was all prior to GA.For example, Azure was a consideration for our "not AWS" cloud provider until we tried to use custom Linux images with at rest encryption on the root volume. It's simply not supported and there was no viable work around. There was also no ETA for adding in support for this non-negotiable customer requirement. This is why we moved onto evaluating GCP. So far it has been pretty great and it is checking a lot of our requirements that Azure either fell short with or flat out didn't support.
Microsoft is repeating what they did with NT — owning identity. They reel you in with Office 365 and when you need controls convert you to Azure AD. Once that happens, Azure is a no-brainer.
Also, every company with an EA has a contract vehicle for Azure... just add water!
Azure VMs are way behind compared to EC2 IMO. Burstable VMs (the cheap ones) aren’t available in all data centers or for all images (which they call SKUs), storage accounts are weird, provisioning times are slower, their cheapest instance types (Standard_A) are REALLY slow, documentation is confusing, extremely spread out and contradictory at times, account management via Azure AD is way more confusing than IAM at least at first, their auto scaling service (scale sets) is less flexible than AWS autoscaling, etc.
Also, Microsoft and HashiCorp are pretty much the only maintainer of the Packer and Terraform providers. The difference in commit history between the ARM provider and the AWS provider are stark. This means that when you encounter bugs or cryptic errors from the Azure API (which will happen), you’ll be waiting a bit for a fix, even if you cut the PR.
The other thing with Azure is that every large company uses it for Windows workloads because they get massive discounts and tons of credits
For example, we were looking for a solid redis hosting service in the EU, and latency is key obviously. There are virtually none with GCP, but quite a few with AWS. I'm pretty sure there are other similar examples with other services.
Google Cloud recently released public beta for Memorystore Redis (https://cloud.google.com/memorystore/) which is a hosted opensource Redis on GCP. It is availabe in europe-west1 region so it might suit your needs.
disclaimer: I am an engineer on Memorystore Redis team.
Last I read in the docs, Memorystore was not accessible through the AppEngine Standard Environment, which was rather disappointing to me.
I was wondering if just providing the App Engine Standard Environment with the Redis instances address and credentials, along with a service account or IP whitelist for access, would not suffice to be usable there as well. If I'm not wrong, something similar is done in Flexible Environment according to the docs.
https://cloud.google.com/memorystore/docs/redis/connect-redi...
Thanks so much!
They have more-or-less gone away now but several months later it leaves me with very little confidence in GLBs and quite scared to change anything around.
It's definitely not as well-integrated as AWS' forums, but it's a starting point.
I've used AWS extensively, and GCP moderately.
If you take the thoughtworks style approach, I would classify AWS as "Adopt", GCP as "Trial".
Google is investing very heavily in GCP, both internally and in sales and marketing. But, it still feels very much like patchwork. At least when you compare to AWS - which, while complex, is very consistent, very well documented, very stable, and performs very well. GCP is also, I believe, a combination of acquisitions, and this bleeds through with differently skinned UIs and URLs.
I'd say the big exception to this is GKE. If you're happy with a managed K8S, there is no comparison to GKE right now. GLB+GKE and you're pretty much set to handle anything with little ops work, and probably more affordable than AWS+Kops. Although, I'm not sure what the support story is like.
monitoring, as the author said, is absurd.
but their logging solution is even worse, specially when compared to solutions like kibana + elasticsearch. i wouldn't recommend stackdriver to anyone.
This post was interesting to me as the training course presenters seemed to make out that StackDriver was the best thing since sliced bread, but I guess they have to being Google partners...
The biggest wall for me in setting up GCP is that the resources and accounts are all so tightly tied to personal Google identities. This is unworkable for me, but I don’t see any clear way around it. We’ve set up our billing account on a shared account but it becomes hard to work with around personal Google accounts (even moreso when they are corporate Google Apps accounts). The AWS IAM identity model makes much more sense to me, but perhaps I just need a new paradigm for thinking about it. Are there any good resources for relearning how to think about resource ownership in the GCP model for AWS heads?
I'm currently hosting websites with it & despite being underpowered it seems to be hanging on thanks to cloudflare caching.
Sure, they are too complicated and unnecessary for small operations, but many large organizations wouldn't even consider a cloud infrastructure provider without those things (comprehensive permission and identity model, discernible region and inter-region model, breadth, and depth in computing types, highly auditable security model, etc...)
I get his points, but all the things he claims are great about GCP are only great for a one-man shop and not for an enterprise operation.
Also network configuration can be complicated if you track down for the detailed form.
Any specific feature you're looking for is missing in GCP?
Look at GKE :)
Would you run all your workloads in a single account? A bit messy if you have tons of resources and people collaborating.
And what if this single account gets compromised?
Thanks for contacting Google AI support. We can't escalate this to a human because we don't have enough skilled support team but our AI can solve any problem and we use you as a training set. /s
(I know its harsh but I'm sure people here can relate to it)
For what it's worth, I'd imagine most people would be hesitant to use a free service for something as critical as their monitoring. Especially since the first term in the terms of service says the service can terminate at any time, without notice.