Announcing Google Cloud Bigtable
googlecloudplatform.blogspot.com
googlecloudplatform.blogspot.com
Lets compare trying to do create an RDS instance on AWS vs Creating one on CloudSQL
AWS:
1. Get AccessToken and AccessSecret from IAM
2. pip install boto
3. conn = boto.rds.connect_to_region("us-west-2")
4. db = conn.create_dbinstance("db-master-1", 10, 'db.m1.small', 'root', 'hunter2')
Done !
1. Get a Client Id and Client Secret
2. pip install google-api-python-client
3. Go through the OAuth Flow and run a server locally to capture the access token
4. Use the discovery api to generate a service object. Good Luck finding this in the documentation
5. use the uninspectable service object to create a cloudsql instance.
The reason I don't have code for steps 3,4 and 5 is because I gave up after wasting time trying to figure this out.
My point is that they've gotten into the habit of doing half assed work so I have no hopes that the've improved this time. Practically no way to automate this The only way to use this would be from the horribly slow GUIs that google provides.
EDIT:
I ended up using google cloud sdk cli and running the automations with subprocess.check_output(['gcloud', 'sql', 'instance' ... ])
1. curl https://sdk.cloud.google.com | bash
2. gcloud auth login (one time thing)
3. gcloud sql instances create db-master-1 curl * | bash
and should be discouraged. It's easy but it's like saying, hey I'm logged in as root all the time - it's easier. wget https://dl.google.com/dl/cloudsdk/release/install_google_cloud_sdk.bash
chmod 775 install_google_cloud_sdk.bash
./install_google_cloud_sdk.bash #!env bash
f () {
...code
...code
...code
}
f
Hence a partial stream will do nothing (syntax error, missing brace to be precise).Otherwise you might as well give up on downloading any software from the internet. Ever.
rm -rf ~/.config/google
and the connection gives out at
rm -rf ~/
Suddenly your script didn't install, and you've blown away your home directory. HTTP(S) is designed for reading documents, where it's OK if you can't read the document in its entirety.
That is if you don't do a
view ./configure
first and go through it. You do take a look at the configure scripts, don't you?Example: expat's autoconf script is over twenty two thousand lines long.
That's a royal pain in the arse to automate with something like ansible... AWS nailed it with a token & secret (not some horrible expiring oauth2 object).
And that's not to mention gsutils sleeping for a second or two every time it tells you there's an update.
It's like these guys have never heard of devops.
Updates are my problem; I don't expect software to sleep whenever they're available... because you never know, I might be running it in a loop under cron and it might just piss me off when I start to lose performance...
Also, if you are running stuff in Google Compute Engine, no auth flow is needed at all: there is a metadata service that you can connect to (and gcloud connects to by default) that provides credentials associated with the machine you're on.
> It's like these guys have never heard of devops.
:\ yeah...I can't imagine trying to automate a web login flow with ansible...but no one would ever suggest you do that.
(disclosure: I wrote most of the client-side auth code for gcloud)
Also, google comes with very good monitoring tools (they bought stackdriver and integrated it).
We are running a big appengine app, with quite some compute engine instances. Together with cloud-storage and biguery its a full blown solution, without having to setup every instance yourselve. I think AWS cant compete with that, you almost need a system-engineer when you work with AWS.
All in all: you are used to AWS. Making it cumbersome to get started with Google Cloud.
I do agree tough that the documention is lacking often.
Google Cloud is a much bigger vendor lock-in, you could not setup the same environment outside of it.
> HBase 1.0.0 has a bug in it. It returns maxSize instead of the buffer size for > getWriteBufferSize. https://issues.apache.org/jira/browse/HBASE-13113
Redshift use Postgres drivers. EMR uses the HBase API.
Whats more important the ease, simplicity and clarity when using a service. The reason I like boto is because its easy to find out what is possible and what is not.
"You're doing it wrong, we deprecated that last year, I know we haven't yet updated the 4 separate places in which the same thing is documented all in a different context and different way. We really should fix that but we're busy coding. Did you know we're all PhDs at Google?"
What?!?!? You're using the Google console version 3? Why? That feature isn't implemented there. Stupid you. You're MEANT to be using our new Google console version X. Why are you using our old one?
Also you missed the 16 hours you'll spend trying work out why something isn't working only to find it's actually been changed or taken out of the Google developer console, without any trace or notice left in the code to say'we moved/removed the feature that you expected to be here.'
Really you should ask Google for help on StackOverflow which is now the official Google channel for ignoring support questions, and where your question will within seconds be down voted, derided and deleted by the StackOverflow Community,saving Google the effort of not reading and ignoring your question about how to resolve the catastrophic failure of your software.
Seriously though, why not entrust your critical systems to such capable hands?
Today I came across an App Engine SDK bug that was reported almost 5 years ago, has a several years old patch in the comments, and still hasn't been remedied by Google. App Engine users who encounter it have to find the bug, apply the patch to the SDK, and repeat that upon each SDK update.
The patch is trivial
Thank you for that!
#1 The documentation is often months and many versions old.
#2 is the fact that app engine instances have to connect to back end compute engine instances using a public IP is unbelievable. The solution for an app engine instance to communicate with your compute engine app is to use their PubSub api system. To get the PubSub api system to work, see #1. The examples are months outdated and you no longer use the appengine.Context to create your OAuth token, but instead now use context.Context which you learn by going through github issue comments.
Google needs to recognize their great deficiency in documentation to be considered in the same tier as AWS.
We recognize that there's out of date stuff in there :\ This is obviously on us, but if you see something wrong, please use that little 'send feedback' link the upper right corner, as demonstrated in this gifly: http://gifly.com/3SA5/
I just set up an alert set up for feedback with a table flip in the description, so those go right to my inbox.
In fact, every piece of feedback goes directly to our bug tracking, and a human being looks at every single one. That's not an excuse, and doesn't mean that we can't do better, but we track it all. The biggest problem is the conversion from item in the bug list to actual changes, and we're working on that!
"Oh the documentation is out of date, why didn't you tell us?" illustrates the problem with Google exactly.
I don't intend to shift any blame here. The feedback link isn't a solution. I know you can't un-see docs and reclaim your lost time.
But, there are a lot of words in there for us to fix, and that feedback does help us prioritize :)
The proof of the pudding is in the eating, right? The only way I can respond to that is by delivering good stuff. It will take time for that evidence to accumulate.
For today I just wanted to communicate that we are paying attention :)
I bet it's frustrating, though, that you can't do more to address the real problems, like lack of (apparent) incentive to keep this stuff up to date. You have to resort to clever sort-of hacks like "put this magic string in your request so someone will read it". A solution which, I might add, sounds a lot like the problems people are talking about when they talk about inscrutability.
Honestly, I plan to read all of the feedback. I worked on the Firebase docs, and I've see how effective a tight feedback loop can be for improving polish.
An example: Chrome on Android goes beyond the spec by adding limitations on when an Audio element will play, requiring an explicit user action. (https://code.google.com/p/chromium/issues/detail?id=178297)
In that example, the limitation (along with other Chrome limitations, such as all the features that behave differently on HTTPS sites) isn't documented anywhere. Developers have to try it on Android, notice that something is amiss, search for what is going wrong, eventually figure out that Google have made an intentional limitation, and then work around it.
I'm not a fan of Google going beyond limitations that are necessary for security when it comes to Chrome. That said, if it's going to happen, it would be nice to at least save folk some time by documenting it properly.
I switched from GCE to AWS when we had a build break because Google released a new command line utility for deploys, immediately yanked the previous CLI causing it to 404, and provided no way of pulling the "latest" CLI. (We needed to reinstall the CLI for each build because we used CircleCI). It was part of a long string of frustrations of getting basic workflows to work.
Thanks for the feedback and stuff. I empathize with your frustration. We're listening to threads like these, and workin' hard to make stuff better.
Also, next time you're frustrated with Google Cloudy stuff, feel free to vent to me. It's like unfiltered usability studies.
I need more context though. Where are you trying to use Python 3 and what problem are you trying to solve with it?
Sorry, but here's my answer: http://i.imgur.com/tZOS8.gif
The across-the-board guidance to go to StackOverflow is not working, and we get that. It's not just that system administration questions get downvoted on StackOverflow (which I think is the parent's point) but that StackExchange isn't good for general discussion.
We (support) are trying to be much more clear that free support is not "StackOverflow or nothing." We recently updated our "community" support page at https://support.google.com/cloud/answer/3466163 to lay out all of the community support options in one place.
Bottom line: everything listed on that page has Googlers actively participating. This includes groups, StackExchange, and issue trackers. There's room for plenty of improvement, the intent is not to ignore you.
StackOverflow should not be on your support list at all, for the reasons raised. Having said that, Google should still be on SO answering questions that end up there.
What Google needs to understand is that its long standing public reputation that it has built is an organisation that actively tries to avoid providing support - ever since Google started it has tried to avoid support. That is now the reputation that Google carries into its efforts to woo the developer community.
Google has to be extraordinarily good at developer support to dispel the baseline assumption that developers have that Google really (genuinely) wants to avoid dealing with support questions.
There's a level of cluelessness to Google's support strategy that is concerning. Why, for goodness sake, would it EVER look like a good idea to push support to StackOverflow? Who is doing the thinking behind that sort of decision? It is self evident that Google's support interests and StackOverflow are not the same thing. It's the sort of decision made without really considering the detail, and that is the point about Google's support - it's an afterthought. Kind of like washing the dishes after dinner is eaten - has to be done but we're not enthused about it.
And in the end, lack of support is a showstopper for using a cloud computing platform. If the support looks sketchy then it just isn't worth risking your business by using that platform.
If Google doesn't want support questions getting through, forcing us to submit them through Stack Overflow should do the trick nicely.
First, I should point out that what we are talking about here is Google's "Bronze Support." This is similar to Amazon's "Basic Support." In both cases it's not what you should be thinking about if you need a case response time SLA, or the ability to wake up engineers at midnight. If your business depends on any platform provider I really hope you buy a support plan which gets you the ability to talk to support and engineering whenever you need. Google definitely offers these. They start at $150 per month. (https://cloud.google.com/support/) End plug.
On StackOverflow: let's stop calling it "support." It's a Q&A site with good SEO. If you have a question which "belongs" there it's a fine place to ask. We're moderating our "go to StackOverflow no matter what" messaging, but I can't see tossing it completely.
Anyway, when it comes to free support, I partially agree with your point about a single forum, in that can be confusing. Again, keep an eye on https://support.google.com/cloud/answer/3466163?hl=en&ref_to...
To your point about a single forum for everything, I respectfully disagree. Free-form discussion is different from bug tracking. Highly structured Q&A has a place too. It seems to me that every product should have - a discussion forum where users can discuss the product and raise issues - an issue tracker to collect bug reports and feature requests - a designated place for Q&A
On each of those three, support staff and (preferably) engineers should participate daily. I think the situation with Compute Engine is closest to our ideal right now: it has a lively Google Group, actively triaged issue tracker and a sponsored tag on ServerFault (which I hope we can agree is a better destination than StackOverflow)
You can get support in 10 minutes in Linode, without any payments for support. Linode can afford it but Google can't, or don't care.
As for Google's support plan thing, it might be a little odd that it costs $150 per month, but if you're spending any significant money GCP, it's well worth it. The support response times even at the lowest level are pretty good and they sometimes fix bugs.
Of course, I think if Google employees use GCP internally, it will improve at a much faster rate.
It seems like they get a bunch of engineers to come and build stuff then once they move on, the project is abandoned or poorly maintained until another team of engineers come and try to rebuild it from scratch.
Here's more proof, Google Wallet for example ... https://news.ycombinator.com/item?id=9498475
You're describing the OAuth 2.0 client dance which has always been cumbersome but necessary for browser-side auth. With Google Cloud Platform related services, you should never have to do this if you're using these services from the context of your server. As someone previously mentioned, Google has Service Accounts which greatly simplifies the auth process.
For example, to auth with gcloud-node it's as easy as:
1. Download the key.json associated with your service account.
2. var gcloud = require('gcloud')({ projectId: 'abc-123', keyFilename: './key.json' });
3. There is no step 3.
We've also made some pretty nice docs [2] with clear examples to help you every step of the way. We have yet to integrate Cloud Bigtable but it's on our radar.For a pythonista like yourself, you may like to consider trying gcloud-python [3], and it should go without saying please feel free to file issues or reach out directly with concerns if something just doesn't feel quite right!
[1]. https://github.com/GoogleCloudPlatform/gcloud-node
[2]. http://googlecloudplatform.github.io/gcloud-node/#/docs/
[3]. https://github.com/GoogleCloudPlatform/gcloud-python
Edit: Formatting.
In comparing against DynamoDB (for example), you'll have to weigh a proprietary single-vendor API against an API with a good open-source implementation (that will get even better with hydrabase), yet that is also available in managed-form on all major clouds.
Edit: although - ouch - the $1500 per month entry price-point does not compare well to DynamoDB's $5 per month minimum.
0.65 x 3 x 24 x 30 = $1404 / month
And that's before any storage costs.
If you need to store highly structured objects, or if
you require support for ACID transactions and SQL-like
queries, consider Cloud Datastore.> If you need to store highly structured objects, or if you require support for ACID transactions and SQL-like queries, consider Cloud Datastore.
http://googlecloudplatform.blogspot.com/2014/03/cassandra-hi...
How did the latency for Cassandra on their cloud platform increase by 200ms from a year ago?
For any given replication factor in Cassandra, overhead remains the pretty much the same irrespective of whether you have 300 or 3 nodes. So should the latency.
On top of that both BigTable and Cassandra use SSTables to store the data on disk (with all the compactiony goodness that goes with them), so I'm even more surprised that the difference in latency is so huge.
Would love to see the scripts for the benchmarks! I don't want to take away from a great product launch and I'm sure BigTable kicks arse in certain areas that Cassandra doesn't... I'm just surprised at the differences in latency.
Worst case, people are going to benchmark this independently and hopefully do a better job being transparent.
That's what I'd be looking for, not so much some basics on the clusters and the workload.
The biggest pain point with the current Datastore is how difficult it can be to predict your costs. Also, there are weird quirks in the pricing model (e.g. "writes" used to cost more than "reads", it's more expensive to delete rows than it is to flag them as tombstoned and continue storing them indefinitely, etc). These quirks have left people with a lot of technical debt from having designed around them.
If this is another database option (alongside the Datastore and CloudSQL) for "classic" App Engine apps, which aren't likely to be re-written for Managed VM's, then it might be interesting. However, if it's only for Compute Engine or Managed VM contexts, where you're not locked-in and are free to choose any technologies you want, then at this point I would need to hear some pretty amazing information on the pricing model before I could be bothered to even test it out. Google lock-in is painful... once you've gone through the trouble of breaking free from the App Engine jail, it's really difficult to even consider adding new lock-in dependencies.
EDIT: Doh. You have to click through a couple of links from the original post to find it, but they have indeed posted pricing specifics already.
https://cloud.google.com/bigtable/#pricing
Looks like it's priced by the number of VM nodes you want in your cluster, storage, and network I/O if you're using it from outside Google's datacenters. No metered pricing on "read ops" and "write ops". This model IS a significant improvement over classic Datastore pricing. Unfortunately, it doesn't look like you can use it as a Datastore-replacement on classic App Engine front-end instances... and I'm not sure that I wouldn't just use Cassandra in other contexts where I have complete control.
Any idea how the service partners were chosen?
Source--I wrote one of the whitepapers on the BigTable homepage.
You are intentionally forgoing things like coprocessors, and other advanced stuff listed here[1]
[1] https://cloud.google.com/bigtable/docs/hbase-differences
Google Wallet for Digital Goods was retired a couple of months back, but they still use it on some Google properties. It's still used for the Chrome Web Store developer fee & I think Android, too. They essentially just made it private.
You might argue that BigTable going through the same thing would be too impactful, but GWDG was quite impactful as well. Businesses that had subscribers paying monthly subscriptions via Google Wallet lost those subscribers.
Are you sure? I'm not a Googler but have been told by other Googlers that Bigtable has essentially been replaced internally (though heard it's still similar). So I wasn't sure how much Bigtable is even used anymore inside of Google.
Of course Google has a whole slew of other storage options optimized for various use cases, but some of these are actually built on top of Bigtable.
Ultimately, I do love App Engine. It's just the really poor support from Google that is letting down what could be an awesome platform. Neglected bugs, outdated docs, key limitations (eg. no https on naked domains, which looks like it will hopefully be fixed soon) and such.
It probably does not (yet) have the generated client for cloud bigtable checked in (but I'm sure it will), but you can always use it to generate a client. You pass it the API to use on the command line, it will go fetch the docs it needs to make your client, and put its source where you tell it to.
> To access Cloud Bigtable, you use a customized version of the Apache HBase 1.0.1 Java client.
go get google.golang.org/cloud/bigtableGovernment entities, for instance.
Google has a dedicated cloud environment for governmental agencies: http://googleforwork.blogspot.pt/2009/09/google-apps-and-gov...
So what about foreign governments?
The former is a valid position - if a fairly harsh one in the current tech landscape. If the latter however - I'd like to see a more fleshed out explanation as I don't see Google being much of an outlier on this front.