Cloud Spanner is now production-ready
cloudplatform.googleblog.com
cloudplatform.googleblog.com
(I work at G)
Define "maintenance window" - the DB becomes read-only? Or unreachable? This sounds insane to me.
If you're running anything production-worthy, you'll be using one of AWS's multi-AZ/failover solutions for your production RDS instances. In that case, the database will perform a failover in those instances so there isn't downtime.
Have you done that? I have and there's been downtime every time (~5m).
At our shop, we haven't had any problems with RDS maintenance at all -- and we would indeed get paged for even 1 minute of database downtime, much less 20 (!). That said, we don't have that sort of monitoring in place for any of the test/staging DBs (which are not multi-AZ).
insert mind expanding gif
* How precisely is it measured? * What happens if it is not met?
As observed by the other posters, scheduled maintenance doesn't count.
Some people get disillusioned at this fact, thinking RDS was something more magical. But AWS does offer the Multi-AZ option, which does the maintenance on the standby, fails over, and then performs the same on the main instance. It's effectively transparent. RDS also makes backups, restores (including point in time), encryption using KMS etc. really easy. If you don't have a full-time DBA, RDS provides for a very easy to use DB with virtually no maintenance required, but Multi-AZ is absolutely required for any kind of production deployment.
AWS's answer to Cloud Spanner is much more likely to be a future cross-region replicating version of DynamoDB or Aurora, not RDS.
I'm sure there's scenarios where it's worked well for people, but I've never seen it happen.
When is it going to be available for non-companies in Ireland?
Disclosure: I work on Google Cloud, but I'm definitely not a lawyer, tax lawyer, accountant or any such financial type.
https://cloud.google.com/spanner/sla
"'Downtime' means, with respect to any Production-Grade Cloud Spanner Instance, more than a five percent Error Rate for the instance. Downtime is measured based on server side Error Rate. 'Downtime Period' means a period of five consecutive minutes of Downtime with a minimum of 60 requests per minute. Intermittent Downtime for a period of less than five minutes will not be counted towards any Downtime Periods."
For monthly downtown between 99.0% and 99.999% the customer is credited with 10% of the monthly bill. The customer is required to track the uptime and request an SLA credit.
That's unfortunate. The most frustrating part about cloud services are exactly these kinds intermittent performance and responsiveness issues.
You're right, but that's basically 'by-design' as far as the internet is concerned.
You have to consider the rate at which the raw data is collected from the replicas ---and the jitter across all of them!-- then the rate at which that data is aggregated, and so on. (The above assumes that Cloud Spanner uses Borgmon for monitoring.) Any hiccups in the collection itself will result in less noise if you look at 5m vs e.g. 1m. You could increase the sampling rate, but then you're making your DB replicas spend more precious cycles on work that is not serving queries or replicating data.
Most enterprises that have to maintain global consistency are better off with Spanner. I doubt there's many, though, that really need that vs splitting stuff into locales that are consistent in a smaller, geographical area that allows regular, DB clusters.
it may be true that spanner is currently better than any free software option, but you are much better off contributing to a free software project that solves this problem than paying google to keep developing their non-free alternative.
"Azure Cosmos DB accounts that are configured to use strong consistency cannot associate more than one Azure region with their Azure Cosmos DB account. "
...it lacks Spanner's ability to work across geographical regions. It's otherwise a full-featured, distributed, DB service. Lots of tradeoffs allowed.
"While Spanner provides linearizability, CockroachDB’s external consistency guarantee is by default only serializability, though with some features that can help bridge the gap in practice."
"A simple statement of the contrast between Spanner and CockroachDB would be: Spanner always waits on writes for a short interval, whereas CockroachDB sometimes waits on reads for a longer interval."
https://www.cockroachlabs.com/blog/living-without-atomic-clo...
Patent lawsuit waiting to happen? SCO Linux all over again?
I used Linux instead of Windows NT as a kid because it was cheaper and ran well on cheaper hardware. Then I became an expert at it. Then it became the OS I used for servers.
I've wanted to use spanner for years, but at this point, the longest my production SQL database will ever be down is for less than 2-3 minutes, and it's not worth more than the cost of a rack full of servers and a massive migration effort to theoretically shave off a couple minutes of potential downtime for me. If I'm going to go that route, I would probably just opt for something like DRBD. The way my current infrastructure is set up, it wouldn't add any costs for me to do this.
> including ANSI 2011 SQL support
which kind of eases this concern a bit. If Azure Cosmos' marketing emphasized this more, this wouldnt be as big of a concern for it either.
(I work for Google Cloud)
Side note: Cosmos is really confusing to me, it has 5 consistency models, supports 4 different database models (I don't think relational is an option?), multiple different APIs, etc.
For example, Azure advertises the p99 latency, but don't specify the consistent model. I'm thinking it's for the eventual consistency, so in that case what's the latency for the strong consistency model? Definitely need to learn more.
For lower QPS, normal RDBMS solutions (like mysql) have lower latency than Spanner by a decent amount. But Spanner scales way beyond what MySQL is capable, which is where it starts to shine.
I'd throw one downside out there though: lock in. While queries use standard SQL syntax [1], all write operations are performed via a very non-standard GRPC API [2]. I think the team certainly considered this and weighed the trade offs, but it does mean that after you're on Spanner, you're on Spanner forever.
Disclosure: I work on Google Cloud, but not on Spanner.
> While we have seen comparatively little demand for this feature internally, supporting such semantics improves compatibility with other SQL systems and their ecosystems and is on our long-term radar.
I'm surprised no one has challenged this yet
Are you saying that a company could lift and shift MySQL then?
> $0.90 per node per hr
That's $648/mo in pure overhead.
I'm not saying there should be no overhead but it shouldn't be that high.
They also optimistically encourage you to try it with a "free" trial. The "free" trial, assuming you haven't used GCP before, gives you $300 credit, which is enough for less than 2 weeks of a single Spanner node.
Plus according to https://cloud.google.com/spanner/docs/instance-configuration, they recommend a minimum of 3 nodes, or $1800/mo.
I know this is targeted at businesses but it'd be nice if we could get smaller "nodes" or something to lower the barrier to entry. You shouldn't have to go straight from nothing to 10k QPS.
https://cloud.google.com/spanner/docs/instance-configuration
A Cloud Spanner “node” isn't a physical machine, per se. It is an allocation of compute capacity and disk access throughput on a set of machines with access to the storage API.
(disclosure: I work at Google Cloud but not on Spanner)
The reality is that $8k/year isn't that much for a company. You can all share it, get quite a bit out even the smallest deployment and so on. However, compared to the work to slim down the minimum deployment for Spanner, this was honestly a reasonable outcome (that is, I certainly wouldn't have wanted to see the team delay a year to build a shared multitenancy version).
Additionally, Spanner does things for you that MySQL et al. don't. Having an automagic Regional (and eventually Global if you'd like) database without dealing with sharding is worth $8k/year even to me. So even if it could fit on $10/month of hardware, I don't begrudge them for charging a service fee, rather than saying "This is how much cores, RAM, disk and flash this eats".
Disclosure: I work on Google Cloud and obviously have a vested interest in you paying us :).
But I think the issue is: making the commitment to use Spanner will take more than 2 weeks ($300 in credit) to determine if it's worth using.
Why not setup a "shared Spanner" with millions of pre-populated rows, and let free accounts have unlimited read-only access to it? That would let people fiddle with the technology before sweating bullets during the 2 week trial. (5 day trial if you run the recommended 3 node minimum)
So by pricing us out of the entry level, you will forgo having startups build their platforms on it and get hooked (while it's still at 1-node mysql scale).
When the sharding headache hits and they're scrambling to scale, will they switch to Spanner? Maybe. Probably. They'll want to, anyway. But there is a big market that will be missed.
Basically yes, I'd like a shared Spanner. If $648/mo gets you 10k read/2k write, then for ~$9/mo you should be able to get roughly 140 qps read and 9 qps write, so long as cost scales roughly linearly. Assuming you can spread this over the month (1000qps here, 0qps there), this would be super appealing to a lot of smaller projects I think.
I think the way it is, it's exclusively a business tool but it has potential to be so much more. Imagine if it was possible for anyone to set up a distributed, fault tolerant, scalable Wordpress instance using Cloud Spanner, Cloud Functions and Cloud CDN any pay per uncached pageload? You could even set it up with a single click in the Cloud Console. Running your own scalable cloud service would be no more difficult than setting up a Tumblr blog.
But yes, we already have the per-op model for Cloud Datastore (and it's part of what makes it so attractive!) so you can understand that we're naturally inclined to do just what you're saying. Making Spanner multitenant in a secure, performant manner is real work. But it's what Datastore, Big Query and our other shared services do. I hope you can appreciate that getting "Dedicated Spanner" out the door was the right first step though.
And that's fair enough, I understand dedicated/shared are very different things, it's just that for me, right now, Spanner isn't viable. Change that and you'll be a whole lot more appealing!
Two node-weeks of Spanner credit, used wisely, might be enough to tell you whether Spanner fits your operations requirements, but it won't be enough to tell you whether you want to commit to creating an application architected around the paradigm it represents.
Technically stunning for sure though.
There will always be some cases of bad customer service but our support interactions have always been quick, relevant and helpful.
https://news.ycombinator.com/item?id=14356409 https://medium.com/@contact_16315/firebase-costs-increased-b...
> Even then outages will occur, in which case Spanner chooses consistency over availability.
In case P does happen (a network break), it's a CP system.
I think Cockroach suits the RDBMS world much better. It understands that there is little reason to make it scale beyond a small amount of nodes in couple of nearby datacenters, but it does offer a lot compared to pre-CAP databases, meaning that it will be cheaper to operate and with no cloud lock-in risk.
But CockroachDB suffers from worse latency limitations than Spanner since it relies on NTP/hybrid clock to serialize transactions.
Disclosure - I work for NuoDB.
When your app needs horizontal scaling DB with multiple regions, and you don't have to come up with your own sharding solution, suddenly $8000 a year sounds like a great price.
Should we really be thinking of sharding as something you can outsource to the infrastructure layer?
It just seems like having a stance about how the data in your specific domain naturally shards is probably going to pay off not even the long run, but immediately in terms of sanity checking your information architecture.
And I wonder if in 2017 we haven't gotten to the point where a cloud application should just shard, because we're trying to think about software as something that runs on a transient instance with access to a subset of data, appearing and reappearing, not a giant box with everything on it.
I get that Google engineers are darn close to abstracting away that "giant box with everything on it" behind an API, but I guess I'm asking if that's really how we should be thinking about our code.
For you and me working with data we already understand and added to our project piece by piece, we can safely shard and compartmentalize.
Writing a LOB app with 4 expected users all located in the same location, and expected storage requirements of a gigabyte a year? Go with a traditional DB, add replication if you need high availability, make and test your backups, done.
Writing the new Facebook weknoweverything app that captures smartphone audio and video continuously for all users at all times, automatically transcribes all conversations and produces AI-driven summaries of all video, all of which are saved forever and highly searchable? You're going to need a highly customized data storage architecture to have a prayer of keeping up with that data.
Somewhere in between? Some of those are good candidates for Spanner. Some aren't.
I guess it is not even meant for me (a single developer with a small project) then!
https://news.ycombinator.com/item?id=13298664
Particularly, the comments about them mixing Go and Rust wondering if they should move to gRPC. The integration complexity concerns me a bit but at least they're smart about what pieces they use. There was also a lot of detail in the architecture although consistency and availability vs Spanner was not clear in my links. I'll have to look into it more. Thanks.
At the scale they're targeting, it makes sense, especially for large companies who will gladly pay for the reduced complexity and easy scaling while removing all the operational overhead and cost of other solutions that don't come close to this.
Linked paper: http://delivery.acm.org/10.1145/3060000/3056103/p331-bacon.p...
Useful parts: "Lessons learned and challenges" p. 11. "Conclusions", p.12.
ACM link: http://dl.acm.org/citation.cfm?id=3056103&CFID=915581248&CFT...
Discussion from 2 days ago: https://news.ycombinator.com/item?id=14337817
Also Bigtable, Spanners predecessor, is not open source too.
https://www.febo.com/cgi-bin/mailman/listinfo/time-nuts (I don't understand a lot discussed there, but many interesting links and info about current deals to be found)
There also are differences between the quality of the timing output of different GPS modules, with some being optimized for timing applications (as opposed to location finding)
> GPS Time is a uniformly counting time scale ... The word "uniformly" is used above to indicate that there are no "leap seconds" in this time system.
> The GPS message contains information that allows a receiver to convert GPS Time into Universal Time
In fact the article you linked corroborates this, because they used GPS time to show how inaccurate NTP servers are.
PHP Support? Does it support views and triggers like MySQL/PostgreSQL?
https://cloud.google.com/spanner/docs/data-definition-langua...
Spanner is a database that merges the horizontal scalability of a NoSQL database with the strong consistency and SQL semantics of a relational database, basically the best of both worlds. People like to call this type of database "NewSQL."
Cassandra is also open source, Spanner is not. CockroachDB is the OSS clone of Spanner, but there are some differences and tradeoffs to be made as with all things.
(I work at Google Cloud)
Looking at Cosmos DB for (at least readonly) regional distribution for this now as well as the other options like scylla/cockroachdb/tidb.
https://github.com/tcncloud/protoc-gen-persist https://github.com/tcncloud/sqlspanner
Both are pretty new projects, and we ran into a snag with the sql driver when it came time to implement delete statements, but both are pretty cool, and we would love contributors!
True
> it is likely to be available for a good long time.
Where does this conclusion come from?
The google mapreduce paper was released in 2004. [0]
"By 2014, Google was no longer using MapReduce as their primary Big Data processing model" [1]
The google spanner paper was released in 2012.
If spanner lives as long as mapreduce it will be largely replaced in 5 years.
[0] https://research.google.com/archive/mapreduce.html
[1] http://www.datacenterknowledge.com/archives/2014/06/25/googl...
However, it is still very much available. There are still many important production MapReduce jobs running AIUI.
[work at Google]
Map Reduce is still supported today. Most teams have chosen to move to newer infrastructure because it is better. Nobody forced those teams to move.
Here Google has not dropped support, but has invented something better.
[Disclaimer - work at Google, but not on anything related to Cloud]
https://cloud.google.com/prediction/docs/end-of-life-faq
"As we've expanded our Cloud Machine Learning services, many of the use-cases supported by Cloud Prediction API can be better served by Cloud Machine Learning Engine."
In this case "invent something better" == "drop support".