Google Cloud Spanner is now half the cost of Amazon DynamoDB
cloud.google.com
cloud.google.com
Google runs their tech stack as if it's a startup that builds their CV. Everything is immature, tons of hacks, undocumented features. If you are on their k8s there are tons of upcoming new versions and features that force you to revisit key hacks you put in your infra because of their misgivings. Our infra team keeps tinkering around our infra and it never ends. It's 50:50. 50% of time making sure we are prepared for their shit and 50 % our ambitious infra plans. Good luck with that.
With AWS our bill is 60% of what GCP used to be running 3 k8s clusters.
AWS support is so nice, you can't believe it.
Nah, I don't trust Google with anything. It's a scam. Google's support is horrendous. They refer you to idiots that drag you through calls until your will for life dies. And you're back to the mercy of some lost engineer that may comment on a github issue you opened 20 days ago. We have a bug reported back in 2020 that got closed recently without any action because it became stale and the API changed so much it doesn't really matter. It's that bad.
The billing day is a monthly reminder you're paying entitled devs to do subpar work other companies do a lot better.
No, we don't miss them already.
This reminds me of the fond days of having weekly customers calls. We develop AWS services, and we answer our customer-support calls directly. No middle man. Just techies to techies. And we made promises to customers on the fly, and customers sometimes project managed us.
This! They even custom-coded their support portal better than those off-the-shelf vendor like Zendesk. I say this as a Zendesk paying customer.
GCP on the other hand, is a F-tier in support. Almost feel like I need to beg them to get any level of help.
We booked a follow up call in the calendar, I spent good time preparing my notes and requirements for the meeting... and then nobody on their side showed up or contacted me again.
I'm of the opinion that focused products created by smaller teams are better.
I think Google Cloud needs someone like Jeff Bezos as their head: Look what your customers actually want and need and understand their requirements. And they usually want good customer support and want a competent key account manager as well.
When we were looking to migrate our analytics database from on-premise to a cloud alternative we were looking at BigQuery and Snowflake. BigQuery is a great product and we were already deeply invested in GCP as well. However the GCP sales team just couldn't sell BigQuery - they just don't know what old corporations want to hear in a sales pitch. So we went with Snowflake in the end. Not because it's the better product but because their sales team is better.
I'm not sure if the cloud business is actually a priority at Google. If it is then I think they don't understand the mistrust Google is facing when it comes to stable long term support of their products.
TL;DR: they broke something and wouldn't fix it.
Recently was woken up by alert about DNS resolution issues.
GCP rolled out new version of SkyDNS and NodeLocalDNS, SkyDNS reports 99% miss, had to quickly hack it.
This is not the „out-of-the-box” experience you want to have.
Which is on brand with Google. They have no problem launching stuff, and no problem killing stuff. But man, then just get out of the cloud business and focus on what you're good at.
I wonder what makes us different, I work in europe on video games; AWS’s handling of me when I was at Ubisoft left a really sour taste - when I moved into Tencent/Sharkmob I tried really hard to love AWS as it was the defacto industry standard and instead I was left with a feeling that most of it is inconsistent garbage papered over with lambda functions. I referred to these weird gotchas as “3am topics”; things that I don't have the mental capacity to deal with at 3am and convinced the studio to switch to GCP- which, incidentally they are still extremely grateful to me for doing.
This sounds more like an indictment of the system design than the cloud provider.
What are some of these “3am” topics that made GCP a better choice?
1) having the project/account your in visible at the top at all times.
We used SSO for “accounts” which is AWS’s way of completely separating resources; the long string that is returned is not unique in the start and the remainder is cut off: so all accounts/projects looked the same, was impossible to tell at a glance if you were in dev, staging or prod.
2) Autoscaling groups with that had human readable incrementing “names”, in AWS instances have hex slugs as instance names and you can give an instance a special “Name” label: but any new machines created with an ASG will just reuse the same name label making them hard or impossible to tell apart.
The AWS official solution for this is to have a lambda function hook on the scale event and give your new node an incremented name label. Given that AWS is pricy to save me time: I do not personally consider this an elegant solution.
3) having all regions on one page.
We spent €6,000~ on a database we didn't know about until we started digging into the bill. Not knowing what resources are available at a glance feels pretty basic to me tbh.
4) the network implementation overall; in Google you can just make a network and it will work without having to mess with zone routing and configuration of that which is put on the user.
If it’s on the user, it’s a variable that has to be checked during an outage; it is terraform code that has to be grokked and so-on.
Why were you even messing with the instance name? This is a ridiculously simple problem to solve with tags on your ASG. And AWS even did the courtesy of propagating those tags across the ASG and all its instances.
https://docs.aws.amazon.com/autoscaling/ec2/userguide/ec2-au...
I agree that this is an annoying issue in the AWS web console.
I assume this is something that could be fixed on your end by a little bit of CSS.
https://docs.aws.amazon.com/IAM/latest/UserGuide/console_acc...
pros: cannot get accounts mixed up.
cons: All sessions are actually 12hr sessions (ASIA not AKIA) and no access to perm keys for cli, security i suppose. Its not too bad though as TVM gives creds for various use cases.
GAE uptime improved, for a little while. Yeah, we're on AWS now too.
My old team was building a system that was half-GCP and half-Borg, and we had to write our own (extremely bad) Cloud Spanner fake for use in tests. In contrast, Infra Spanner is extremely well supported for tests. Same with BigQuery vs Dremel and many other systems.
Instead many aspects of GCP's management console are handled by different internal tools, often command line driven. IME they are often far more unwieldy than GCP.
Sometimes this makes sense (far tighter access controls and configuration change controls than a typical company), and some times it's just because of legacy ways of doing things.
I worked on a team at Google that used the internal GCP to serve some code/content for a specific feature, and it was in some ways it was more frustrating than using just either the normal internal systems or just vanilla GCP.
My team has a sort of a sandbox where we can use almost any Azure product we want (our IT is supportive and permissive as far as that sandbox goes, which is a blessing), but even then it's just painful in comparison.
AWS is probably better though.
It's actually sort of ridiculous. AWS has the best support I have ever interacted with. I mean, our org certainly pays enough for it but it's so completely unusual in tech, or really any sector to get great support even when you're paying for it.
Every time, I am also proven wrong as someone competent on their side both actually understands my issue and finds a resolution.
One of my pet hates is the (ab)use by repo maintainers of the auto-close-when-stale feature on Github.
What useful purpose does it serve beyond making the repo maintainers look good because they have a low number of open issues ?
It doesn't actually address the issue. Its the virtual equivalent of brushing under the carpet.
Azure is a dumpster fire from the ground up.
My experience has been the opposite though not without issues, Azure has some of the best corporate and security features of any cloud and it's only getting better. The zero trust model fits in so nicely with their identity platforms it's a sight to behold compared to other cloud providers which likely use some form of AAD or AD DS anyway.
Their support is responsive and they seem to know what they're talking about. (AKS)
Please provide some specifics on your experience?
I have friends forced to use Azure and they routinely report issues with provisioning resources, things taking a very long time to spin up or simply being rejected because Azure doesn't have any capacity.
Functions use containers under the hood. Each invocation created a new container, and when enough of them ran long enough, the host disk would fill up. (Pretty sure our workload wrote almost nothing to disk.)
An internal Azure disk clean-up routine kicked in, which deleted image layers for running Functions. This deleted the filesystems for containers that were still running, yanking them out from underneath the running processes. It also meant the host couldn't launch new instances of our Functions.
At this point the host was poisoned and couldn't launch any new work, even after the workload was reduced, It had to be terminated and replaced, after we detected the problem manually.
Azure support never weemed to take the problem seriously, and after we migrated our workload off of Functions they decided the problem must be resolved since we weren't complaining anymore.
Our reason for going all in with GCP was the k8s. We've been using GCP for 2+ years. The trouble we have is with stability and so many of the features being constantly rolled out.
Our experience was that K8s cost more on GCP than AWS.
Just on LoadBalancers alone, you have tons of tricks that are specific to GCP implementation. And we needed a few extra because you couldn't run all the features we wanted on 1-2 per cluster. For example, we have a 3rd party that required all our requests to always originate and respond back from a fixed IP address. We could only pick one not a range, not a list. This was a hard requirement. The service was important so we had to do it.
It took our team several days to find how to do it using online documentation and support. Tech support was useless. We had one guy in our team that spent 2 days on the phone with a paid, local GCP implementation partner trying to get this problem sorted. Nothing came out of it other than being pitched on our dime a lot of services and architecture we didn't need. Eventually we figure it out on our own. I don't even remember speaking about this when we transitioned to AWS.
This comparison seems to be not exactly fair? Amazon’s 126 million queries per second was purely for Amazon-related services serving Prime Day generating this on DynamoDB, and not all of AWS is my read.
What would have perhaps been a more fair comparison is to share the peak load that Google services running Cloud Spanner, and not the sum of all Spanner services across all of GCP and all of Google (Spanner on non-GCP infra).
I will say that it would show a massive of confidence to say that Photos, Gmail and Ads heavily rely on GCP infra: which would be brand new information for me! It would add to confidence to learn more on how they use it, and if Cloud Spanner is on the critical path for those services.
What is confusing, however, is how in this article "Cloud Spanner" is consistently used... except for when talking about Gmail, Ads and Photos, where it's stated that "Spanner" is used by these products, not "Cloud Spanner!". Like if they were not using the Cloud Spanner infra, but their own. It would help to know what is the case, and what the load of Cloud Spanner is: and not Spanner running on internal Google infra that is not GCP.
At Amazon, practically every service is built on top of AWS - a proper vote of confidence! - and my impression was that GCP had historically been far less utilised by Google for their own services. Even in this post, I'm still confused and unable to tell if those Google products listed use Cloud Spanner or their own infra running Spanner.
It’s a pretty big deal if Gmail migrated to GCP-provided Spanner(not to an internal Spanner instance) and sounds like he kind of vote of confidence GCP and Cloud Spanner could benefit from: might I suggest to write about it? It’s easier to digest and harder to miss than an hour-long keynote video with no time stamps.
And so just to confirm: Gmail is on Cloud Spanner for the backend?
https://www.youtube.com/live/268jdNwH6AM?si=WkgnvqaIwFidt-hc...
I don't think they would've migrated again to GCP Spanner (even if it would've been a show of faith).
When I worked at Google I tried to get more services to migrate to the cloud but the internal environment that was built up over 25 years is much better at supporting billion+ users with private data.
It doesn't make much sense to have a 'better' version of a product you sell but keep it internal.
Source: I was on the last team running our own [[redacted]].
link with time-stamp:
This wasn't the first time Gmail has replaced the storage backend in-flight. The last time, around 2011, they didn't hype it up, they called it "a storage software update" in public comms. And that other migration is the origin of the term "spannacle", because during that migration the accounts that resisted moving from [[redacted]] to [[redacted]] we called barnacles.
I'm not sure? I guess I'm mostly not sure what "gcp infra" means there. The blog post says
"Spanner is used ubiquitously inside of Google, supporting services such as; Ads, Gmail and Photos."
But there's google-internal spanner, and gcp spanner. A service using spanner at Google isn't necessarily using gcp. (No clue about photos, Gmail, etc)
Granted, from what I gather, there's a lot more similarity between spanner & gcp spanner than e.g. borg and kubernetes.
I really want to give Google the benefit of the doubt: but it doesn't help that they did not write that eg Gmail is using "Cloud Spanner." They wrote that it uses Spanner.
Apparently Cloud Spanner doesn't support protobuf columns? It would be hard for any internal Google product to use it under that restriction.
gcp spanner and normal spanner are different deployments of the same code.
Which can be the difference between 99.99% availability and 99% availability with data corruption issues. Not saying that's the case here but one should not downplay the difference deployments can make.
There are many engineering directors at Google.
swish
Not necessarily about volume of transactions, but this is similar to one of my pet-peeves with statements that use aggregated numbers of compute power.
"Our system has great performance, dealing 5 billion requests per second" means nothing if you don't break down how many RPS per instance of compute unit (e.g. CPU).
Scales of performance are relative, and on a distributed architecture, most systems can scale just by throwing more compute power.
Probably dealing with thousands of requests per seconds, but wants to say they're building something that can scale to billions of requests per second to justify their choices, so there they go.
Large parts of AWS itself uses DDB - both control plane and data plane. For instance, every message sent to AWS IoT will internally translate to multiple calls to DDB (reads and writes) as the message flows through the different parts of the system. IoT itself is millions of RPS and that is just one small-ish AWS service.
Source: Worked at AWS for 12 years.
> DynamoDB powers multiple high-traffic Amazon properties and systems including Alexa, the Amazon.com sites, and all Amazon fulfillment centers. Over the course of Prime Day, these sources made trillions of calls to the DynamoDB API. DynamoDB maintained high availability while delivering single-digit millisecond responses and peaking at 126 million requests per second.
Amazon was very, very clear on this. For Google to use that number without the caveat is just completely underhanded and dishonest. Whoever wrote this is absolutely lacking in integrity.
Worst case scenario, it's Google you're buying, not a random startup etc.
If Amazon said they didn't use DAX that day I would say they were lying.
The average consumer or startup is not going to squeeze out the performance of Dynamo that AWS is claiming that they have achieved.
In fact, it might have been fairer in Ruby if they didn't hard-code the net client (Net/HTTP). I imagine performance could have been boosted by injecting an alternative.
I don’t have p99 times in front of me right this second but it’s definitely lower than 20ms for reads and likely lower for writes. (EC2 in VPC).
This sounds like you're blaming dynamo for you/your stack's inability to handle connections / connection pooling.
I am running https://cloud-canary.com a service where I monitor AWS primary services for latency and availability.
It comes with a lot of data.
For instance this is the latency I see doing operations against Dynamo.
https://cloudcanary.grafana.net/public-dashboards/c53e2092d6...
I'll definitely update it!
Numbers without units are dangerous in my opinion.
Little bit of well meaning advice: This needs copy editing -- inconsistent use of periods, typos, grammar. Little crap that doesn't matter in the big picture, but will block some from opening their wallets. :) ("OpenTeletry", "performances", etc.)
All in all this is quite cool, and I hope you get some customers and gather more data! (a 4k object size in S3 doesn't make sense to measure, but 1MB might be interesting. Also, check out HDRHistogram, it might be relevant to your interests)
Any feedback is appreciated!
I pick 4k as a no-op against S3, something that very little time but still does some work.
I will definitely consider to increase it!
That said, I can’t imagine these numbers mean much to anyone after a certain point. It’s not like either company is running a single service handling them. The scale is limited by their budget and access to servers because my traffic shouldn’t impact yours. I feel like the better number is RPS/QPS per table or per logical database or whatever.
There's no indication that google is talking about ALL of spanner either? The examples they list are all internal google services, and they specifically say "inside google".
I'm also dubious that even with all of the AWS usage accounted for that DynamoDB tops Spanner if Amazon themselves are only at 126 million queries per second on Prime Day.
Not only this, but practically most, if not all, of the AWS services use DynamoDB, including use cases that are usually not for databases, such as multi-tenant job queues (just search "Database as a Queue" to get the sentiment). In fact, it is really really hard to use any relational DB in AWS. I mean, a team would have to go through a CEO approval to get exceptions, which says a lot about the robustness of DDB.
Edit: It's possible you're limiting your statement specifically to AWS teams, which would make it more accurate, but I read the use of "Amazon" in the quote you were replying to as including things like retail as well, etc.
is that true finally? It sure wasn't in the 2020-2021 timeframe.
Yet I still see the very deep stack of technically incapable middle manager sorts dutifully posting "come join us" nonsense on LinkedIn.
(I had the luxury of having worked in one of the inner sanctums of Apple hardware for years prior, so was immune to nonsense, and didn't need the job.)
AWS and GCP (and Azure, and Oracle cloud, and bare Kubernetes via an operator, and...) support Postgres really well. Just...use Postgres.
I think they're directly comparable with this context.
The entire point is that every cloud provider has a managed postgres offering, and there's no vendor lock-in. Though, technically, Dynamo does have a docker image you could run in other cloud providers if it came down to that, you'd get no support for it.
There's a reason why Google installed and built their own atomic clocks and put them in their datacenters, it is to facilitate global timekeeping for this type of services. Most likely 99.9% of the time this type of database is overkill, and also likely way more expensive than you need.
https://cloud.google.com/spanner/docs/true-time-external-con...
I think just doing some (not necessary millions) ACID transactions over the globe and have consistent DB is strong value proposition even for small users.
There are a couple of databases out there with ddb compatible interfaces, like scylladb
I totally believe you, I just can't see how it becomes easier than chucking a container on Fargate or something. Maybe I've just been scarred by lambda rat's nests in the past.
CDK for everything else.
EC2-backed ECS has a great use case for things that you can run ephemerally in a container but require a persistent data store.
Seems like a leaner setup than using ECS/Fargate + LBs to me. Have I overlooked something?
Meanwhile at work I have a cowoker who loves to create AWS soup where they use an assortment of lambdas/api gateways/sqs queues/sns topics to accomplish tasks such as taking files from one s3 bucket and putting them in another s3 bucket owned by a different team. Their justification of this was that it was generic so other teams could use it, but it is a pain to maintain and make changes to.
ok? and sqlite3 in memory is even cheaper than postgres!
if you can use (and support correctly) postgres then you should use it, obviously there's no point using a globally scalable P-level database if you can just fit all your data on one machine with posthgres.
If I'm writing a chat app with millions of messages and very little in the way of "relationships", should I use Postgres or some flavor of NoSQL? Honest question.
You do have to think about how to model the data in each system, but there are very few cases IMO where one is strictly 'better.'
PostgreSQL is extremely good at append-mostly data, i.e like a chat log and has powerful partitioning features that allow you to keep said chat logs for quite some time (with some caveats) while keeping queries fast.
Generally speaking though PostgreSQL has powerful features for pretty much every workload, hence the Golden Rule.
Normally you'd make a decision like this by figuring out what your peak demand is going to look like, what your latency requirements are, how distributed are the parties, how are you handling attachments, what social graph features will you offer, what's acceptable for message dropping, what is historical retention going to look like...[continues for 10 pages]
But if you don't have anything like that, just use something simple and ergonomic, and focus on getting your first few users. There's a long gap between when the simple choice will stop scaling and those first few users.
I'm using SenseDeep's OneTable which was pretty interesting to learn https://doc.onetable.io/ in case others reading this are curious.
example=> CREATE TABLE tab(k int PRIMARY KEY, data jsonb NOT NULL);
CREATE TABLE
We can fill this with heterogeneous values: example=> INSERT INTO tab(k, data) SELECT i, format('{"mod":%s, "v%s":true}', i % 1000, i)::jsonb FROM generate_series(1,10000) q(i);
INSERT 0 10000
example=> INSERT INTO tab(k, data) SELECT i, '{"different":"abc"}'::jsonb FROM generate_series(10001,20000) q(i);
INSERT 0 10000
Now, keys in the range 1–10000 correspond to values with a JSON key "mod". We can create an index on that property of the JSON object: example=> CREATE INDEX idx ON tab((data->'mod'));
CREATE INDEX
Then, we can query over it: example=> SELECT k, data FROM tab WHERE data->'mod' = '7';
k | data
------+---------------------------
7 | {"v7": true, "mod": 7}
1007 | {"mod": 7, "v1007": true}
2007 | {"mod": 7, "v2007": true}
3007 | {"mod": 7, "v3007": true}
4007 | {"mod": 7, "v4007": true}
5007 | {"mod": 7, "v5007": true}
6007 | {"mod": 7, "v6007": true}
7007 | {"mod": 7, "v7007": true}
8007 | {"mod": 7, "v8007": true}
9007 | {"mod": 7, "v9007": true}
(10 rows)
And we can check that the query is indexed, and only ever reads 10 rows: example=> EXPLAIN ANALYZE SELECT k, data FROM tab WHERE data->'mod' = '7';
QUERY PLAN
---------------------------------------------------------------------------------------------------------------
Bitmap Heap Scan on tab (cost=5.06..157.71 rows=100 width=40) (actual time=0.035..0.052 rows=10 loops=1)
Recheck Cond: ((data -> 'mod'::text) = '7'::jsonb)
Heap Blocks: exact=10
-> Bitmap Index Scan on idx (cost=0.00..5.04 rows=100 width=0) (actual time=0.026..0.027 rows=10 loops=1)
Index Cond: ((data -> 'mod'::text) = '7'::jsonb)
Planning Time: 0.086 ms
Execution Time: 0.078 ms
If we did not have an index, the query would be slower: example=> DROP INDEX idx;
DROP INDEX
example=> EXPLAIN ANALYZE SELECT k, data FROM tab WHERE data->'mod' = '7';
QUERY PLAN
---------------------------------------------------------------------------------------------------
Seq Scan on tab (cost=0.00..467.00 rows=100 width=34) (actual time=0.019..9.968 rows=10 loops=1)
Filter: ((data -> 'mod'::text) = '7'::jsonb)
Rows Removed by Filter: 19990
Planning Time: 0.157 ms
Execution Time: 9.989 ms
Hence, "arbitrary indices on derived functions of your JSONB data". So the query is fast, and there's no problem with the JSON shapes of `data` being different for different rows.See docs for expression indices: https://www.postgresql.org/docs/16/indexes-expressional.html
As with all data storage, the question is usually how do you want to access that data. I don't have experience with Postgres, but a lot of (older) experience with MySQL, and MySQL makes a pretty reasonable key-value storage engine, so I'd expect Postgres to do ok at that too.
I'm a big fan of pushing the messages to the clients, so the server is only holding messages in transit. Each client won't typically have millions of messages or even close, so you have freedom to store things how you want there, and the servers have more of a queue per user than a database --- but you can use a RDBMS as a queue if you want, especially if you have more important things to work on.
The super-power of Postgres is that it supports everything. It's a best-in-class relational database, but it's also a decent key-value store, it's a decent full-text search engine, it's a decent vector database, it's a decent analytics engine. So if there's a chance you want to do something else, Postgres can act as a one-stop-shop and doesn't suck at anything but horizontal scaling. With partitioning improving, you can deal with that pretty well.
If you're writing fresh, there is basically no reason not to use Postgres to start with. It's only when you already know your scale won't work with Postgres that you should reach for a specialized database. And if you think you know because of published wisdom, I'd recommend you set up your own little benchmark, generate the volume of data you want to support, and then query it with Postgres and see if that is fast enough for you. It probably will be.
My advise to you is to use Postgresql or, heck, don't over think it, sqlite if it helps you get a MVP done sooner. Do NOT prematurely optimize your architecture. Whatever choice results in you spending less time thinking about this now is the right choice.
In the unlikely event you someday have to deal with billions of messages and scaling problems, a great problem to have, there are people like me who are eager to help in exchange for money.
Lots of people like to throw around the term "big data" just like lots of people incorrectly think that just because google or amazon need XYZ solution that they too need XYZ solution. Lots of people are wrong.
If there exists a motherboard that money can buy, where your entire dataset fits in RAM, it's not "big data".
I think the truth is that you should use the simplest, most effective tech possible until you are absolutely certain you need something more niche.
Once you have millions of messages, maybe consider moving the data intensive parts out if postgres, if necessary.
The criticism is often that people look for big data solutions, before they have big data.
If you scale out of postgres, you probably have enough users and money that you can fix it :)
But moving to a NoSQL before you have to, might just slow down development velocity -- also you haven't yet learned what patterns users have.
Having said that I guess I broadly agree with your comment. It seems like a lot of people like to plan for massive scale while they have a handful of actual users.
In the first case, the database created as many problems as it solved (which is true of any large application running at scale; your data store will _always_ be suboptimal). A fancy, expensive NoSQL database won't save you from solving hard engineering problems. At smaller scales (on the order of tens-hundreds of RPS), it's hard to go wrong with any established SQL (or open source NoSQL if that floats your boat) database, and IMO Postgres is the most stable and best bang for your engineering buck feature wise.
Using Spanner is giving up a lot for the scalability, and if you ever reach the scale where a single node DB doesn't make sense anymore, I don't know if Spanner is still the answer, let alone Spanner with your old design still intact. For one, Postgres has scaling options like Citus. Or maybe you don't need a scalable DB even at scale, cause you shard at a higher layer instead.
At a previous job, we ended up creating a very complicated write through cache system in front of spanner that dynamically added memory/CPU capacity as needed to prevent hot shards; our application was extremely read heavy, and writes were relatively low RPS, so this ended up working OK, but we were paying tens of thousands of dollars a month for Spanner plus tens of thousands of dollars a month for all the compute sitting in front of it. I don't think we ended up doing much better than if we had bitten the bullet and run clustered Postgres because our write volume ended up being just a few hundred RPS, even though the read volume was 1000x that. Postgres behind this cache system would have handled the load just as well and cost less than half as much.
The other thing that frustrates me personally about Spanner is that Google's docs are incomplete (as usual); there are lots of performance gotchas like this that exist throughout the entire service, and they aren't clearly documented (unlike, to their credit, AWS with Dynamo, who explains this entire problem very clearly and has an [expensive] prebuilt solution for it in the form of the DynamoDB accelerator).
Also https://www.postgresql.org/support/professional_hosting/
vs AWS Free Tier:
"25 GB of data storage ... 2.5 million stream read requests ..."
https://aws.amazon.com/dynamodb/pricing/
So, there's probably somewhere the lines on the graph cross, but Google's headline seems misleading.
Edit: actualy Spanner looks like another CockroachDB. You use sql to interact with it. In which case I can see many people who would want to use this with a free tier for hobby projects. ie. in between education and production development.
Pedantically, cockroachDB is another spanner. It was made by Google devs who left Google having previously used spanner, and intentionally made something similar to spanner (ish, lots of handwaving happening here)
Yeah it does :). CockroachDB set out to be an open source version of Spanner by ex-Google engineers.
lol
cockroachdb is an external reimplementation of some of the ideas of spanner but without depending on excellent clocks.
Ha. Remember Gary Bernhardt of WAT fame? https://twitter.com/garybernhardt/status/600783770925420546
> Consulting service: you bring your big data problems to me, I say "your data set fits in RAM", you pay me $10,000 for saving you $500,000.
Columnar databases typically get a 10:1 compression ratio over raw data = 240 TB effectively.
That’s a lot of data.
If you squint, any database engine is “in memory” if there is more buffer than data.
Or just use a RAM disk!
That is sadly not true, I remember one lonely night debugging a MSSQL 2012 instance that was _very_ slow, and it turned out that for a simple query (one join, 100 rows in one table and 10 in the other, 100 result in total, one where clause) it forced writing the result to disk before evaluating the WHERE condition. Unable to fight the scheduler I've ended up making a ramdisk for this data.
"Google should offer intro discounts" is IMHO a very valid point (absolutely no idea why this doesn't exist), but it doesn't really speak to whether or not the real product is more or less expensive.
Those are more minor services in the long run, but it makes me a little nervous to go in again on Google for a critical service. Before I invest my time and effort into using it I have to ask myself "Will Google someday sell off or end the cloud spanner service? Will I be in trouble if they do so?".
https://steve-yegge.medium.com/dear-google-cloud-your-deprec...
The products/services that “Google” the search company launches are different than “Google Cloud”. While the discontinuation of Google products is annoying it has nothing to do with Google Cloud products/services. I don’t think Google Cloud abruptly announces discontinuing products/services as they have paid customers.
Regarding Google Domains that is a Google product. The equivalent product from Google is “Google Cloud Domains” which is available to Google Cloud customers.
Also sold: https://cloud.google.com/domains/docs/faq
How did Google become like this?
And I guess this is the kind of marketing that attracts the customers they want.
(But I bet that Spanner is much easier than DynamoDB to develop with...)
I built a system that relies on a high-performance database and tested with both AWS DynamoDB and Google Cloud Spanner (see disclaimer) and was able to scale Google Cloud Spanner much higher than AWS DynamoDB.
DynamoDB is limited to 1000 WRUs per node, and there isn't an obvious way to get more than 100 nodes per table, so you're limited to 100,000 WRUs per table (= 102400000 bytes/sec = 97 MiB/sec = 776 Mib/sec) -- even if you reserve more than 100,000 WRUs in capacity for the table. The obvious workaround would be to shard the data across multiple tables, but that would have made the software more difficult to use.
Google Cloud Spanner was able to do much more than 97 MiB/sec in traffic (though the exact amount isn't yet public), and also was capable of much larger transactions (100 MiB versus DynamoDB's 25 (now it is 100) items * 400KiB of ~10 MiB) which was a bonus.
Disclaimer: The work was funded by a former Google CEO and I worked with the Google Spanner team on setting it up, while I am a former AWS employee I didn't work with AWS on the DynamoDB part of it, though I did normal quota adjustments.
By node you mean partition, and a partition is limited to 10GB. Store more than 1TB and you will have at least 100 partitions
https://aws.amazon.com/about-aws/whats-new/2019/05/amazon-dy...
I don't know about "got bored of"; I'd say more "effectively deprecated, using increasing costs† as an implicit push toward rewriting your service for more modern parts of their platform."
Specifically, Google want you to rewrite your GAE apps for Cloud Run (https://cloud.google.com/appengine/migration-center/run/comp...):
> Cloud Run is the latest evolution of Google Cloud Serverless, building on the experience of running App Engine for more than a decade. Cloud Run runs on much of the same infrastructure as App Engine standard environment, so there are many similarities between these two platforms.
> Cloud Run is designed to improve upon the App Engine experience, incorporating many of the best features of both App Engine standard environment and App Engine flexible environment. Cloud Run services can handle the same workloads as App Engine services, but Cloud Run offers customers much more flexibility in implementing these services. This flexibility, along with improved integrations with both Google Cloud and third-party services, also enables Cloud Run to handle workloads that cannot run on App Engine.
Anyone who's still on GAE (rather than having moved over to Cloud Run) at this point is a "legacy enterprise customer"; and so Google have at this point moved GAE pricing beyond just a monetary disincentive to use, to being "fired-customer pricing" — i.e. the price you charge when you don't really want to work with a customer any more, a price that says "go away", but if they still want to pay you even at that price-point, then sure, why not?
BigQuery initially was a lot more powerful, then they started adding a bunch of resource limits on queries that you could only overcome by paying up. Not a direct fee increase, rather you paid the same fee for a worse product.
As far as I know, AWS only ever decreases prices on services for example
https://techcrunch.com/2022/03/14/inflation-is-real-google-c...
Discussion: https://news.ycombinator.com/item?id=30671997
I don't know if Google App Engine falls under Google Cloud services, but either way that's just a technicality; the sentiment remains the same.
Edit: more information here: https://github.com/stickfigure/blog/wiki/The-Unofficial-Goog...
https://aws.amazon.com/blogs/aws/new-aws-public-ipv4-address...
We are introducing a new charge for public IPv4 addresses. Effective February 1, 2024 there will be a charge of $0.005 per IP per hour for all public IPv4 addresses, whether attached to a service or not
And for 65/month, you can get a VERY beefy Hetzner server.
You’ll have to wade through the crazy thicket that is the menu of cloud offerings.
I gave that one look, and decided I might as well give up and learn the basics of Linux admin once and apply it for life.
Comparing Postgres to Spanner is kind of like comparing a delivery van to a train. The train is always going to have higher overhead costs.
Linux admin is a useful skill, but I know my Linux admin skills can’t compete with the reliability, availability, and scalability of cloud systems… like Dynamo, S3, Spanner, etc.
Linux admin + hosting a server = my data, on my terms, until I keel over, and possibly long after that.
- in the next few decades, my Linux servers will have been updated completely multiple times
- software updates happen on my schedule and at my behest
- I can move to newer hardware whenever the mood strikes me
- I maintain full de jure and de facto ownership of my data (AKA I control it completely)
- Since I own the data, I can always upload it to some vendor in future. Due to vendor lock-in, non-standard data formats, and my least favourite: data egress fees, it's not straightforward to go from a vendor to another vendor, or from a vendor to DIY. I maintain maximum optionality
- Since I committed to the private server path, I can take full advantage of the server being a general computing device. I can combine web-hosting, databases, and other things on the same device / a stable of devices. I end up having ridiculous performance, full control of my entire stack, and at a huge discount, and it's a very simple system.
Security concerns are addressed in a couple of ways:
- By having everything on one server, or by architecting things just so, I can stand up a database that does everything I need, including serving my web-apps, without ever facing the public internet directly.
- Maintaining a secure server is admittedly more of an ongoing chore, but it's not a significant timesink at all
- Every online service by AWS et al ultimately runs on a server much like mine, so if there's some serious widespread Linux vulnerability, it'll affect managed services just as much as my server.
- The managed services themselves are not only juicy targets but are themselves vulnerable to both hacking and phishing. I'm convinced SSH'ing into Postgres + Linux is a safer option than a more complicated structure.
All of the above assumes my apps will never be planet-scale, which even in the most bullish case, they never need to be.
What an asshole.
Yeah I think this raises an underappreciated drawback of working on heavily AWS/GCP native projects. So much of the time ends up being spent on service level config and troubleshooting that has little relevance elsewhere.
don't bother with mongo or mysql or dynamo or cassandra or bigtable or spanner or ... until your lack of profitability or size means you can't afford to just use postgres.
if I have 10TB of hot data, can I afford two machines with 10T of RAM each? how about 100T?
> I thought postgres cost per query in many cases is cheaper than competitors.
that's not really a useful metric without size/latency/etc attached to it, being cheap for 0.1qps might be fine for a YC company but that's no good for my successful company etc
Could someone share a use case where you truly benefited migrating from Postgres/MySQL to Spanner/Citus/Cockroach, and there was no better solution? I'd like my hunch to be wrong.
> 7. Benchmarking. Customer may conduct benchmark tests of the Services (each a "Test"). Customer may only publicly disclose the results of such Tests if (a) the public disclosure includes all necessary information to replicate the Tests, and (b) Customer allows Google to conduct benchmark tests of Customer's publicly available products or services and publicly disclose the results of such tests. Notwithstanding the foregoing, Customer may not do either of the following on behalf of a hyperscale public cloud provider without Google's prior written consent: (i) conduct (directly or through a third party) any Test or (ii) disclose the results of any such Test.
It looks like that is less restrictive than it used to be. I found a blog post from last year mentioning an additional requirement to obtain Google's prior written consent prior to publishing (for all customers, not just fellow cloud providers) which no longer is included. https://cube.dev/blog/dewitt-clause-or-can-you-benchmark-a-d...
I think the only thing I pay for is API Gateway, and that's around $2 a month max.
• Spanner is more in the domain of CockroachDB — distributed strong consistency and ACID compliance. Both are ANSI SQL.
• DynamoDB and ScyllaDB are in the NoSQL key-value store / wide column domain. ScyllaDB is API compatible with DynamoDB, but is also API compatible with Cassandra Query Language (CQL). These databases are for eventual consistency use cases.
As to the "99%" comment, there are over 415 databases currently tracked on DB-engines.com (https://db-engines.com/en/ranking). It is a very competitive and specialized industry. Very few databases have more than 1% marketshare these days.
According to this: https://6sense.com/tech/relational-databases
...only the top 9 options show >1% of marketshare.
MySQL being so far in the lead (42%) is far more likely due to it being baked into so many OEM deals (over 2,000 as per https://www.mysql.com/oem/). Like, every WordPress site in the world than anything else.
Yet that doesn't really make MySQL great for every use cases. It has obvious limitations.
We also have to note that different studies have far different results.
In this video: https://statisticsanddata.org/data/the-most-popular-database...
It shows a different set of top databases, and has 11 of them showing >2% of marketshare. It has MySQL at only 13.91% of marketshare, with Oracle being on top at 31.19%.
The source of this poll is TOPDB Top Database index:
https://pypl.github.io/DB.html
It's not based on money marketshare, nor on poll results, but simply is counting how often a database is searched on Google. Note that anything that could garner 1% of the market would suddenly catapult that database into the top 16th place on the list.
At the end of the day, database "market share" — however it is calculated — is not representative of exclusive percentages. Many companies run more than 1 type of database. There might even be multiple SQL and NoSQL databases, each designed for purpose, at large enterprises. OLTP, OLAP. I would not be surprisedi in the least to learn that large corporations had well over 100 different databases running at any given time.
So, back to your basic comment: if a database is useful for even 1% of the use cases, that would still place it in the top 20 databases in the industry.
Moreover, a lowest-common denominator database will not be able to support many critical use cases where you really do need a specialized data model, or index type, or query language, or workload type, or latency, or distribution, or scale of QPS/TPS/OPS, or total data size or query payload size, etc.
I have likened the database industry to the state of Christianity after the Reformation. Even if there are still plenty of "Catholics" (classic SQL adherent, like Oracle), there are also Reformed faiths ("NewSQL" / "Distributed SQL", e.g., PostgreSQL), plus any number of Protestant reformation alternatives (NoSQL). I'm not sure what the "Orthodox" in this analogy refers to. Maybe SQL data warehouses? The schism between OLTP and OLAP? Dunno.
Anyway, it's an analogy. Bound to break down at some point.
At the end of the day, my advice: use the database most closely aligned to the dataset and workload and use case you have at hand. Don't just throw something at a problem because it's a popular choice, or because it's the database you happen to be most familiar with.
That GCP interface is cancer. You need a Google-to-English dictionary to understand it.
Google Maps is the key lesson here. 10 to 20 times price increase, just because someone had a meeting. No justification or coherent strategy, just "sorry here's a shiv to the gut".
At least with Amazon, as chaotic as it is, you know that they just mark stuff up to the margin they want and let the market sort out what gets used. Google is into secretive grand strategy that changes completely anytime 4 product managers get together in a room.
https://support.google.com/analytics/answer/11583528?hl=en
But the API to access the replacement GA4 data is still in Beta: (for dotnet / go / python and node):
It was very difficult to understand their thought process, and we were pretty much forced to find alternative services that had a more reasonable pricing structure. The amount they wanted to charge us for a few million geocode requests per month was bananas.
Was there really no value in keeping a customer like us around at a more reasonable rate?
Spanner is a pretty solid database, it checks all the boxes: consistency, geographic replication and transactional updates are not present at the same time in most large scale db's out there.
I think the "google shuts things down" trope is justified and something that I wince at every time I see something like this, but GCP isn't a random pet project that loses money and has no users. It's a very successful enterprise business.
(disclaimer: I work at google in cloud, so obviously I have personal bias, but this is just my opinion based on the pure numbers.)
Spanner is used for a lot of internal Google services, too. Spanner has basically been eating all the other databases Google uses internally. Nom nom nom, every database that used to be something else is now Spanner.