Discovering Azure's unannounced breaking change with Cosmos DB
metrist.io
metrist.io
If your traffic pattern is exactly right, and you always scale traffic up and never ever down and do not have spikes, I guess it is probably OK. The main problem is the docs are (or, at least were 2 years ago) not clear about all the caveats and restrictions but pretend it is a generic database that just works. So one has to discover all the caveats oneself.
Microsoft thinks the exact workings of the partitioning is something that should work so well you don't need to know it in detail. But, if your usecase is slightly off you end up really needing to know. I know at least one team who routinely copy all their data from one Cosmos instance to another and switch over traffic to the copy just to get a partitioning reset; it is one thing to have to do it; another to discover in production yourself it has to be done with no prior warning..
Also: The ipython+portal+Cosmos security meltdown from 1 1/2 years ago alone should be reason to look elsewhere.
(No, not a competitor, just have spent way way way too much engineering time moving first on and then off Cosmos and yes I am bitter)
It could have been easy. We could have used Postgres.
And then all that would have been left is just to make that Postgres in a globally distributed DB and manage it (with something like Citus). Postgres' native scaling out (not even talking about globally distributed) capabilities are basically nonexistent so you need third party tooling.
For you young uns, back in the 1990s Microsoft was so convinced that NTFS made file fragmentation impossible that they didn’t provide a way to defrag for a very long time.
FAT12/FAT16/FAT32's earliest-free-block-no-matter-what allocation method⁴ trained people to believe that fragmentation is a universally rampant problem, but with a better designed allocation strategy it really isn't.
----
[1] for instance simultaneously growing files on near-full volumes
[2] more so ext4 with delayed allocation turned on
[3] at least not ones without attendant "are you really sure you want to use this on data you care about?" warnings
[4] I wonder if anyone retrofitted a brighter heuristic into an implementation of these or exFAT. Could be a (not massively useful but) interesting learning exercise for a budding OS developer.
Can you share what you migrated onto and the results?
Cosmos is used because some tables we use are larger than 64TB. It’s decently useful for large chunks of data.
Cosmos is at best unfinished. I used Google Data Store for years and it was finished in the sense that it always did what it says on the tin and didn't cause lots of problems for you (although I think higher latency than Cosmos, so may not be for every usecase -- there is also BigTable and Spanner).
And I don't believe Google would ever have done something like deploying a shared multitenant Jupyter system with full access to all customer's DB. Microsoft actually did that.
As I said the main problem is mainly how Cosmos is marketed, as a general purpose NoSQL. If it was marketed with all the caveats you learn about after you go live, and people only used it if they really had to, it would be a bit different.
Though realistically it would drive most people to other cloud providers.
https://devblogs.microsoft.com/cosmosdb/distributed-postgres...
It’s worth noting how cloud vendors have fallen back on tried and true databases such as Postgres (mind you, often replacing the guts of those databases with new implementations.. but still).
Have they? Amazon created Aurora from scratch, with compatibility layers for different database engines (MySQL, Postgres), and GCP did the same with Spanner, which i wouldn't call "fallen back".
That said, you didn't mention things like RDS or Google Cloud SQL, which are cloud-based versions of standard open source dbs.
I cannot comprehend what organizational process lead not only to its creation, but also to its continued existence.
Is existence is a constant reminder that the road to hell is paved with good intentions, and that you better do your due diligence before adopting any piece of tech that’s not easy to replace. Even if it comes with a pedigree.
I wrote up some of the horrors stories a while back on HN: https://news.ycombinator.com/item?id=29295871
AWS DynamoDB used to do that. When I was working for a team in AWS we ended up discovering an unexpected bad hash choice, had a hot partition and ended up with a crazy number of partitions and a need to have really high r/w allocation on it. The only option at the time was to roll over to a new table (which we could do and fix our hash choice, thankfully, without too much hassle).
They fixed that stuff a few years ago now so most (all?) of my "here be dragons" concerns about DynamoDB have been addressed.
I’ve learned to stay away from services that are high abstractions. Huge lock in, expensive and much more risk of it being eol’d.
In our platform we’ve actually dropped SQL in favor of Az blob and table storage. It’s going to reduce our database costs by 5x.
We had a service that had a list API that was paginated. It returned a nextToken to specify the start of the next page of the results.
Internally we were doing a database migration to a completely different system and migrating one customer at a time. The problem was that if a customer was in the middle of a list call and had the next token with them which was generated from the previous database system, after migrating to the new database the older token would not be able to start from exactly where it should had the customer not been migrated. This was because it would not have all the information of the service to resume at the exact offset.
One option was to throw an error and let the customer retry the request; another option was to return some possibly duplicate items in the next page; none of these were good enough for both engineers and PMs and instead we decide to take up a bunch of additional work so that no customer would be impacted. This was 10+ weeks of additional work for the whole team but we did it because culturally it felt the right thing to do for the customer.
Note that the impact would have been tiny if at all. A customer would have to be in the middle of a paginated request and out migration system would have had to migrate that particular customer at that exact time and the impact would have been a few possibly duplicate items. But we didn’t know the actual impact of those temporary duplicated for a single call and we all agreed breaking changes like this are unexpected and cause customer to lose trust with us.
There are some Microsoft products I genuinely love, but some are terrible. To an extent it is a reflection of the inconsistency in internal teams. Culture, values, skill level, and quality bar are all over the place depending on who you talk to, even compared to other large companies.
From what I had heard, CosmosDB was not a healthy team, and I would not consider using it as a product.
I suspect the way Microsoft does interviewing and performance management (very local to the specific team) contributes to the inconsistency.
MSFT has also been fairly open to its employees that it does not try to compete with competitors like Google, Meta, or even Amazon, in terms of compensation. So it isn't really trying to get the best engineers, so long as it can continue to print money.
There are still folks there who are incredible, but the floor is shockingly low at times. Folks will self-select, so you will then get teams which are more homogeneously good or bad.
I think pay is certainly a factor in talent bleeding to G/Meta in general, which they refuse to address. I imagine this isn't unique to MSFT.
Perhaps it’s because Microsoft is still catching up with Azure, and as such prefers moving fast and occasionally breaking things?
This really feels like a bug to me, and probably didn’t trip monitors due to 2 reasons: 1) Given that portal does the right thing (and probably ARM template samples), this was very small percentage of traffic.
2) the failures would look like client side errors, making it less likely to trip monitors.
*PS my comment is not an official response (I don’t even remotely work on CosmosDB) but I’ll forward this internally
I'm finding a lot of the reliability guarantees of Azure PaaS services are overblown or come with big caveats when you start to work with them in a serious way. For example I've had some bad reliability issues with Azure Functions not firing, or their premium function runtimes becoming unresponsive. And it seems like that's just the start of the outstanding issues with them https://github.com/Azure/azure-functions-host/issues
I think people need to look more carefully at these PaaS guarantees and look at what that 99.999% reliability Microsoft are claiming actually means.
https://www.wiz.io/blog/chaosdb-explained-azures-cosmos-db-v... https://msrc-blog.microsoft.com/2021/08/27/update-on-vulnera...
How can this be a premium iaas/paas? Azure feels like the MS teams of tele conference. Companies buy in because they are already in the MS world. Not because azure is better.
> And stronger yet when the database is unusable due to an incident the cpu is maxed out and it doesnt allow any successful connection, nothing is detected
Apparently Azure's storage system that backs this uses some sort of thread pool and the thread pool can lock up/become exhausted leading to I/O starvation. When this happens, connection attempts fail. When the connection attempts fail, it can lead to a connection storm where all these new connections rolling in exhaust the CPU. The telltale indicator is Postgres checkpoints getting behind.
All the while, the DB I/O metrics look like they're completely fine because it's not hitting an I/O limit, it's hitting thread pool exhaustion in the some storage system under the instance, outside of Postgres.
You can also get some clues if this is the problem by enabling Performance Insights and checking the Waits tab. If all the top waits are related to I/O activity, that's another dead giveaway the storage system is locked up again. You can just web search the name of the waits to see what causes them. AWS has some nice docs detailing Postgres waits
Since we have premium support (P1?), we had some internal azure postgresql engineer look at the issue and they pushed the problem back to us. Blaming our app not built correctly. That has been ping-ponging for over a year now.
Finally i saw this semi-acknowledgment in their health status yesterday.
Do you happen to know a proper solution? Are you waiting for them to fix this issue or moved to a different db service?
Perhaps the flexible server is better?
Personally, I'd look into a 3rd party if you want managed Postgres (assuming you don't have contractual obligations that might complicate 3rd party access). There's vendors like EnterpriseDB, Scalegrid, etc that provide various solutions (I don't have any recomendations here--Postgres has a list of managed providers by country https://www.postgresql.org/support/professional_hosting/nort...)
Absolutely agree on a third party. Azure is just a let down overall.
And the funny thing? Status.azure.com is all green. No events in activity overview. No service health within the affected instance.
Workaround advised by azure? Upgrade to next plan. We already reached the maximum size. Maybe time for Citus . More $$$ for M$
Hypercloud managed service SLAs: all the fun of novel complex, technical solutions in production + the transparency of cast iron + the pendanticism of being a contract lawyer
Which leaves exactly zero people who are excited to be at that intersection.
That's a couple months after the Ubuntu/systemd incident (Azure's "blessed" Linux image is Ubuntu and it has unatttended-upgrades enabled including on managed infrastructure like AKS (where you can't turn it off without dirty hacks). A bad Ubuntu update caused hosts to lose their DNS from DHCP config rendering massive amounts of machines in partially broken states)
https://thenewstack.io/ubuntu-linux-and-azure-dns-problem-gi...
It sounds cool, but I was surprised when after what I think should be the worst and dumbest security design flaw breach [0] there wasn’t much uproar.
I thought maybe no one is using it so there wasn’t much impact.
Pushing out breaking changes without telling your customers also gets explained by there not being any (or many since these folks found it) users.
Could you image how big of a deal it would be if a breaking change or elevated privs bug were in actually used products.
[0] https://www.techtarget.com/searchsecurity/news/252505973/Res...
Azure being a shade of blue, you should've called it "Blue Monday"[0]. Could've even rigged up something to play the song when integration tests mysteriously failed. How does it feel/ to treat me like you do?/ When you've laid your hands upon me/ and told me who you are?...
Alternatives? LOL
Basically just looking for geo-redundant, high read & write throughput. Our intention was to leverage Azure Event Grid/Kafka Connect to have event streaming used to coordinate writes between Redis (cache), Cosmos (transactional DB), and our systems of record (legacy). Majority of read/writes would occur via our API, but some would occur via the systems of record, hence the use of a log-based architecture.
Do you have any specific requirements for which cloud provider you use, or any particular interface you really need?
It seems like there are a lot of options for large scale analytics, but I don't know a lot for high throughout geo-redundant transaction processing.
Use Confluent and connect it with a managed postgres on aws or gcp. If you really need more, use Google Cloud Spanner.
So, let's say I'm woken up in the middle of the night because my black box database as a service suddenly returns errors. If I'm not incompetent, I should have error messages and stacktraces available in a few seconds. If I'm a rich cloud customer, I can call the premium cloud support and ask for an explanation. If not, I would probably have to debug it myself.
With your service, I understand that I can blame the cloud provider faster. Maybe it can make the debugging session slightly faster when your monitoring also returns errors. End users don't care whether it's my code or the cloud provider code crashing, so it's a developer tool for emergencies. Did I understand well?
It is not clear to me if the issue was with an old SDK using the newest api version in calls or was it something else?
[1] https://learn.microsoft.com/en-us/rest/api/cosmos-db-resourc...
They needed years to finally introduce PATCH in CosmosDB, Request Units feel like they're obscure on purpose to hide insane cost of using this storage, being able to use Stored Procedures only on one Partition Key(while /id is being the default...), requests would often fail with 429 Too Many Requests when the container was set to Autoscale with obscene limits that were never hit.
Just setup Marten with Postgres and get it over with for fraction of the cost.
This is a change without the package version even changing (even newest package didn’t work)
The two have nothing in common (and trust me, it sure is fun having to constantly make sure which of the two someone is actually referring to every time...).