Object storage is all you need
tigrisdata.com
tigrisdata.com
To me it reads a bit like, how we built our office without desks: it turns out if you stack two chairs on top of each other, you can balance your laptop on the top and you’ll also have a shelf on the bottom for your things.
For those working with databases, it's easy to come up with a bunch of reasons. ORM is a necessary evil when working with RDBMS, as are things like schema migrations that can easily result in loss of data (I.e., dropping columns or tables).
What if all we need is dumping a big old JSON in a container?
This idea is very enticing. The scale of this whole NoSQL thing is pretty telling.
The blog presents a thought provoking question: what if we don't actually need a full-blown database, and instead we only care about things like unique constraints, transactions, indices, and history tables.
What if we only need a subset of those?
Hmm, neither of those been necessary for me. ORMs is a choice, you can choose not to and still have a proper design and architecture and not suffer from that choice. Requires you have programmers who know how to use SQL, which seems less and less common as the days go by though, so understandable most reach for an ORM.
Schema migrations that can "easily result" in data loss is even easier to avoid though. Don't drop the column/table in the same step you copy data to the new place, or do the migration via application code which gets cleaned up later, to catch most of the stuff before one bigger copy, and then the eventually DB cleanup.
Guess it depends on how careful you want to be, or how careful your agent lets you be, I guess. But it's definitely possible to avoid both of those "necessary evil", which I guess makes them just "evil".
See:
This sound like: what if we dont actually need a table and instead we only care about four legs and a wooden plate.
But that sounds kinda unstable and doesn't really provide the full benefits we expect from a...
Oh.
Your time and attention is precious as a developer. I'm absolutely sure it's possible to implement uniqueness constraints, transactions, indices, and history yourself, but is that really the most valuable use of your time? There's probably not a need for you to have a unique solution, so you're quite literally just re-inventing something someone already had for not a lot of benefit.
Wouldn't your time be better spent actually solving the problems that whatever you're building is supposed to solve?
If you want to kill someone's dream of publishing a game, encourage them to build their own engine from scratch. I cannot think of a more malicious piece of advice given how effective it is, statistically speaking.
The exceptions to this are so incredibly rare. Virtually all of the in-house engine work at indie scale has been replaced by Godot in recent years. It used to be something like 10-11% of studios were in-house fully custom. Now it's probably closer to 1-2% fully custom and 8-10% on Godot. Having a reasonably stable OSS option has removed the last major argument that I am familiar with for simply using what already exists.
I was interested in gamedev as a kid. I always went with a custom engine. You're right, I never released/finished a large game.
But now I think I was just into game engine dev. Every time I finished my game engine and started working on game content I ingredient got bored and went to another project.
So I guess this is the reason, engine development is fun :).
Same vibe here imo. Use zerofs, juicefs. Use picomq, automq, s3stream. Use slatedb. Use lancedb, fusion, tonbo, iceberg, vertex. Use duckdb, data fusion, polars. These systems all natively speak to object storage. Use Qwikwit, rising wave / hummock, greptine mito, use warp stream, neon. Use celld. Over half of the projects here are very specifically about object storage.
Many deal with uniqueness, transactions, indicies. Which is extra impressive because its across machines. Managing data on a box creates a huge array of challenges and difficulties, backup and scaling and HA, etc etc. A huge reason to face object storage head on is that often you are headed towards that fray one way or another, and object storage makes an excellent easy to scale and manage disaggregation layer that keeps your problems from complecting together, is a clear, well known separation point. With an incredible ecosystem around it.
What if you just ran FoundationDB instead?
Would that let you do collaboration blog-posts with another VC-funded startup though?
Also, the database is the last thing you want to re-invent unless your business is explicitly building a db (even then it's best to re-use a db like all the postgresql forks that have existed)
good news, FoundationDB is already invented, and just slightly more tested than others on the market.
In our current design we use FDB as our metadata store which is the sort of workload it is really good at.
That might be a very specific assumption. What about serializability? Replication? Materialized views? Procedures? Locking? Access control?
It’s cool to experiment and try new approaches. Neat one here.
Could still end up moving to Postgres.
Did it take 10 seconds of prompting? Or hours of thought, trial error and revision? Especially when asking people to read for 20+ minutes...
If you only have binary encoded protos to worry about (which is typical) then you can rename fields.
So, they built a thing that pretends to, but does not actually properly handle transactions?
I guess I should be glad they are not a fintech startup...
For about a decade, I've been using flat files on disk or object storage for most of my side projects. There's even a python library that handles some of the plumbing for you [1].
If you don't have strong record-level concurrency needs then it's a lot nicer, easier, cheaper than a relational or document database. And if you do need that, you can design your data model around what defines a record.
I generally don't read low-quality AI slop though, so its possible I missed an explanation that didn't contain those key words.
it would be interesting to see their experiences with using FoundationDB. this feels like a tech that is amazing if only it had more information and practical examples of how to use, leverage and manage it. the client is complicated and needs expertise to use correctly. would love it if there was more info about it all.
Article flagged as AI slop.
I had a sensible chuckle when I got to this part. This kind of article "____ is all you need" is like another case of Betteridge's Law of Headlines. The answer is "that's not true" every time.
Two things the post doesn't cover and I'm happy to get into. What closing the cross-region lost update actually took: annotating the RPCs that depend on a compare-and-swap and replaying those to a single region, with a client-side guard. And the read amplification, which we have a plan for but waiting on a clear signal for when it’s needed.
Cross-posted with thanks to the Tigris folks; the original is at ampbase.io.
From the article:
> In practice, when you reach for a database engine you're actually reaching for four basic features: unique constraints, transactions, indices, and history tables. In order to use Tigris' global object storage as a database, we had to implement all of these primitives ourselves.
Not sure why the "without a database" is or isn't so important, why is it mentioned so often and why the article flip-flopping between "we don't have a DB" and "we're effectively building our own DB"?
Set of my Claude alarm bell, and lo-and-behold, Pangram judges this comment to be 100% AI-generated, albeit with limited confidence.
edit. Clearly the joke about "lo-and-behold" didn't land well with the bots...
CMPXCHG
For compare-and-swap I have even got a lock less queue that I implemented with it.
This is Claude between the lines admitting none of it was needed. Postgres on a VM would be doing just fine right now.