The system that I'm currently responsible for made this exact decision. The database is the API, and all the consuming services dip directly into each other's data. This is all within one system with one organisation in charge, and it's an unmanageable mess. The pattern suggested here is exactly the same, but with each of the consuming services owned by different organisations, so it will only be worse.
Change in a software system is inevitable, and in order to safety manage change you require a level of abstraction between inside a domain and outside and a strictly defined API contract with the outside that you can version control.
Could you create this with a layer of stored procedures on top of database replicas as described here? Theoretically yes, but in practice no. In exactly the same way that you can theoretically service any car with only a set of mole-grips.
IME what data pipelines do is they implement versioning with namespaces/schemas/versioned tables. Clients are then free to use whatever version they like. You then have the same policy of support/maintenance as you would for any software package or API.
There is a big difference. The types of an API can be changed independently of your schema.
In reality other ways of solving the same problem have a decade of industry knowledge, frameworks and tooling behind them.
Is the marginal gain from this approach being a slightly better conceptual match for a given problem than the “normal way” worth throwing away all of that and starting again for?
Definitely not in my opinion. You’ll need to spend so much effort on the tooling and lessons before you’re at the point where you can see that marginal gain appear.
I've worked on production systems where this kind of stuff worked very well. I think there's weirdly a big wall between software and data, which is a shame, because the data world has a lot to offer SWEs (I've certainly learned tons, anyway).
> In reality other ways of solving the same problem have a decade of industry knowledge, frameworks and tooling behind them.
It's pretty likely that any database you're working with is as old or older than any software stack. Java, PHP, and MySQL were all released in '95 (Java and MySQL on the very same day, which is wild), PostgreSQL was '96. Commercial DBs are even older, SQL Server is '89, Oracle is '79, DB2 and SQL itself is 70s. There's a rich history on the data side too.
> Is the marginal gain from this approach being a slightly better conceptual match for a given problem than the “normal way” worth throwing away all of that and starting again for?
The gain is pretty tremendous: you don't need an app server, or at least you only need a very thin one. Tech has probably spent billions of dollars building app servers over the last 30 years. They're hard to build and even harder to maintain. Frankly, I'm tired of stacking up huge piles of code just to transpile JSON/gRPC to SQL and back again.
> Definitely not in my opinion. You’ll need to spend so much effort on the tooling and lessons before you’re at the point where you can see that marginal gain appear.
There's a lot of tooling, it's generally just built into the DB itself. And a lot of software tools work great with DBs. You can store your schemas and query libraries in git. You can hook up your CI/CD pipeline right into your database.
I also can't recommend dbt enough [0]; it's basically the best on-ramp for SWEs into data engineering out there.
Stored procedures technically can do anything, I guess, but at that point you would be better with traditional services which will give you more flexibility.
https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/Postg...
Also, using views and stored procedures with source control is a pain.
Deploying these into prod is also much more cumbersome than just normal backend code.
Accessing a view will also be slower than accessing an “original” table since the view needs to be aggregated.
Where does it say anything needs aggregating. You can have a view that exists just for security.
> Also, using views and stored procedures with source control is a pain. Deploying these into prod is also much more cumbersome than just normal backend code.
Uh? This is normal backend code.
Are modern developers allergic to SQL or what is the issue?
versioned views and materialized views are essentially api endpoints in this context. just developed in sql instead of some sane language.
SQL is a legendary language, it's powerful enough to build entire APIs out of (check out PostGraphile, PostgREST, and Hasura) but also somehow simple enough that non-technical business analysts can use it. It's definitely worth spending time on.
Wouldn't it be far simpler to just create a service providing access to those views with something like OData?
So when you provide an API - you don't make all functions in your code available - just carefully selected ones.
If you use the DB schema as a contract you simply do the same - you don't let people access all functions - just the views/tables they need/you can support.
Just like API's, databases have tools to allow you to evolve - for example, maintaining views that keep a contract while changing the underlying schema.
In the end - if your schema dramatically changes - in particular changes like 1:1 relation moving to a 1:many - it's pretty hard to stop that rippling throughout your entire stack - however many layers you have.
What are the database tools for access logs, metrics on throughput, latency, tracing etc.? Not to mention other topics like A/B tests, shadow traffic, authorization, input validation, maintaining invariants across multiple rows or even tables...
Databases often either have no tools for this or they are not quite as good.
- Throughput/latency: pg_stat_statements [1] or Prometheus' exporter [2]
- A/B tests: aren't these frontend things? recording which version a user got is an INSERT
- Auth: row-level security [3] and session variables
- Tracing, shadow traffic: I don't think these are relevant in a "ship your database" setup.
- Valdation: check constraints [4] and triggers [5]
Maybe by some measures they're "not quite as good", but on the other hand you get them for free with PostgreSQL. I've lost count of how many bad internal versions of this stuff I've built.
[0]: https://severalnines.com/blog/postgresql-audit-logging-best-...
[1]: https://www.postgresql.org/docs/current/pgstatstatements.htm...
[2]: https://grafana.com/oss/prometheus/exporters/postgres-export...
[3]: https://www.postgresql.org/docs/15/ddl-rowsecurity.html
[4]: https://www.postgresql.org/docs/15/ddl-constraints.html
[5]: https://www.postgresql.org/docs/15/plpgsql-trigger.html
The important crux of the counterpoint to this article is "if you ship your database, it's now the API" and everything that comes along with that.
All the problems you _think_ you're sidestepping by not building an API, you're actually just compounding further down the line when you need to do things to do your database other than simply "adding columns to a table". :\
Edit: re-reading, the point I didn't make is that having your database be your API _is_ viable, so long as you actually treat it as an API instead of an internal data structure.
Are you just taking about the expected shape of the data - the consumer of the database can do that either in SQL or at some later layer they control.
If you are talking about my 1:1 -> 1:N problem. I'd argue that can ripple all the way though to your UI ( you now need to show a list, where once it was a single value etc ) - not something you can actually fix at the API level per se.
Bottom line, the more layers of indirection, the more opportunities you have to transform - but potentially also the more layers you do have to transform if the change is so big that you can't contain it.
Let's be clear - I'd typically favour APIs -especially if I don't control the other end. But I'm saying it's about the principals of surface area and evolvability, not really whether it's an API or SQL access.
tbf, the idea isn't as novel. Data warehouses, for instance, provide SQL as a direct API atop it.
If all the users of an API were bound to the shape of the data returned, but you wanted to add an extra field, you'd have exactly the same problem surely?
Sounds like the problem was with too much magic in the layers above - as in the end the shape of the data returned from a query on a table is up to the client - you can control it directly with SQL - in fact dealing with an extra column or not is completely trivial in SQL.
My point is that if people build brittle stuff on top that's not a problem of the DB being accessible per se.
That could just as easily happen against an API.
I assume you had was some problem with ORM's and automatically built data structures etc - I would argue that's a problem with those, not with the DB.
In a past life, I worked for a large (non-Amazon) online retailer, and "shipping the DB" was a massive boat anchor the company had to drag around for a long time. They still might be, for all I know. So much tech and infra sprung up to work around this, but at some point everything came back to the some database with countless tables and columns where no one knew the purpose, but couldn't change because it might break some random team's work.
From the article:
> A less obvious downside is that the contract for a database can be less strict than an API. One benefit to an API layer is that you can change the underlying database structure but still massage data to look the same to clients. When you’re shipping the raw database, that becomes more difficult. Fortunately, many database changes, such as adding columns to a table, are backwards compatible so clients don’t need to change their code. Database views are also a great way to reshape data so it stays consistent—even when the underlying tables change.
Neither solution is perfect (raw read replica vs API). Pros and Cons to both. Knowing when to use which comes down to one's needs.
My last customer used an ETL tool to orchestrate their data loads between applications, but the only out of the box solution was a DB-Reader.
Eventually, no system could be changed without breaking another system and the central GIS system had to be gradually phased out. This also meant that everybody must had to use Oracle databases, since this was the "best supported platform".
When that thing fails again they will hopefully settle on a sane monolithic API.
Don't get me wrong, the amount of time it saves is massive compared to rolling your own equivalent, but it doesn't take long before you've dug yourself a big hole that would conventionally be solved with a thin API layer.
That's why API exists at first place.
[0]: https://www.postgresql.org/docs/15/ddl-rowsecurity.html
I'm not tooo familiar with DBs, but I know customers. They're going to present custom views to your client SDK. They're going to mirror your read-only DB into their own and implement stuff there. They're going to depend on every kind of implementation detail of your DB's specific version ("It worked with last version and YOU broke it!"). They're going to run the slowest Joins you've ever seen just to get data that belongs together anyway and that you would have written a performant resolver for.
Oh, and of course, you will need 30 client libraries. Python, Java, Swift, C++, JavaScript and 6+ versions each. Compare that to "hit our CRUD REST API with a JSON object, simply send the Authorization Bearer ey token and you're fine."
Building an API for a new application is a pretty simple undertaking and gives you an abstraction layer between your data model and API consumers. Building a suite of tests against that API that run continuously with merges to a develop/test environment will help ensure quality. Why would anyone advise to just blatantly skip out on solid application design principles? (clicks probably)
> Building an API for a new application is a pretty simple undertaking
This is super untrue, backend engineering is pretty hard and complicated, and there aren't enough people to do it. And this is coming from someone who thinks it should be replaced with SaaS stuff like Hasura and not a manual process anymore.
> Building a suite of tests against that API that run continuously with merges to a develop/test environment will help ensure quality.
You can test your data pipelines too; we do at my job and it's a lot easier than managing thousands of lines of PyTest (or whatever) tests.
> Why would anyone advise to just blatantly skip out on solid application design principles?
Because building an API takes a lot of time and money, and maintaining it takes even more. It would be cool if we didn't have to do it.