I Don't Want to Teach My Garbage DSL, Either
github.com
github.com
> What’s worse than data silos? Data silos that invent their own query language.
and:
> I just want my SQL back. It’s a language everyone understands, it’s been around since the seventies, and it’s reasonably standardized.
Huh? SQL is just as DSL-y as anything else. It's so non-standardized that I can't take any program written for any RDBMS and run it against any other RDBMS, unless the author specifically extracted all the DSL into an interface layer and ported it to that other RDBMS already. Even something as common as "CREATE INDEX" has different syntax on every database I've ever used.
It's even worse than HTML/JS/CSS 10 or 20 years ago, where at least it had a chance of minimally working. And you can't seriously tell me that it's harder to make SQL implementations (where you control the database) portable, compared to, say, C++ compilers (where you don't control the hardware architecture).
Yeah, data silos with their own query languages are bad. But RDBMSs are some of the worst offenders, because they pretend this doesn't apply to them. It's easy for newcomers to justify creating their own SQL-like languages, because every existing database is already merely an SQL-like language.
All these database vendors need to get together and make some "SQL5" that finally works consistently.
There are lots of weird parallels between SQL and Smalltalk. For example, both SQL and Smalltalk are constructed with Douglas Hofstader's "strange loops." All Smalltalk object instances have a class, and the classes themselves are object instances. SQL tables are defined using metadata stored in SQL tables. Another parallel, is that the language variants are very similar, but ultimately incompatible to the point where translation is non-trivial. In Smalltalk, this is due to a toothless language standard, resulting from political maneuvering by various language vendors. Not sure what it is in the case of SQL.
There is a problematic relationship between DSLs and libraries in languages of sufficient power. In programming languages with a certain level of expressive power, it's easy to write a DSL on top of a library, or on top of another DSL.
I used to joke about "lady/gentleman computer scientists." What's the difference? A computer scientist knows how to implement another computer language. A lady/gentleman computer scientist knows when they should know better.
A DSL should be used to super-charge your project's very idiosyncratic special context. The trick is knowing better when you shouldn't. Accessing a database isn't your project's very idiosyncratic special context.
Arduino "whatever it is called but really c++"-language or Warcraft 3 scenario language are kinda nice for what they are.
People get attached to their creations and especially if the thing starts as a hobby project, the author can get carried away pretty far. When they eventually want to share their creation with the world they might get a really cold shower.
IMO us the programmers started becoming less practical and more tinker-y. Which is not a bad thing; it's healthy for the psyche. But when it comes to pieces of tech that can end up getting used on vast scales, we should be more responsible.
In this lane of thought, DSLs should be reserved to niche business domains and not as the first tool you reach for when you can't quite describe a problem with your programming language of choice.
Personally I think the best solution is using a static site generator like Hugo or Zola to control how you generate your content, and then host that using Caddy + Docker. A database is overkill.
You can also use caddy's git plugin to kick off automatic builds build from any git repo, not just a github one. It just needs to support webhooks, I believe.[1]
[0]: https://utteranc.es/
But there is nearly zero walled garden downside. I own the domain, I don’t use github.io, so I don’t fear link rot if I move.
Because it’s git, I always have a copy of everything locally, I don’t depend on github for storage.
Because it’s jekyll, I can generate my blog on my own system and upload it somewhere else whenever I want.
I don’t support comments at all, but that’s a personal choice. I’m not in the community business, I outsource comments to Hacker News. Which is also how most of my readers want to discuss my writing.
With the setup I mentioned above, I just build, commit, then git push, which triggers a pull on both of my servers due to webhooks.
Getting linked to reddit or hn and being discussed is fine enough.
The trouble with SQL is that it doesn't easily allow for basic building blocks that ORMs benefit from, like composability. This leads many ORM authors to build their own query language, which support features that SQL lacks or does poorly, that compile to SQL in order to simplify the rest of the development of the ORM.
This comes as a result of SQL not being a very good language (not to be confused with the application of declarative querying of relational data, which is beneficial and could benefit greatly from a good query language) by modern standards.
The application of the language does not make the language itself great. As we have seen with DSLs that often come bundled with ORMs, there are other languages to query databases with (even if they ultimately compile to SQL), and I would argue that some of them do a lot better job than SQL does for providing a comfortable and cohesive environment for developers to write queries in.
In the imperative language space, we have one hundred and one different languages all trying to make things slightly more comfortable to developers. C, Go, and Rust can all be used to write the same kind of application, more or less, but that does not mean all of those languages are equally great. The same is true of declarative queries. Just because SQL is popular does not mean it is great.
> When you know how to write efficient SQL statements, ORM feels like having a hand tied behind your back.
ORMs and SQL are orthogonal concepts, really. There is no reason an ORM couldn't require you to hand-roll every single SQL statement. An ORM's concern is simply mapping the results of that query into the application's objects. That some ORM implementations also include functionality to build queries for you, often on top of the aforementioned DSLs, to make that mapping require less effort on the developer is, I would argue, largely a result of SQL being a bad language.
I think, invariably when you use ORM, you end up having query-specific object structures. For example [1].
You also might be passed objects with deferred fields [2]. This will be completely opaque to someone consuming the resulting object. You'll eventually run into this problem [3]. Solving the lazy load problem requires an understanding of how SQL works in the first place. And if you look at the solution in that example, it's an ORM-wrapped series of joins.
From 3: > For good measure, we add a raiseload to throw an exception if we try to load anything that we didn’t load here.
Who wants to live in this world?
1: https://stackoverflow.com/a/45905714/5573538 2: https://docs.sqlalchemy.org/en/13/orm/loading_columns.html#c... 3: https://engineering.shopspring.com/speed-up-with-eager-loadi...
Which language(s) are you using as a point of comparison and why is SQL better than those other languages? SQL is no doubt better than nothing, but that is not in the spirit of our discussion.
But two things come to mind on reading your comment. 1. For the problem of Object-Relational-Mapping (ORM) which query language is used is not that important. If one accepts that there is one general purpose language for application development and another for managing persistent data, you will be left with a situation where there is a bunch of redundant code to write in whatever query language and application language you have available, and it would be nice to not have to write all that by hand. 2. Over the years, I have encountered many projects (languages) that attempt to address various deficiencies in SQL. I feel like a lot of these projects make life too hard on themselves. Rather than building extensions to existing database engines, they want you to adopt not just their new language, but their whole persistence stack. If I was smarter, and I had a design for a "better SQL", I would try to integrate it into Postgres or some other existing database engine. I guess that's what some of the ORM authors are trying to do, but by compiling their language to SQL supporting all databases at once. But returning to point 1, that seems pretty separate from the issue of Object-Relational-Mapping.
I don't know if it's better or worse - the original blog post asking governments to regulate data formats/code!
I agree that some DSLs make SQL better and easier to reason about -- and that Ecto is one of them.
As a single non-representative example, ActiveRecord has `before_insert` hooks you can simply add as methods to your model class. Ecto doesn't have those.
> Data Mapper: A layer of Mappers (473) that moves data between objects and a database...
(https://martinfowler.com/eaaCatalog/dataMapper.html)
Sounds like an ORM to me.
Some people try to distinguish between the "data mapper" pattern and the "active record" pattern (it's not just the name of the Rails library, it's a pattern... which the Rails library may or may not implement very well). Both are ORMs, because both are ways of mapping from an rdbms to an object system.
(Neither of which actually has to do with query DSL. We could imagine just taking the part of ActiveRecord that produces queries, but having it return simple hash/string literals. It wouldn't really be an "ORM" (except in the most technical sense that even hashes are objects in ruby), but it would still have the parts you don't like. The nature of query building is actually not related to 'data mapper' vs 'active record' -- you could have an instance of either in which you wrote raw SQL queries, or an instance of either which used the same non-SQL DSL)).
But even distinguishing between "data mapper" and "active record", in actual practice, I don't think there are two completely separate, distinct, and unified camps. I don't think these categories actually serve well to deliniate the ORMs we've got. Instead, there are a just a whole bunch of approaches, some more light weight than others, some more mature/reliable than others, some 'leakier' than others, differing on all sorts of additional dimensions. I agree that some ORMs are better than others -- and some may disagree on which these are -- I don't think saying "data mapper" is actually useful for understanding which these are.
[1] https://www.goodreads.com/book/show/8082269-domain-specific-...
Overall, I agree that it's important to think about how to make sure that operations compose in an unintuitive, expected way. I find this hard to figure out without thinking of it in terms of a grammar and semantics, whether it's embedded or not.
Now onto why I mention this:
> free of all sorts of fun bugs that happen when odd constructs are put together
I actually had an interesting experience with this, but in a good way. As I said above, I hoist the code for functions out of the AST (the final AST is basically a map of function to function body AST, where one of these is the main script code) and the declarations are replaced by the transformation step with a local variable binding that gets set to some data describing how to call the function (key into the map, list of parameters, other stuff), so when you do 'f()', it looks up 'f' in local state, gets this data and figures out how to call the function from that. What this meant was that I could pass functions around. And so I got higher order functions almost for free, without having planned for it, just because I thought "hmm, what if I do this?" (the current version did add some special logic to capture state, for closures, but I did that after when I decided why not support it for real).
As an aside, my little language is a synchronous language[1], which acts as if its runtime is instantaneous (ie, the world does not change until it completes), which makes a lot of things easier (both for me and for users) and makes sense in the domain, since its basically commands that get triggered when certain conditions are met. This almost makes it worth writing a custom language over using an existing one, but if I were to start again, I'd probably just use Lua or Javascript and save some time, even if the final language isn't so easy for my target audience.
[1] https://en.wikipedia.org/wiki/Synchronous_programming_langua...
We can't write software without making APIs...
How would one go about securing this thing anyways?
(I agree with the basic sentiment, but find it somewhat unrealistic)
If along you expose not the the real tables but views I honestly don't see what could go wrong.
Then again, no doubt there are some cases where this is the best solution. But it's worth being cautious before adopting an approach like this.
com·plex·i·ty n. the state or quality of being intricate or complicated.