a * b Cross-product / Join if column names match
a + b Union
a(first=="joe" && salary>100000) Select rows
a[first,last,ssno] Select columns / project
I hate SQL syntax and wish relational algebra was used instead. a * b Cross-product / Join if column names match
a + b Union
a(first=="joe" && salary>100000) Select rows
a[first,last,ssno] Select columns / project
I hate SQL syntax and wish relational algebra was used instead.The argument that SQL allows you to say what the result should be without any reference to the order in which the steps should be done is fine, except that it is hard to say some things in SQL. If C compilers can transform sequential, imperative code to an equivalent optimized form, I don't see why relational algebra compilers / optimizers cannot.
Optimal join order usually is a function of both the query and the data, and a query optimizer inside your database doesn't necessarily find the optimal join order. It uses heuristics (which often include particulars about the data currently stored in the database) to find a join order that's hopefully better than a naive query plan, in much the same way that optimization passes in a C compiler use heuristics to generate machine code that's hopefully faster than a naive translation to machine code.
> The argument that SQL allows you to say what the result should be without any reference to the order in which the steps should be done is fine, except that it is hard to say some things in SQL.
The GP is talking about a surface syntax change, not a change in the underlying computation model. Using a relational algebra notation, the query optimizer would have just as much freedom as an SQL query optimizer. Relational algebra isn't any more inherently imperative than SQL is.
> If C compilers can transform sequential, imperative code to an equivalent optimized form, I don't see why relational algebra compilers / optimizers cannot.
It sounds like you want a semantic change that gives the query optimizer less freedom than an SQL query optimizer. The GP is suggesting only a syntactic change. The thing it sounds like you want does sort-of exist... many SQL databases will let you inspect the query plan that their optimizer has generated.
However, it sounds like you want some statically defined query plan. The problem with this is that the optimal plan depends on the data that's in the database at the time the query is run. For instance, a query optimizer can look at a complex query with multiple constant WHERE clauses on indexed columns, and use the indexes to quickly determine the size of intermediate tables when deciding the order in which to perform joins. A query language that statically defines a query plan cannot take advantage of this information, unless you want to "re-compile"/"re-optimize" once a day or something. However, if you trust the database to automatically re-optimize on some schedule, then you've lost your static query plan and it seems you might as well let the query optimizer regenerate the plan based on heuristics created by the database developers rather than sticking to some static schedule of recompilation/reoptimization points.
... and this notation almost seems more similar to the underlying math than SQL does.