But in all seriousness, DataFrame centric operations are a superset of SQL (you can always do a df.sql("...") if you want to ) and have a lot more efficient implementation of both OLTP/ORM requirements and OLAP/DS/BI requirements.
They also encourage composability, modularization, reuse, unit and data testing ...
So it's ironic I feel like SQL's replacement is another declarative language - its original inspiration - plain English.
Just a natural language transpiler (like Palantir's Ontology plus Looker's Malloy (reverse disclaimer: I do not work for or enjoy either of these products but these underlying concepts are correct)) with some fancy domain heuristics and light AI (I suspect a Pareto like model that supports 80% of use cases only needs a semantic graph with a few thousand nodes and vertices)
There are some problems though:
- Single query can't write to multiple tables.
- Single query can't return multiple resultsets.
You can write a stored procedure, but this is no longer "an SQL query", strictly speaking. And once you start writing stored procedures, you are no longer using just SQL, but whatever Ada-inspired procedural extension to SQL was implemented by the database vendor. In other words, you are using "a traditional programming language to work with the data".
You could compose a SQL query that allows you to map multiple resultsets to 1 resultset, although that feels a bit awkward.
WITH a AS (
insert into a (k, v) values ('a', 1.0) returning *
), b AS (
insert into b (k, v) values ('b', 2.0) returning *
)
SELECT
row_to_json(a)
FROM
a
UNION ALL
SELECT
row_to_json(b)
FROM
b;
Returns: row_to_json
--------------------------
{"a_id":1,"k":"a","v":1}
{"b_id":1,"k":"b","v":2}
(2 rows)I'm currently on SQL Server and it doesn't support INSERT as a CTE (and I think most DBMSes out there still don't). It would definitely make my life easier if it did...
https://docs.snowflake.com/en/sql-reference/sql/insert-multi...
ah, the insert is a CTE because it produces a value ('returning' I guess). Hmm. This is very odd. Doesn't seem to work in mssql.
Well thanks for the can of worms...
with x as (...)
update x set ...
> you can't nest cteok but you can linearise them
with x as (...), y as (...)
> the query planner has no understanding of joins that cross cte boundaries so they are totally unoptimizedutter, reeking garbage.
> you can't use distinct or group
more garbage. I have. Show me an example of it not working.
(Edited for less rudeness)
WITH t AS (
DELETE FROM foo
)
DELETE FROM bar;
Or can you only use SELECT for WITH queries? Did you not realize that other databases and the SQL standard allow you to do this?Yes on the final query of the CTE you can do all sorts of things, but that's way less useful if you can't do them in all the component queries.
> Or can you only use SELECT for WITH queries
you can only use a select inside a cte (or should be able to) because the 'e' stands for 'expression'. It seems postgres does allow an insert with and output which sort of makes sense but I doubt it's in the standard.
Your last sentence makes no sense to me. Give an example.
Also you failed to give an example that group by/distiinct weren't allowed in ctes.
That may be technically true (or not, I don't know), but in many cases some complex data manipulation (especially when it's done in multiple passes) practically needs a lot of ram and a programming language to be more time efficient.