A LLM+OLAP Solution
doris.apache.org
doris.apache.org
I wish it spent time on talking about how they trained their LLM to reliably generate parsable queries for the semantic layer, and what the accuracy rate of what the user intended vs what they got.
I do think the only way a LLM based analytics tool can succeed is via a semantic layer rather than direct SQL, since database schemas fail to encode a lot of information about the data (EG a warehouse might not even know user.customer_id = customer.id).
Malloy could be an interesting target here.
eg Snowflake lets you declare all the foreign keys you want, but does nothing with that info except let you use it.
No OLAP database I know of would let you encode other semantic layer things like aggregations or metrics. EG defining a DAU/MAU metric as "The distinct number of users logged in that day vs the distinct number of users in the 28 days before that day."
Those types of definitions usually live in the semantic layer or bi layer, which a LLM analysis tool would need to solve for.
One thing you really need with LLMs is consistency. Text-to-SQL kind of lets the LLM do whatever it wants - join tables that shouldn't be joined, define aggregates one way in one query and another way in the next.
Because semantic layers define how tables should join, measure definitions, etc., they mean people get consistent results from one query to the next, which builds trust in the LLM.
Cube (which was mentioned in another comment and has a great open-source semantic layer) has a good article about that here: https://cube.dev/blog/semantic-layer-the-backbone-of-ai-powe....
Others include AtScale[0], dbt's MetricFlow[1], Google's Looker[2] (also a BI tool but powered by a semantic layer), and Propel[3].
[1] https://www.getdbt.com/product/semantic-layer
[2] https://cloud.google.com/blog/products/data-analytics/introd...
[3] https://www.propeldata.com
They're kind of an updated version of OLAP cubes if you're familiar with those.
Typically semantic layers sit on top of a data warehouse, let you define metrics using code or a UI, and provide APIs or SQL connectors so that you can query them.
Less about one-shotting the answer, and more about showing its work, if it errors, letting it self-correct. Latency goes up, but quality of the entire conversation also goes up, and feels like it builds more trust with the user. Key steps are asking it to "check its work", and watching it work through new code etc. (I open-sourced one version of this: https://github.com/approximatelabs/datadm that can be run entirely locally / privately)
From their article: I'm surprised they got something working well by going through an intermediate DSL -- thats moving even further away from the source-material that the LLMs are trained on, so it's an entirely new thing to either teach or assume is part of the in-context learning.
All that said, interesting: I'll definitely have to try out tencentmusic/supersonic and see how it feels myself.
When Veezoo connects to a database / dwh for the first time, an initial Semantic Layer / Knowledge Graph gets built automatically based on the data itself. We try to recognize how the columns link to other tables, try to identify units, and other semantic information e.g. if something is a "Location" or a "Country" and so on.
The whole conversational "plain english" querying then operates on top of the semantic layer, ensuring business logic (and other governance topics) are always respected.
Same goes for calendar vs. fiscal year for companies that have different fiscal and calendar begin dates. Something as simple as "2023 YTD" will mean different things depending on the audience within an organization.
I’m prototyping a distributed DuckDB in the same vain as LiteStream for SQLite and I wonder if it would be a good fit for something like this.
About distributed DuckDB, have you checked Boiling Data? [2]
1- https://cube.dev/blog/introducing-duckdb-and-motherduck-inte... 2- https://boilingdata.medium.com/lightning-fast-aggregations-b...
I feel like I never have novel ideas sigh.
Interesting links though, thank you!