We also have the ability to with our visual language join data from multiple different queries/datasources. You can see how that happens here https://chartio.com/docs/visual-sql/merge-queries/
2,294 karma · joined March 27, 2007
blog at
http://thingsilearned.com
We also have the ability to with our visual language join data from multiple different queries/datasources. You can see how that happens here https://chartio.com/docs/visual-sql/merge-queries/
We do this though as a very flexible visual language, where you can create a pipeline of actions including merging queries from multiple different datasets and then doing some post query computation.
I believe most data exploration just shouldn't be SQL/text based. It's just faster to do it visually if there aren't extra steps added and the interface is good/flexible enough.
Even as developers we spend most of our time in a GUI environment vs the command line. It's not just that it's easier/better for business users - it's also just better for everyone.
The interface right now only works for structured data sources and my brain also starts to sputter when I think about making a visual language for anything not structured :).
If security is your main reason for wanting self-hosting you may be interested to know we're also SOC2, HIPAA, and GDPR compliant.
Right now the SQL we write is very proper - with quotes around all of the column names and table names listed before each column name. It's not what a human would write. I'd love to one day make it a little more human so that the switching into SQL mode will feel even better. It'll be a fun project.
DataCamp.com has some great courses
Tableau has some great online education - many people learned BI from them
Many free resources for getting started with data at different levels (note: this is one of my sites) https://dataschool.com
Description of common roles in BI/data: https://chartio.com/learn/data-analytics/distinguishing-data...
It covers BI a bit, but mostly the stack that BI sits on top of. It's an open book so we're always looking for suggestions and experiences such as those shared here.
The T being after the L means you can do that stage more simply in just SQL (with views or materialized views - possibly with the help of DBT), as opposed to some vendor interface, or python/R/etc script.
It also means that rebuilding your warehouse is much less significant of an ordeal. When the structure of source data changes or if you want to make some migration of the schemas of the warehouse you don't need to also re-run your ETL jobs and start over from scratch.
We try to explain some of that here - https://dataschool.com/data-modeling-101/row-vs-column-orien...
And Fivetran did a great benchmarking of it here - https://fivetran.com/blog/obt-star-schema
The architecture of the C-Store warehouses often removes the benefits of materialized views. This is why for a very long time Redshift didn't even support them - they insisted they weren't needed as they didn't improve performance over regular redshift significantly.
And though ELT is a very common standard now - I don't know of a single book that recommends it or explains why that change has happened. Just one of the reasons for writing a new data book.
2. C-Store does largely solve pervious performance and cost issues with denormalization. We also write a bit in this book (more to come) on how to avoid doing all that copying.
It allows you to not do E & T & L all together. It's really nice (less complex, easier to implement, less costly, and more flexible) to have that T part pulled out and done after.
In the book we use much of the old terms and recommendations. Most of the high level organizing is still totally right - but a lot of the optimizing and work done for performance and cost reasons is very different now.
For example ELT makes now much more sense than ETL for the reasons Kostas wrote about here: https://dataschool.com/data-governance/etl-vs-elt/
And many things previously done for cost and performance reasons are just not relevant anymore thanks to the big innovations in C-Store warehouses.
There are a number of people starting to talk about Star Schemas having little gain on modern stacks. The performance and costs gains are automatically done now with C-Store warehouses. Fivetran has a great post on this https://fivetran.com/blog/obt-star-schema
Besides the need for a new data book, we realized that it needed to be of a different format, as the space is moving so fast the expertise is very distributed.
Also, it's definitely a challenge to support 100's and 1000's of users digesting data, especially in the democratized fashion that we're typically used in. It takes good data governance, support, and admin tools. I gotta chime in and say for the record though that Chartio well supports many such customers.
It may be weird for you to chat as you worked at Looker, but I'd love to hear anytime on why you see Chartio capping out at a lower # than the others listed. You can reach me at dave-at-chartio.com!
https://chartio.com/learn/sql/
With the writeup on why i made it here: https://medium.com/@__dave/why-i-wrote-yet-another-sql-tutor...
Do let me know if there's any collaboration we could do!
There are challenges built in and practice sections at the end of each segment. Any feedback/corrections/suggestions are hugely appreciated.
More on why I wrote it here: https://medium.com/@__dave/why-i-wrote-yet-another-sql-tutor...
I'm the founder of Chartio.com here. Not sure if you've given us a look yet (or lately) but we pride ourselves on being as usable and flexible as possible. Would love to show you more: dave@chartio.com