HNHacker News
TopNewBestAskShowJobs

thingsilearned

2,294 karma · joined March 27, 2007

Dave @ Chartio (chartio.com) http://chartio.com/

blog at

http://thingsilearned.com

submissionscomments
thingsilearned··on Show HN: Visual SQL
Yup - if you grab columns from multiple different connected tables (or schemas) the join will happen automatically for you and can be adjusted to different paths. This is actually a really hard problem, especially since 70% of datasources these days don't have foreign key information and we have to detect that automatically.

We also have the ability to with our visual language join data from multiple different queries/datasources. You can see how that happens here https://chartio.com/docs/visual-sql/merge-queries/

thingsilearned··on Show HN: Visual SQL
Sigma and Chartio have similar missions but a very different product approach and functionality. Sigma also tries to bring the influence of spreadsheets into the query building/data exploration phase.

We do this though as a very flexible visual language, where you can create a pipeline of actions including merging queries from multiple different datasets and then doing some post query computation.

thingsilearned··on Show HN: Visual SQL
Yeah we're somewhat targeting a different audience here but I also think this will end up being a much better experience even for the power users than an IDE.

I believe most data exploration just shouldn't be SQL/text based. It's just faster to do it visually if there aren't extra steps added and the interface is good/flexible enough.

Even as developers we spend most of our time in a GUI environment vs the command line. It's not just that it's easier/better for business users - it's also just better for everyone.

thingsilearned··on Show HN: Visual SQL
To be clear - this isn't about making visualizations from SQL results (though that does happen). This was us making a visual query language. It's a very flexible interface that writes SQL for you - and has been proven to be intuitive enough for 80% of business users, so they can write SQL now too.

The interface right now only works for structured data sources and my brain also starts to sputter when I think about making a visual language for anything not structured :).

thingsilearned··on Show HN: Visual SQL
I appreciate the feedback. We are discussing a few different pricing models, and if your team size is below the 5 seat minimum you can reach out and we're quite flexible there.
thingsilearned··on Show HN: Visual SQL
Ah, unfortunately there isn't. We're 100% cloud hosted - though we call ourselves a Hybrid cloud solution because through our reverse SSH tunnel connections (where your servers SSH into ours) we also work well with on-prem data.

If security is your main reason for wanting self-hosting you may be interested to know we're also SOC2, HIPAA, and GDPR compliant.

thingsilearned··on Show HN: Visual SQL
Yeah - changing date formats, double typing things you're grouping and ordering by, remembering the oddities of each dialect, typing out full join paths - it's a nice experience to have that done automatically.

Right now the SQL we write is very proper - with quotes around all of the column names and table names listed before each column name. It's not what a human would write. I'd love to one day make it a little more human so that the switching into SQL mode will feel even better. It'll be a fun project.

thingsilearned··on Show HN: Visual SQL
Dave, founder of Chartio here. We're so excited to launch what we call Visual SQL today. It's been a lot of work based on customer feedback and extensive prototyping and user testing. If you'd rather skip the story attached here you can also check out our product walkthrough video or give it a spin yourself here:

https://chartio.com/product/visual-sql/

thingsilearned··on Ask HN: What does your BI stack look like?
If you're already doing engineering work you may be able to just go get a job focusing on data - if you're upfront about wanting to learn. I've found most people have learned the skills on the job. Besides the books mentioned I commonly recommend these resources

DataCamp.com has some great courses

Tableau has some great online education - many people learned BI from them

Many free resources for getting started with data at different levels (note: this is one of my sites) https://dataschool.com

Description of common roles in BI/data: https://chartio.com/learn/data-analytics/distinguishing-data...

thingsilearned··on Ask HN: What does your BI stack look like?
Dave from Chartio here, wanting to share our new book describing the 4 stages of setting up your ideal data stack here - https://dataschool.com/data-governance/.

It covers BI a bit, but mostly the stack that BI sits on top of. It's an open book so we're always looking for suggestions and experiences such as those shared here.

thingsilearned··on We need new data books, so we started one: Cloud Data Management
Yes, and also the whole process is much simpler to pull the T out. The reason T had to be done at the same time as E&L was because of those storage and performance costs. Now you don't have to - and it separates the stages and simplifies.

The T being after the L means you can do that stage more simply in just SQL (with views or materialized views - possibly with the help of DBT), as opposed to some vendor interface, or python/R/etc script.

It also means that rebuilding your warehouse is much less significant of an ordeal. When the structure of source data changes or if you want to make some migration of the schemas of the warehouse you don't need to also re-run your ETL jobs and start over from scratch.

thingsilearned··on We need new data books, so we started one: Cloud Data Management
This is what I mean by updates due to C-Store warehouse engines (Redshift, BigQuery, Snowflake, etc). It's not just that the cloud providers are happy to rent you more space - it's also that they're C-Store and because of that redundancy in columns is well compressed automatically.

We try to explain some of that here - https://dataschool.com/data-modeling-101/row-vs-column-orien...

And Fivetran did a great benchmarking of it here - https://fivetran.com/blog/obt-star-schema

The architecture of the C-Store warehouses often removes the benefits of materialized views. This is why for a very long time Redshift didn't even support them - they insisted they weren't needed as they didn't improve performance over regular redshift significantly.

thingsilearned··on We need new data books, so we started one: Cloud Data Management
I totally agree. There has been some progress here recently. Have you checked out DBT and their testing features?

https://docs.getdbt.com/docs/testing

thingsilearned··on We need new data books, so we started one: Cloud Data Management
1. It is - and that's why I chose that example (a less controversial one) here.

And though ELT is a very common standard now - I don't know of a single book that recommends it or explains why that change has happened. Just one of the reasons for writing a new data book.

2. C-Store does largely solve pervious performance and cost issues with denormalization. We also write a bit in this book (more to come) on how to avoid doing all that copying.

thingsilearned··on We need new data books, so we started one: Cloud Data Management
Yup! And that commoditization of hardware has made it really inexpensive to have a Data Lake, where you first put all your data in raw format (so you only need to do EL - and not T all in the same step). And then, because of the way C-Store sources like Redshift are built it makes a ton of sense to just do your T step as a set of Views (materialized or not) onto of that Data Lake.

It allows you to not do E & T & L all together. It's really nice (less complex, easier to implement, less costly, and more flexible) to have that T part pulled out and done after.

thingsilearned··on We need new data books, so we started one: Cloud Data Management
not yet. Just PDF. But we'll be working on getting it into print eventually.
thingsilearned··on We need new data books, so we started one: Cloud Data Management
We're definitely not trying to start from scratch or throw out all the old knowledge/practices - just update them for the common data stacks used today.

In the book we use much of the old terms and recommendations. Most of the high level organizing is still totally right - but a lot of the optimizing and work done for performance and cost reasons is very different now.

For example ELT makes now much more sense than ETL for the reasons Kostas wrote about here: https://dataschool.com/data-governance/etl-vs-elt/

And many things previously done for cost and performance reasons are just not relevant anymore thanks to the big innovations in C-Store warehouses.

thingsilearned··on We need new data books, so we started one: Cloud Data Management
Yeah exactly. Much of the "overkill" was done because of performance and cost reasons, that frankly just don't apply anymore. Now the largest expense by far is time.

There are a number of people starting to talk about Star Schemas having little gain on modern stacks. The performance and costs gains are automatically done now with C-Store warehouses. Fivetran has a great post on this https://fivetran.com/blog/obt-star-schema

thingsilearned··on We need new data books, so we started one: Cloud Data Management
We use Jekyll. It's a custom design from our own awesome Steven Lewis.
thingsilearned··on We need new data books, so we started one: Cloud Data Management
Thanks! We're truly looking for this to be community driven as well. So if you see places to contribute, or where you might disagree, or where you could share a story - do let us know or make a pull request on GitHub!

Besides the need for a new data book, we realized that it needed to be of a different format, as the space is moving so fast the expertise is very distributed.

thingsilearned··on Show HN: Open-Source Business Intelligence for BigQuery – Looker Alternative
Hey @segah, founder of Chartio here. I can deeply second all the comments on how long of a feature tail BI is, and the amount of work it takes to have real product depth and stability vs. an impressive demo. It takes years, and ongoing maintenance that is impossible to estimate in the beginning.

Also, it's definitely a challenge to support 100's and 1000's of users digesting data, especially in the democratized fashion that we're typically used in. It takes good data governance, support, and admin tools. I gotta chime in and say for the record though that Chartio well supports many such customers.

It may be weird for you to chat as you worked at Looker, but I'd love to hear anytime on why you see Chartio capping out at a lower # than the others listed. You can reach me at dave-at-chartio.com!

thingsilearned··on SQL: One of the most valuable skills
I wrote this interactive SQL tutorial (you can write SQL in the page, and it tells you if you're right or not) last year! Always looking for feedback:

https://chartio.com/learn/sql/

thingsilearned··on Show HN: Select Star SQL, an interactive SQL book
Awesome work Kao!! Earlier this year I launched a similar interactive SQL tutorial with similar goals.

https://chartio.com/learn/sql/

With the writeup on why i made it here: https://medium.com/@__dave/why-i-wrote-yet-another-sql-tutor...

Do let me know if there's any collaboration we could do!

thingsilearned··on Let's Encrypt Root Trusted by All Major Root Programs
Congrats Josh!
thingsilearned··on Show HN: Interactive PostgreSQL Tutorial
Thanks! I've actually got plans to write a post or two about how it was built and open source it. It's a nice minimal example of a python server and a database connection.
thingsilearned··on Show HN: Interactive PostgreSQL Tutorial
I wasn't comfortable recommending existing SQL tutorials to any friends or customers, especially for the non-engineers. It's such an important tool to learn that I've made a crack and making it much more approachable, both in the writing style and in the tool (SQLBox) embedded in the page that lets readers follow along with writing SQL against a live PostgreSQL instance with no setup required.

There are challenges built in and practice sections at the end of each segment. Any feedback/corrections/suggestions are hugely appreciated.

More on why I wrote it here: https://medium.com/@__dave/why-i-wrote-yet-another-sql-tutor...

thingsilearned··on Triplebyte Raises $10M from Initialized Capital, Marissa Mayer and Paul Graham
Such a great service. Recruiting with real value add for both candidates and the companies. We love Triplebyte at Chartio!
thingsilearned··on Amazon QuickSight – Business Intelligence by AWS
Both Chartio and BIME do support BigQuery as well.
thingsilearned··on Amazon QuickSight – Business Intelligence by AWS
Hi Vicki,

I'm the founder of Chartio.com here. Not sure if you've given us a look yet (or lately) but we pride ourselves on being as usable and flexible as possible. Would love to show you more: dave@chartio.com

thingsilearned··on Our First Certificate Is Now Live
Congrats Josh and team!!!!
← PreviousPage 3 of 10Next →