This is a great group already doing this with bison, and he’s got a great book as well
2,294 karma · joined March 27, 2007
blog at
http://thingsilearned.com
This is a great group already doing this with bison, and he’s got a great book as well
This book though is to be clear (and that's in the first about sections) does not have much on product/event based analytics nor giant scale data as judged as 100+ TB data management.
The majority of the important things for a startup are usually already being tracked with timestamps in the database (new signups, new users, churned users, new todo items, etc). If a key metric is not directly tracked, there's usually a good enough proxy available somewhere in the database. A Busineess Intelligence product, or a SQL to chart tool is much more applicable and affordable in the startup stage.
There are likely going to be responses here "But what about..." and I'll premptively respond to those with the question: Could that be answered/solved almost just as well with a query to the database, without sacrificing the agility, extra setup and data collection time, and expense overhead? Those are all very costly things for a startup.
We're now iterating quite quickly on expanding and improving the options and would love any feedback!
1. A cleaner universal more natural syntax for analytics: I love writing python as it is to me such a cleaner syntax than C or Java. We could do the same for SQL and make something that feels more natural. Turning a common query like
> SELECT count(*), TO_CHAR(created_at, 'YYYY-MM-DD') FROM Accounts GROUP BY TO_CHAR(created_at, 'YYYY-MM-DD') ORDER BY TO_CHAR(created_at, 'YYYY-MM-DD');
into something much more natural like
> count by Day(Accounts.created_at)
2. A Visual SQL: for analytics it's so much faster to query and explore visually. Building queries visually means you don't make common typo or syntax or structure errors, joins happen smoothly, you can browse the data as you build, you don't need to google for syntax (what's that date function again?), and it works across dialects and databases. We've built and launched this a few months ago at Chartio https://chartio.com/blog/why-we-made-sql-visual-and-how-we-f...
1. Chart based commenting system for Chartio https://chartio.com/blog/charts-worth-commenting-on/
2. Helping the https://howwefeel.org team with their data and dashboards https://how-we-feel-chartio.herokuapp.com/the-how-we-feel-pr...
3. A number of hydroponic experiments including with lettuce https://img.chartio.com/nOueDAy2 and even trying corn https://img.chartio.com/NQugK2G0
4. Finishing a book on Data Management that'll have a dead tree edition published later this year https://chartio.com/blog/cloud-data-management-book-launch/
5. A few small wood projects like a planter bed and a projector mount
We've got ours also well setup with a big library of assets like buttons, and illustrations we've had made, and different example charts, so if what you're discussing is a new feature or landing page it's really easy to drag in a bunch of ready-made components to express your idea.
We're really excited about how this came out - like comments on google docs, it makes for an entirely different and more collaborative experience. I hope it helps teams create better dashboards together and more easily discuss their data.
As for large scale datasets - we do cache those sample tables and we only grab the first 10 rows. Also, BigQuery actually has a great API for fetching a set of sample data that we utilize heavily. They made it because of exactly the potential concerns you outlined around columnar stores.
So what you describe is somewhat built in to what we have now. Users still have to choose what columns they want to look at (there's no real way for us to guess that) and then we do apply some knowledge on what type of data they're looking at to help them get to what they're likely looking for.
We've also tried at times to make default dashboards for data sources when people connect. We can do this to some extent with known Schemas like connecting GA, SalesForce, Hubspot, etc, but for databases - that's proven to be a largely impossible task so far. Everyone's data is so different, and have such odd conditions to consider filtering by, that the auto dashboards end up being quite useless.
As for an established service - Chartio's been around for almost 10 years, we're profitable, and are the main data interface behind some really great companies and brands https://chartio.com/customers/
This video is maybe the closest you can get without trying yourself - https://www.youtube.com/watch?v=YBXMTipHGfQ
https://dataschool.com/data-governance/
Our next phase is to help people get to that cleaner source of truth much more quickly than traditional dimensional modeling approaches. Tools like Visual SQL and DBT (https://www.getdbt.com) are really changing the complexities here.
Every product needs some education (excel is still of the most popular online courses) and we knew ours would be no exception which is why we've also last year launched DataSchool - our free online community driven courses on data