Airbnb open-sources Caravel: data exploration and visualization platform
github.com
github.com
I'm not knocking Caravel (it looks amazing) just curious why build vs buy in this case.
[1] Tableau, Looker, Periscope, Chartio, Qlikview, Gooddata are just some that come to mind.
Anyway, the ones where bring your own database can scale as far as the database can bring you.
But the JDBC -> Druid connectors that exist look pretty janky. So if someone builds a stable connector, I suspect we'd support it. But for the moment, no Druid.
Larger, data-driven companies with significant engineering teams prefer not relying on 3rd party, closed-sourced vendors. That can represent a significant risk and a blockage for deeper integration with other internal applications when needed.
Not that building always wins over buying, but the balance shifts relatively to the size of the company.
Also, when using open source on the receiving end of the equation, you want to be a good citizen and contribute back to the ecosystem. It ties to pride, passion, and reflect a strong engineering culture, which can help with recruiting.
I work on Google Cloud, where we have both closed-source - BigQuery, open source - Dataproc, and closed source that makes open source rock - Dataflow/Beam. There are merits to each.
BigQuery is serverless and multi-tenant and can't really exist outside of Google Cloud due to its intrinsic dependency on low-level services that don't exist elsewhere (and for the same reasons we couldn't directly externalize Borg, choosing to create an OSS clone in Kubernetes).
It is not unusual for me to hear from folks that they spent 6+ months building a performant and sizable, say Presto cluster. Then there's continuous management, tinkering, configuration, and optimization projects. I hear this from companies that one would consider sophisticated technologically.
By contrast, 40% of all of BigQuery's Petabyte customers scale to these levels without ever talking to us. We just find them on consumption reports. On multiple occasions we've had "surprise" load tests of millions of rows per second streamed into BigQuery, and it just works. BigQuery is also HA out of the box at no additional cost, which is a great luxury.
So sometimes if you need to scale analytics to Petabytes, the option is to just consume a managed and cost-effective service, or tinker with OSS, where there are significant operational tradeoffs. On the other hand, as you said, you build pride and culture. It's also far from an automatic shoo-in that OSS gives you better TCO (against the old closed-source guard, yes, but not so much BQ). Thus, the relationship between company size and value of technology can invert at higher levels.
With all that, I'd love to see a Caravel-BQ plugin :)
(PS. Kudos to Druid for introducing a Streaming ingest. BigQuery also sees value in Streaming ingest, GA-ing our own Streaming API in March of 2015, and Kudos to Airflow).
Free as in beer is one incentive as licenses are not cheap, and vendors know when they have you locked down and tend to milk everything they can.
More importantly, software for which we don't have control over the source is a risk. In this day and age anyone that cares enough should be able to push a bugfix/hotfix overnight. What if you'd have to wait for entire quarters or years for Tableau to parallelize their "live mode", or to get connectivity to Presto to work?
What if you want to integrate a new type of visualization that isn't supported? What if you want to integrate with your anomaly detection framework or your A/B testing framework or other internal or external facing applications?
Since this is a common need for most companies, it makes sense to have an open source solution that we can all use and collaborate on.
Free as in beer is never the answer. A project like this takes multiple engineer-years to build and maintain. That's hundreds of thousands of dollars, at least. How much is a site license? Are you sure? Have you negotiated the rate?
Even for "expensive" services, buying it from someone else is almost always cheaper than paying someone to maintaining it yourself, because expensive services are usually expensive for a good reason: they're niche, and finding someone with the expertise to build it is expensive. And having the source is for a product so that you can customize it is certainly a better answer, but it rarely happens, in practice. It's why we have gobs of open-source Apache-foundation products that nobody in their right mind wants to host in-house, unless they absolutely have to.
Developers have a real, well-documented resistance to paying for things, and it sucks. Because in reality, most development of open-source tools happens when someone gets paid to maintain the tool. If they don't, the tool falls into disrepair. Open-source software isn't free -- it's just paid for by someone else.
Building a community can also help distributing the costs. But more importantly it's a labour of collaboration instead of a client/vendor struggle.
But if you find this to be a frequent occurrence, you're almost certainly evaluating costs incorrectly.
We also use some of the previously mentioned visualization technologies and get generally poor results by all measures.
If AirBNB has a few engineers around who are part of solving a unique challenge AirBNB has, then they are probably getting a better deal than we are. I'd wager their productivity per resource is significantly better, not to mention licensing, infrastructure and control over product features. That said, not everyone is AirBNB and has the ability to attract and retain that type of talent.
I'm not saying that you wouldn't choose a good open source alternative if it exists. I'm saying that you shouldn't run off and implement something that exists commercially if your only reason is that the commercial thing costs money.
1. Our problem set is not novel, we're more or less like many others out there and collectively we are a market that a vendor can build a product for
2. Since we have significantly more revenue than a startup like AirBNB we can rationalize and amortize an ongoing expense
3. We actively try not to do things that aren't our core competency
4. Even if we wanted to we have plenty of human resources but not the talent to build our own analytic DB or data visualization tools
I think AirBNB is likely a much different company than most. Their approach to talent attraction/retention, novelty of problem set, ability and tolerance to make a multimillion dollar capital investment in analytic database and visualization tools, etc, etc. In their case the upstart cost of building a custom thing and the value prop likely makes sense.
This is why you never try to target developers and if you do, you have to make sure the initial cost to use it is negligible. This is what Atlassian does very well, with their pricing structure.
In the past, I developed in house tools for massive Enterprise companies and the common theme has always been, we can do it better and cheaper. Sometimes that's true, sometimes it's not. I noticed this trend slowly tapering off before I left to start my startup and this was due to the availability of better open source solutions.
Technical people always overestimate what they are capable of, present company included, and if you are set on selling to developers, you have to make sure the barrier to use is negligible. By at least trying, they will either come to the conclusion that what you have created is a piece of shit or they realize it's not worth their time and effort to duplicate/maintain.
You do realize that this is probably THE reason to open-source a piece of custom software after you've built it, right?
...I imagine that they expect the community will do X% of the maintenance work now, so they can save Y%, so win-win for both the company and the community. Also, you increase the probability that any new hires will be at least mildly familiar with your in-house stack if you open-source some of it (think TensorFlow).
Free as in beer is the best answer when you have reasonable expectations that the community will really embrace the product and you'll get back "free as in beer" upgrades and bugfixes to the software you won't otherwise have the budget to properly maintain (and properly document, btw! think of all the free tutorials that will be written for this after it gets popular!) :)
[0] https://en.wikipedia.org/wiki/Online_analytical_processing#M...
Disclaimer: I work on Metabase.
71% of those who chose to build the BI tools said they built because "We can customize the functionality better"
51% of those who buy say "Buying enables us to provide best-in-class BI functionality"
The study: http://www.jaspersoft.com/sites/default/files/confirmation_f...
Grouping imports into: standard lib, third party, local is a strong pattern that I don't see done consistently in many repos. Likewise with your use of wrapping long imports with ()s and a single tab.
Any chance of sharing your Python style guide? My startup is Python based (Django and Flask) and would really appreciate it!
https://camo.githubusercontent.com/c22acad6c1302c5da3236cb8e...
Here is the original demo[1] from Mike Bostock, D3's author.
I have done a fair bit of Ruby a few years ago but I'm new to python CRUD apps and trying to improve my knowledge here. Is defining all models in the same file[1] conventional in python apps? Rails used to have separate files for each model. And most Ruby apps that I have seen advocate the one-class-one-file convention.
[1] https://github.com/airbnb/caravel/blob/master/caravel/models...
For example you have a comment app that could contain several models: Comment, Thread, Report, etc those can be in the same file. To continue on the django example, I would personally prefer having a models folder in the comments app and one file per model as some can get really big.
I also do 1 file / model in Flask, minus some specific cases where it just makes sense to have them in the same file
If you use multiple apps within one Django project or the equivalent in Flask (Blueprints), that extends to one models.py per app (where a "project" is a collection of "apps").
Sometimes you'll see one file model per (with a models/__init__.py that imports them for use). While I think it keep dependency imports for each model very cleanly separated, you end up having a lot of redundancy importing the same basic pieces in every model file.
If you have a small number of models (e.g. <= 5), then it's fine to have them all in one file, as you will not benefit from multiple files, really.
When your application is growing, you have split the models into multiple files, grouped by features, etc (e.g. users.py, content.py, etc).
I prefer this as models usually a very small, and switching from one file to another can become quickly annoying when working on related models. However, it may be different for large classes.
Our site is here: https://www.periscopedata.com/ and if you have any questions, shoot me an email at jon@periscopedata.com.
A side note: I love that Caravel is written in Python. Metabase switched from Python to Clojure, and I personally believe that to be a barrier to entry for contributing. I've started down the Clojure path a number of times only to stop because I see it as something which would be difficult to impose on my team...Lisp is just so different from what most enterprisey development teams are used to. Python, on the other hand, is easier to justify. I've found things I wanted to help fix in Metabase, but having to learn idiomatic Clojure just to submit a patch is a turn-off.
If you don't mind, I'd love to hear why you felt we were far from being ready for production.
Our main challenge has trading off ease of installation and use with the inevitable feature creep that causes the lots of moving parts and learning curve you find meh about Pentaho. You're right in that we've focused on making the most basic of sql queries (via our non-sql tool) usable by anyone in the company on their own vs essentially producing an analytics SDK like Pentaho/Jaspersoft.
On the language front, we made a conscious decision to optimize for ease of installation and low maintenance overhead. While I agree, it's made contributing less accessible, it's been amazing how porting has made Metabase more stable and easier to install. We run a bunch of instances for people, and our ops footprint has been silly small.
If I recall, it was difficult to understand how to properly format the results of a raw sql query to graph. Also, even once properly graphing, saving the graph to a dashboard wouldn't scale properly and would render in a very jumbled manner. Perhaps these issues have been fixed with newer versions....I'll give it another look next week.
Disclaimer: I am one of the founders.
If you're ok pulling the data out of Postgres into memory locally and mostly care about manipulation and beautiful dataviz, then look at Tableau.
If you're mostly interested in more data sciency/ML stuff, then Shiny or something else that's R-based is a good option.
If you're interested in being able to embed your business logic into the tool so that non-SQL folks can build their own queries and everybody's relying on the same data definitions, that's where Looker (disclosure: where I work) excels.
On your other point, though, to echo the build vs. buy discussion from above, I think it's a bit misleading to say "oh, we'll just use an open-source solution and that'll be cheaper." Because if open source means a couple of internal developers and an analyst, that's easily $300k+/year in salaries that you might not spend if you were using a vendor.
Anyway, given your particular statement of the problem you're facing, I'd humbly suggest you take a look at Looker. The data modeling layer that's core to Looker is meant to solve EXACTLY that problem, by leaving your data where it lives and then embedding your business logic in the layer that sits between end users and the data.
Written documentation is vastly superior to videos in my opinion.
I absolutely loathe video for any analytics-related documentation. It rarely adds any real value over text outside of live webinars where I can ask questions.
Got it up and running easily enough, and connected to Redshift. But seemed like creating a new "slice" required custom JSON params to define it. Unless I missed something?
edit: yep, missed something. Can "explore" a table by clicking it's link in the table listing.
<yourPythonInstallDir>\Lib\site-packages\caravel\bin
then run as
python caravel db upgrade
Database Support
Caravel was originally designed on top of Druid.io, but quickly broadened its scope to support other databases through the use of SqlAlchemy, a Python ORM that is compatible with most common databases[1].
A tutorial on how to link it to a mysql database would be greatly appreciated :)
So, you just need to set the config param SQLALCHEMY_DATABASE_URI like this:
https://github.com/airbnb/caravel/blob/1b4e750b2aa111445703d...
The configuration guide explains it further:
https://github.com/airbnb/caravel/blob/master/docs/installat...
This is just a data visualization platform. You need to bring your own data store and data.