Franchise – An Open-Source SQL Notebook
franchise.cloud
franchise.cloud
* A simple, flexible SQL client
* A notebook interface that lets you drag stuff around and view things side-by-side
* Query CSVs, JSON, and XLSX documents fully in-browser with Emscripten + SQLite
* Postgres, MySQL and BigQuery through a local connection bridge. Your data never touches a third party server.
I’ll be here with @antimatter15 for the rest of the morning answering questions!Any plan to support jdbc so that it’s possible to connect to things such is HiveServer2/SparkSql?
I see that you are using Reactjs. If you can build something that can connect to spark and python, I will pay for this and so will lots of people.
One of the big unsolved issues where we want to really throw money is have a notebook that my analysts can work on and then I can take straight to production and have it run multi-user.
There are lots of startups attempting to do this on top of Jupyter, but I believe building this from scratch is the right way.
I'm literally crying right now for a reactjs + d3js dashboard running on pyspark that multiple people can use simultaneously.
A few questions/comments:
- Are you planning to put this in an Electron app? I think the more you can make it feel like an actual desktop app, the more likely you are to get more "regular" users. Just my opinion.
- While I love the simplicity of making this sound like a cloud service, I think that the nature of it being provided through the user's own infrastructure should be emphasized even more. As far as I can tell (and I haven't looked extensively), there is nothing that is actually "cloud oriented" about your service. Is it because people make there own "cloud" out of your bridge adapter? For people who want these tools on-premise and under their own control, you have a solution.
Yeah the domain name is pretty much a complete misnomer (it was a cheap gTLD :/).
It's sort of a mini ETL powertool.
* I really like the editor. It beats metabases ACE editor anyday.
* Here's the line graph comparison between the two: https://imgur.com/a/gd5P8 I think that metabase wins here because I can get more information off that graph rather than an area under the graph kinda thing that Franchise is trying to do. Is it configurable somewhere?
* Can I link other users to the question somehow?
Great job OP and I'll be sure to try this out as we are having lots of difficulties managing data at my workplace.
1. I'm dealing with a dataset with >10M rows. Is there any way to kill a query that is currently running? I noticed at least with postgres that a long running query would stop me from performing any other queries.
2. The hashtag notation is very nice but would it be possible to cache the results of the hashtag without having to run the same query over again. This would more closely resemble the behavior of jupyter notebook, where results in previous cells can be recalculated by running those cells, but can be quickly accessed without recalculation.
For existing SQL tools there's currently basically two approaches: where either the full application is hosted, or where the full application is local.
Jupyter and Zeppelin generally fall into the second category— you have to set up Docker (or the myriad of dependencies), and edit a bunch of configuration files to launch a local server before getting started. If someone sends you a notebook, and you just want to read its contents and see the interactive charts, you still have to go through this ordeal.
On the other hand, there's Mode, Redash, and PopSQL, which are hosted solutions. Editing configuration files is replaced with a one-time setup and registration, but you have to entrust these startups with access to all your data.
Franchise takes a hybrid approach— it's a hosted static web page that contains all of the display logic, so if you save a notebook and share it with a friend, you can open it and it just works. To connect to a database, you start a local bridge which tunnels data from Franchise directly to your database through your computer.
We spent a lot of time thinking about the interface and trying to make it easy to use, but still powerful. We built a new notebook layout engine that allows you to run two queries and see the results side-by-side. Deleting a cell sends it to an archive, so you don't have to worry about losing data. Data visualizations are accessible with one click. And if you're just trying to query data on a CSV or JSON file, you can just drag and drop it onto the page.
One correction about Redash: Redash is open source and you can host it yourself too, in case you have an issue trusting a 3rd party with your data.
The proxy approach is something I thought about several times, but I am aiming at providing a collaborative tool, where it's easy to share the results and work with others. Depending on a proxy running on a user machine breaks this.
But having said that, I'm sure there are many whose needs Franchise will fit very well. So kudos on launching a great tool! :)
I'd like to confirm, we are hosting redash 2.0 (recently updated from 1.0) to our own server and use it as a query portal and collaboration tool to a number of databases (including mysql, postgres and oracle). Query sharing is really important there are users that don't know SQL but know how to run "peter-21" or "jim-32" with redash.
Everything is smooth till now thank you for the great product!
> Disconnected from PostgreSQL (Error: Too many result rows to serialize: Try using a LIMIT statement.)
This is on my console (my credentials replaced with ellipses):
~> npx franchise-client@0.2.2
npx: installed 140 in 12.533s
franchise-client listening on port 14645
opened connection
received: {"action":"get_postgres_credentials","id":1}
received: {"action":"open","db":"postgres","credentials":{"id":1,"host":"localhost","user":"...","database":"...","port":"...","autofilled":true,"password":"..."},"id":2}
received: {"action":"exec","sql":"SELECT table_schema, table_name, column_name\n FROM information_schema.columns \n WHERE table_schema not in ('pg_catalog', 'information_schema', 'pg_internal')","id":3}
Error: Too many result rows to serialize: Try using a LIMIT statement.
at Object.query (~/.npm/_npx/736/lib/node_modules/franchise-client/response.js:128:23)
at <anonymous>
at process._tickCallback (internal/process/next_tick.js:188:7)
I ran the SELECT query in Valentina Studio and got 11000+ rows in 4s. Maybe raise the acceptable row count to 100,000?The app mentions requiring the latest version of node but it's not clear whether you need the standard LTS (6.11) or the bleeding edge (8.5).
Thanks!
> npx franchise-client@0.2.2
franchise-client\server.js:16
ws.on('message', async message => {
^^^^^
SyntaxError: missing ) after argument list
at createScript (vm.js:56:10)
at Object.runInThisContext (vm.js:97:10)
at Module._compile (module.js:542:28)
at Object.Module._extensions..js (module.js:579:10)
at Module.load (module.js:487:32)
at tryModuleLoad (module.js:446:12)
at Function.Module._load (module.js:438:3)
at Module.runMain (module.js:604:10)
at run (bootstrap_node.js:389:7)
at startup (bootstrap_node.js:149:9)Basically, each visualization is a react widget with some static properties that define its icon / when to show it / etc.
I've been meaning to learn SQL for a long time, but I'm unfortunately in an unrelated field. Could anyone recommend some resources for learning SQL? Preferable targeted towards an experienced programmer
You will need to set aside all your imperative, procedural habits. To use SQL effectively you need to think in terms of sets, unions, intersections, and declarative statements.
If you find yourself thinking "row at a time" you might need to stop yourself. It's not always wrong, but a common mistake many programmers make when first learning SQL.
Basically, it looks for geo-ish column names.
That seems to be a pure local web-app. Kind of a remix of the local aspect of Tiddly Wiki and DabbleDB's approach to data massaging.
Too bad, I was looking forward to that mapping demo. It might be just what I need. I'll try again in a few hours.
node_modules: 329 715 665 bytes (435,4 MB on disk) for 37 881 items
I know it looks cool and animates and all that, but... really?!