Show HN: Mito – Write Python/Pandas faster by editing a spreadsheet, in Jupyter
trymito.io
trymito.io
We started building Mito 6 months ago after finishing up undergrad engineering + business school (aka Excel School). We became comfortable doing data analysis in visual environments, but were held back by Excel's 1M row limit + the inability to create repeatable processes. Doing analyses in Python was wayyy more powerful, but also required tons of trips to Stack Overflow + pandas documentation.
After a few months of building, pivoting (our vision) + lots of refactoring, Mito now supports writing spreadsheet formulas, pivoting (dataframes), merging, saving + applying macros, and the tiniest bit of graphing. And it generates the equivalent pandas code in real time for all of it :)
You can download Mito by following our always-WIP documentation [0]. We'd love to hear your first impressions of the tool (especially if you download it) + your experiences in/building for the data science/analytics community.
Open-Source:
- dtale: Very similar to Mito. On top of mito, it provides basic data exploration and can also be launched from within VSCode. Also exports pandas code but not inline into a cell.
- pandasgui: alternative that does not export code, yet.
Not Open-Source:
- bamboolib: very similar to Mito e.g. also has code export into a cell. The basic version is free on local Jupyter Notebook and Lab. bamboolib does NOT allow inline writing of Excel-style formulas like ROUND(A1, 2) like Mito does. On top of Mito, it supports more pandas functions e.g. also datetime handling. Data explorations for the whole table and columns. It has a plot creator for creating Plotly graphs. It does not log any user data - neither about feature usage nor about the actual data. It also works in Jupyter Notebook. Enterprise customers love the ability to extend bamboolib with plugins in order to add their own custom plots or data transformations. Also, bamboolib supports data loaders e.g. to load CSV files from a GUI - Mito currently seems only to work when the data already is available in a Dataframe variable. With bamboolib the user does not have to code anything in order to spawn the UI. The user can just type the name of the dataframe. For Mito the user needs to type mitosheet.sheet(df_name). bamboolib is more mature because it is roughly 2,5 years in development and has many enterprise customers like Spotify, Bain & Company, Procter&Gamble and 2 of the top 10 global asset managers.
Full disclosure: I am a co-founder of bamboolib
Not questioning the need for this, rather being continually surprised how diverse the Python userspace is, even within a "field", like "data science".
If I can ask - why do you try and avoid point and click tools?
Maybe it’s my perception, or maybe a function of the kind of tasks I do, dunno, but I spend a fair amount of time to go away from GUIs if only I can
Maybe I'm too snobbish, but I don't trust a company with my data that doesn't know basic HTML (I would install the local version though).
Also the github links don't work (npmjs and python repositories show that it's using BSD license): https://github.com/mito/mito
Or if you don't have Jupyter already set up on your computer, we have a hosted version of Jupyter Lab that you can make an account on.
https://pypi.org/project/mitosheet/#files
You can get the source code for something from that page, but it seems you are trying to steer people away from it as much as possible. At least, it doesn't seem to be on github (the horror!).
Is this just because you haven't cleaned up the source code as much as you would like, or does the open-source portion need some proprietary component to work?
However, I installed the package locally, and I see that importing it causes a request to be sent to segment.io.
If you're interested in using the tool without any logging, just let us know and we can make sure to turn it off for you!
For the spreadsheet formula part, we parse the formula in the Jupyter Python kernel and convert it to Pandas code. We actually found a bit of a trick to make this easier. TLDR: We defined Python functions that have the same name as the spreadsheet formulas, it helps us avoid building writing a formal language grammar + syntax tree. If you're interested in reading about it a bit more, we actually have a short blog post that you can checkout: https://trymito.io/blog/transpiler
The second type of translation (merging, pivoting, etc) sends a message to the backend with the parameters configured from the point and click tool. And then executes the equivalent Pandas code in the Python kernel and writes the code to the Jupyter Cell.
There's no mention of the Segment tracking in the docs, and I don't see anyway for the user to opt out of it, which I think is an immediate GDPR issue.
Given that you are logging metadata about the dataframes in use along with the user email and name of the logged in user, I can't see this ever being used in an environment where sensitive data is being processed, since it could potentially leak PII that's easily tied to a given company via the email address.
This is a great idea, and I think if you can go with the BSD license and provide a way for people to opt out of tracking (or ideally flip it and allow them to opt in) this could be used in any number of industries. As it stands currently I just don't think this will ever pass a data audit at any large company which is a real shame.
For our current users who have told us that they are not comfortable with logging, we have been able to turn off logging for their specific accounts. So if you're interested in continuing to checkout the tool while we make those improvements, just let us know.
Since we only support Jupyter right now, about half of the early Mito users are using it locally and the other half are using it on a hosted version of Jupyter Lab, which just makes it really easy to get setup without worrying about Python installations.