Voila – From notebooks to standalone web applications and dashboards
voila.readthedocs.io
voila.readthedocs.io
I can create a GUI for a tool that looks nice faster than I can make a CLI. I've built useful production systems (ok, sure, for internal use) in literally minutes.
You're a bit limited in what kinds of apps you can make but the tradeoffs it makes here means that it's astoundingly easy to make a wide range of very useful tools.
https://data-dive.com/multi-os-deployment-in-cloud-using-pyi...
It uses pyinstaller to build and even pushes the build as a zip into your release page on github and appears to be working quite well.
https://news.ycombinator.com/item?id=19859913
As a note to OP, messaging your solution as “turn your notebook into an app” may not be optimal — you will loose many who abhor working in notebooks.
[1] https://panel.holoviz.org/getting_started/index.html
[2] https://github.com/holoviz/panel/pull/1983
So far, the only low-code PyData framework we saw that avoids most "JS in Python" is StreamLit. However, even there, it is still awkward in practice, so we still see limited adoption by folks who are fine with notebooks, so rarely goes beyond a champion. So there is room to grow.
Here is an example https://panel.holoviz.org/gallery/layout/distribution_tabs.h...
Ex: It took me awhile to appreciate StreamLit's builtin layout: it largely eliminates "HTML-in-Python", so one less thing. Likewise, decorators are weird magic, so yet another educational hurdle. If a tool could do excel -> dashboard, most would rather that! While I love that stuff, and I can recommend it to coders, I've learned to not recommend it to teams that can't guarantee everyone is... which is most. It sounds like Panel is slowly reinventing StreamLit, but as StreamLit isn't even there yet, for most corporate use, I'd be trying to do much more than catchup on this specific aspect if you want it to be relevant here.
Fun story: a PhD friend for a much-lauded company on HN led a team of ~20 analysts. About ~2 people loved Python, and the rest would write pages of SQL to avoid it.
Regarding your second point, looking at streamlits announcements page [1] it seems many features being added, layouts/themes and session state/callbacks for example, indicate to me that streamlit is heading in the direction of panel/dash more than the other way. Streamlit also emphasizes using decorators for caching [2], which I agree can end up with some overall state that is a bit magic.
Overall I think the optimal use cases are a bit different for the sets of tools, and I find the approach of dynamically calculating a full script for every change of a slight widget quite onerous, and actually just not feasible for the use cases I have.
[1] https://blog.streamlit.io/tag/announcements/
[2] https://blog.streamlit.io/six-tips-for-improving-your-stream...
Though again, I'm not a zealot: the StreamLit starting point is still too high in my experience for most teams, so both are wrong. For most people, the default should be no Python, at most SQL or whatever DB lang, and optionally drop down to Python for some cool bits. Ironically, I just got off a call a few hours ago where this exact issue makes us excited about starting with StreamLit, yet we're also already scheduling tools to replace it with something more realistic for 10X+ wider enterprise adoption.
I want my physicists to spend their mental effort on physics, not on software architecture.
Again, if it's not possible, then you accept the imperfection of the result, or provide better tooling.
There is no blame to put on them whatsoever.
Similarly, they should hire experienced, qualified software engineers to write/check the software in their papers.
They don't because 'everyone can code - its just logic'.
You might be mixing up the terms. I don't think that the point was about "scientists" in general, it was about "data scientists". The first is a common term used to describe someone who does science in some professional capacity. The second one is a very broad job title within software which very often includes writing code that ends up in production - at some data science roles that might even be your main responsiblity.
A data scientist is not even a somebody trained as a programmer. Their strong suit is data analysis, and it turns out one of the tool to manipulate data today are programming languages so their do it.
But I as a Python trainer, I train data analyst regularly, and they don't have a clue about language ecosystems, how the OS work, data formats or reliable software architecture.
They mainly want to output their graph, pdf report or other media to serve their conclusion. They may want to create some reusable algo, or machine learning model, but that's the limit most of them hit.
If one take their code and put it in prod (which I know happens, don't get me wrong), that's not the data scientist fault. They are doing their job, in which programming is just one of the many means to an end, and is not their specialty.
This is like teaching some JavaScript to complete beginners at a bootcamp and then declaring that front-end developers aren't real programmers because they know so little.
I'm not talking about complete beginners. I don't train beginners.
Beginners not being aware of some best practices doesn't automatically make them not-programmers.
I work as a data scientist and I see it as part of software development. It's just a different domain - some people do front-end, some do mobile or embedded, I do data science.
Just like I'm not a data scientists, I'm a programmer.
Now, I can use pandas in a pinch and makes pretty graphs, but my statistical analysis will never be on part with yours.
Just like a pianist hobbyist will have a hard time to rival somebody who does that 40 hours a week, although he may be able to play a few fantastic pieces.
Hell, even a web dev programmers, if ask to code a GUI desktop app, is not going to do a good job.
IT is becoming a very large field.
And scientists are not even from this field.
That's not my point.
Physicists are not mathematicians, and yet they are required to acquire a relatively high degree of proficiency in maths because maths is a fundamental tool in their job, and nobody would argue otherwise.
The attitude of considering programming a mundane craft to be picked up as-you-go is the main reason why the scientific software landscape is such a shitshow.
/rant
Just like if you are standing on your two legs most of the day and sprint once in a while, you can be considered a runner. Sure, you can play with semantics, but most people cannot run a marathon.
> The attitude of considering programming a mundane craft to be picked up as-you-go is the main reason why the scientific software landscape is such a shitshow.
Err... That's kinda my point?
> And as such you should be expected to become a decently proficient programmer.
It's very, very hard to be good in 2 different fields. Most people won't have the ability or the context to do so. Even if they did, the time and energy spent to do so would be taken from their main activity, which is why we employ them in the first place.
It's not reasonable to ask a data scientist, geographer, biologist or physicist to follow up with the right practices to deploy the latest sci-stack on a linux server, understand the trade off between GIL locked python thread, asyncio and multiprocessing or spell out what WSGI stands for.
Hell, I know a lot of professional programmers that don't know those things
> Physicists are not mathematicians, and yet they are required to acquire a relatively high degree of proficiency in maths
The quantity of information required to be learned is of one or two orders of magnitude, because the field of maths required to perform physics is quite stable, and well understood.
IT is a very young field, in constant flux. The scientific stack is a moving target, not to even mention the web one. Nobody can expect them to understand python, numpy, pandas, then a web framework, then css, then js, and html, probably some frameworks for them, a builder or two, how to deploy all that stuff in dev, in prod and architectural concerns for linking all that stuff.
That's crazy talk.
If you don't understand the tools you're using, or the environment you're in - you're not any more of a "data scientist" than pretty much everybody else. My carpenter is a data scientist going by this logic.
>The quantity of information required to be learned is of one or two orders of magnitude, because the field of maths required to perform physics is quite stable, and well understood.
Apart from the fact that some areas of physics are really at the forefront of maths, this also ignores the fact that learning the level of proficiency required for graduate work in physics is significantly more involved than learning about some best practices in programming.
I've seen data scientists handling big code bases. The problem was not they couldn't use the language features. The problem is that they would be always lacking essential information for their mission because their is not enough time in a day for a regular human being.
They would put a md5 hashed password in their db, create an xml format to be reusable only to realize they'll need to hard code some value later, or have a gunicorn running to a crawl because they didn't know how to calibrate the number of workers.
It's just too many things to know. Once they mastered that, other things would come to bite them.
It’s definitely a part of it. This isn’t an all or nothing thing, one can learn good practices without encumbering their scientific work.
However, the programming education in science degrees is absolutely appalling. Just show them how program a newton raphson method in matlab (without any considerations for performance) and expect them to know how to program.
It doesn't mean of course that everybody is supposed to be an expert programmer, but a minimum effort to help your colleague is surely not too much to be asked.
Whether you choose to call it "good" or "not too bad", there is a minimum bar of competence that data scientists need to meet in statistics/ML/AI, programming, and their domain of application. And the ability to move from Jupyter notebooks to Python modules/packages is a basic.
I'm a physicist (not "data scientist" though I work with plenty of data), and I've been programming since 1981. Anything I do, I want to do well, especially if I do it regularly or it could cause problems if done badly. I've made an effort throughout my career to keep up with good programming practices. I do that out of a combination of pride, curiosity, and professional ethics.
But I'm not a software developer, meaning that I don't create software for widespread or long term use by others. We have an entire department for that, and many of their techniques are quite specialized.
Naturally it wouldn't surprise me if further improving my skills also moves me closer to being capable of software development, and I'm happy to learn and apply their techniques at a pace that works for me. I think that a scientist who is capable of learning to program should receive guidance on how to do it better, but perhaps in stages, such as:
1. Writing code that has a better chance of working, even as it gets bigger and more complex.
2. Working with others on projects that involve sharing code, meaning that it has to be readable and conform to agreed upon standards.
3. Creating code that can be confidently "shipped" for widespread or long term use.
Later it was implemented in production with a regular stack (Flask + Vue).
Voila is really empowering for e.g. data scientists that are comfortable in a jupyter environment but aren't js wizards. Running locally, I just love the reactivity it provides: you don't worry about sync between front-end and back-end, everything is propagated through websockets I believe (Jupyter is Tornado-based).
However, for production you might want to use another tool, since it (currently) executes every session in isolation, so every time an user connects it re-runs everything from scratch. Moreover, the round-trips to the server can be slow if you are e.g. in a different continent so this degrades the UX.
Here is an example of a small ML app I built with Voilà (this will probably crash due to HN hug of death™), and JAX on the backend: http://grad-descent.herokuapp.com/
Suppose I write some educational Jupyter notebooks, which are not particularly resource intensive, say 100 seconds of compute time per notebook. I host them on some cloud server, using something like OP, and get a 1000 people to learn from it. Maybe they end up using say,
1000 people x 5 notebooks x 100 seconds/run x 20 runs of each notebook = 10 million seconds of compute time.
How much would such a server cost to host, where "many" of these people are working on the notebooks together? Just need a rough estimate.
How many users are accessing the notebooks concurrently (e.g. all 1000 or only a dozen at a given time)? Is there any downtime, i.e. do the users come from the same time zone, so that app can have inactive hours (say it's OK to be unreachable during the night)?
Depending on the specifics, free hosting may be available (e.g. via Heroku, Google Colab, AWS Free Tier etc.).
As far as paid offers go, this is way too unspecific to be answered in a meaningful way. The answer depends on the actual resource requirements (RAM, storage, data transfer, CPU cores), estimated usage patterns (concurrent users), and your location.
TBH, if no commercial interest is involved, just hosting the notebooks on Github or making them accessible via Google Colab would be the easiest option.
We don't want options like Google Colab, because we want the experience to be tightly integrated (there are also issues around GDPR). So we want to run our own server.
If we can run a ten-day workshop of 500-1000 people, where most people work everyday at the same time in a 5 hour slot, and keep costs under 50-100 usd, we would make the switch. But I understand that it is difficult to make estimates without trying how much resource usage there actually is.
Stuff like Google cloud managed notebooks are also pretty cheap, you can create a template, and one-click on demand create, kill it at the end of the day. There is an option for 7 cent/hour per user. 5 hours = roughly 50 cent, so above your budget but soo easy from a management point of view. And infinitely scalable.
You could also start messing indeed by just installing it on a cloud server and only turning it on for those 5 hours you actually need it.
And if you still want to install yourselves, look at tlhj (the littlest jupyter hub). We use that succesfully internally, but it takes a few hours to get everything configured the way you like it. Sill a lot better than actually installing a proper jupyter server imo.
Just contact local(!) hosting providers and ask if they could sponsor such events. This would mean advertisement for them (maybe even tax deductible depending on the legal status of your organisation) and free resources for you.
I can't think of any kind of on-demand service that will handle 500-1000 concurrent users for 50 hours that's under 100 USD. AWS nano instances are 0.256 USD per user per 50 hours (e.g. your workshop scenario), but that's still above your budget even with 500 users. Basically you'd need to find an on-demand hyperscaler that offers instances with ~1GiB + 1vCPU for less than 0.004 USD/h (500 users) or 0.002 USD/h (1000 users).
Working on providing a simplified local install method (e.g. a docker image or a VM image hosted somewhere cheap) is the only realistic way to stay within your budget.
We provide this no matter the season and usage is very seasonal for us ;) We also provide way more computational bandwidth that would be necessary, so I guess you can provision this for half the cost, just be sure to put out the right restrictions for resource usage.
At half cost that's 250 dollars a month, which unfortunately is currently outside our non-profit budget.
There's Starboard (which I'm building, it's built specifically for the browser and can integrate into a larger app deeply) and JupyterLite (the closest you will get to JupyterLab in the browser), either can be a good choice depending on your requirements. Both use Pyodide for the Python runtime.
[1]: https://github.com/gzuidhof/starboard-notebook, demo: https://starboard.gg
"Problem: package xeus-cling-0.12.0-h5a79028_0 requires xtl >=0.7.0,<0.8.0a0, but none of the providers can be installed"
sigh
I don't know when it started, but it seems to be a recent trend to add dependencies for anything and to package everything on demand. It's probably for security or something. But I do miss the days when people would link a static binary that "just works" even without internet and that'll keep working a week later, because it includes all of its dependencies as opposed to downloading and updating 500 packages on-demand.
That's why tools like Anaconda and Docker were created and now even a simple utility can use gigabytes of disk space...
FYI, xeus-cling is a Jupyter kernel for the C++ programming language.
I often think that with the current state of Javascript in the browser, we could build an awesome, super fast, Jupyter style notebook software that runs completely in the browser. With the modules implemented as native Javascript modules which are dynamically loaded.
Is anybody working on this?
I have built a rough version of this idea for myself and been using it for my own statistic needs for a few months now. It is far from being polished/flexible enough to be useful as a general purpose notebook though.
Agreed that for some things, it would be great to be able to explicitly offload to the browser.
I have not yet tried to load a lot of data into it. Would be interesting to see when the load time starts to outweight the benefits of instant calculations. Maybe at something like 10 million datasets? Hard to say.
It's a very promising technology IMO. Not necessarily because it is faster (which is only 0.3x-2x according to my findings) but because it is a nice, simple compilation target.
Here is an example instance: https://notebook.basthon.fr/
The thing is, creating a whole stack in pure JS would be very hard, since the current scientific stacks uses a lot of fortran, assembly and C with python to bind them all. Or julia. It's millions of man hours we are talking about.
So compiling Python into WASM is probably the best deal for such an app.
For the regular web, it would be a deal breaker: you don't want to load 15 mo of runtime before being able to interact with a web page. But for such a scientific app, it's not a problem. Besides, were you to write it entirely in JS, the size would be huge as well.
Language-wise, the other issue JS has is that it lacks operator overloading, which makes array handling a lot nicer. Python has magic methods for doing that. Julia, R and Matlab have that built into their languages. Julia also has it's own version of reactive notebooks similar to Observable.
And then you have a lot of scientists who already know and use Python, R or Julia. I can't really imagine a statistician who's well versed in R finding much value in JS. Javascript just isn't made for complex statistics the way R is.
JS now feels like Java back at the turn of the millennium, when people were thinking Java could just be used for everything, and it sort of was. Or any language could run on the JVM (the web has largely replaced that idea). But there's a reason for the various different programming languages. Some are just better at doing certain things. And JS is not a scientific computing language.
Yes, because it's called "Voilà" :-)
But I agree, starting the title with "Turns" is pretty bizarre.