Automated PDF Reports with Python Notebooks
mljar.com
mljar.com
jupyter nbconvert --execute example_report.ipynb --no-input --to html
I have never been able to get a nice looking report in both html/PDF though via the tools to print html to PDF (as opposed to building a nice looking PDF doc using LaTeX directly). So default just worry about making the html look nice.- you can add widgets (parameters) to the notebook, without rewriting the notebook,
- you can easily schedule the notebook,
- you can convert notebook to PDF with mouse click,
- you can serve interactive HTML notebooks with a Django-based server.
- Convert dataframe and other stuff to HTML render them to usable HTML using Jinja templates.
- Generate PDF from html using Pandoc with Xlatex
I'm generating beautiful html/pdf reports with a simple bashscript, and minimal markup added directly to the source code. Works like a charm.
But if I did have to pick "the one" (I see you trying to pull a Scott Adams, good stuff) it would be that it has the semantics of Literate Programming completely backwards.
Literate programming is about 'code' first. It should be that you have a codebase which honours code semantics first, and allows you to create a meaningful report from that code second. This guarantees that you have a codebase structure that is easy to traverse and read, while still following good software engineering practices and usable as code in itself, as well as being able to trivially create a report from it.
Jupyter notebooks is an app (not even just markup, but an app, and a clunky one at that) for writing reports first, in the form of snippets from which you may or may not be able to generate useful code or outputs later on. And ultimately, even if you do, that code is crappy and unusable in any context other than the notebook it was written for, because of the way jupyter notebooks are designed. They force the programmer to create a monolithic spaghetti structure, preventing modularity or meaningful code hierarchies to take place in the codebase, promoting imports and definitions appearing just before the point of their use rather than in reasonable scopes, promoting modifiability instead of extensibility, while preventing programmers from using a vast majority of appropriate software tools for versioning, testing, etc.
And for some bizzare reason they are becoming the de facto standard for data scientists sharing code. It's like we're encouraging people to return to "Academic Matlab" code-quality standards all over again after years and years of trying to teach academics proper software engineering practices.
I agree that in terms of literate programming as Knuth defined it, Notebooks are not great. There are tools to improve that story; I wrote https://github.com/agoose77/literary which at least lets you do a bit more "tangling and weaving" than you can out of the box. It doesn't let you define functions in arbitrary order, or implement fragments of a code block, but it does let you "boil down" a literate representation into something that is zero-cost at runtime and imports. There's also nbdev, although it's not my cup of tea.
The real point, though, is that most data-scientists aren't using (imo) notebooks to write and share libraries of code. Instead, they're using notebooks as semi-reproducible reports. I'm a physicist, and that's what I've been using Jupyter for. For me, Jupyter Notebooks are fantastic - the cell mechanism lends itself to rich-outputs that augment the narrative, and present the information in-line with the code that wrote it.
For me, the biggest gap here is writing _libraries_ that are leveraged in these notebooks. That's why I wrote Literary - to try and resolve some of the pain points that currently require you to use two tools (Jupyter Lab & e.g. PyCharm). I'm not saying it will work for everyone, or solve all of the problems, but for me it's enough to write my analysis as a package, so that's a limited success in my book.
When it may take 5 minutes to simply load the data, you can no longer rerun some Python scripts to mess around and experiment with stuff. You need that kernel with data and everything else to have 100% uptime.
Similarly, you may want your temporary experiment results that may have taken a while to compute to stay within the kernel even if you're already working on something else.
That, plus an ease of inline visualisation, displaying tables and all that.
Standalone HTML files provide a really nice alternative to PDF as they maintain interactively: you can host them static sites, allow people to download data, use plots interactively, click through pages, similar to a statically generated website. That said, there is still a definite blocker in non-technical people receiving a .HTML file over email and immediately thinking it's suspicious or a virus (doesn't help that gmail has such poor support for them.) It's a shame, because PDFs have so many warts and HTML can be used as a really nice distributable file format, especially as you can make them fully standalone by baking in datasets, plots, libraries, etc. so they can be used without network.
IMO Jupyter is great at what it is - a REPL - but, outside of sharing a step-by-step "here are the steps I took to come to this answer", isn't the ideal format for sharing insights, as there is no reason a report would follow the same narrative as the analysis itself.
with open('template.tex', 'rt') as f_in:
template = f_in.read()
template = template.replace('%PLACEHOLDER%', data)
fname = f'rendered-{dt.now().isoformat()}'
with open(fname, 'wt') as f_out:
f_out.write(template)
fileout = 'outdir'
subprocess.call(['pdflatex', '--output-directory', fileout, fname])What is more, I'm thinking about including Voila into Mercury. So you will be able to serve Voila apps with Mercury. The end-goal is to make notebooks sharing easy.
- Ingest some client data.
- Process client data - so the bit that adds the value for client, normally this will involve some cleaning, normalization, algos, external enrichment.
- Produce human readable output as a pdf (tables, graphs, charts etc)
- Repeat this on a set schedule.
This certainty looks like a solid foundation to use to solve for these types of scenarios
Probably too resource hungry, or some flippery making Jupiter work with it. I've seen Jupiter all over the place but haven't played with it.
Recently I did similar for displaying numbers in the notebook as a good looking boxes. I created a small Python package that takes the number and creates a box with borders with HTML+CSS. It is pretty handy for building dashboards in Python. The package name is Bloxs https://github.com/mljar/bloxs