Tiny Treeshaker: JavaScript tree shaking in 200 lines of code
github.com
github.com
I want a graph that shows me all the branches that got culled.
Even better would be a dependency graph that shows me what is causing things not to be culled.
I had an issue where Antd was pulling in 200KB of icons and I couldn’t figure out why until I spent hours bisecting it by commenting out entire sections of my application.
But yeah, I guess that would also require a VM like understanding of timeouts, intervals and the Browser's triggered events so it's kinda hard to do.
On a super basic level you can just build with and without tree-shaking and diff the results. If you turn off transpiling and minification, and move your bundled NPM modules to their own separate file, you can see what parts of your own code got shaken out pretty easily.
I'm tempted to have a go at writing a tool to do that actually. It would be useful.
I'm sure there are other places where it's applicable, but modern javascript and webapps have to be the quintessential use case.
To apply this to python is interesting - if you were creating a packaged version, I could see "compiling" the code to a separate package with only the required imports.
Those are both dead code elimination. Webpack even says:
> Tree shaking is a term commonly used in the JavaScript context for dead-code elimination.
For example a Lisp system may consist of starting a memory image from disk, which leads to a live running memory heap. The tree shaker then may break unused links in memory - either automatic or under programmer direction. The garbage collector then frees the unused memory. Lisp then dumps a new (typically smaller) image, which has unused code and data removed.
Such features appeared no later than end 1980s.
That's exactly how all dead code elimination works. How else would you do it?
Tree shaking happens at the import level. If I say `import { abc } from 'some-lib'`, tree shaking won’t include any other objects some-lib may export.
Removing a branch that can never be true, like `if ('production' === 'production') { … } else { /* dead code */ }` is dead code elimination, but webpack wouldn’t consider that tree shaking.
It could be, but I don't see it as too practical as dependencies are very likely quite transparent and don't themselves rely on many dependencies. But for JS?
There's no reason you can't write JS the same way you would write Java or C++. (And if you were writing "serious" applications in JS that weren't web apps before NodeJS came along, you did. Of course, you still can, too.)
I couldn't get Proguard to work with Clojure. Granted proguard is quite baroque so I probably did something wrong. But maybe there is some simpler solution out there for the JVM? Or would this simply be impossible b/c Clojure uses reflection? Though so does Javascript from what I understand..
I imagine though usage of something like exec throws DCE right out the window, poisoning the whole program
As far as I understand, and I'm murky on the details, but if you're language supports reflection then you can call any dependency during runtime. And so this is not amenable to static analysis. So the Java compiler doesn't have the same guarantees as say a C++ compiler.
Most people are running JVM on the server so they just don't care about executable size (and on Android people use proguard). But reflection in general is a source of headaches..
If you use something like GraalVM then it will try to prune dead code but it will not work well with reflections.
That all being said, on the specifics, I'm actually not entirely sure how you'd use reflection to arbitrarily call library code at run time. If anyone knows, I'd be curious to see an example
Which is exactly why ProGuard and R8 have so many flags and options, and basically every use of it requires customization: https://developer.android.com/studio/build/shrink-code
It doesn't seem like that powerful of a language feature (and more of a code smell) and it creates this mess down the line. I wonder what fraction of libraries even require it.. I'm guessing it's b/c of the flag soup i never got ProGuard working. I know the Clojure language/runtime itself relies on some reflection - unfortunately. That probably complicates things
Android specifically though: an unbelievable number of libraries use reflection at least a little, if not deeply. E.g. the Android UI framework is absolutely riddled with it - ever used XML to inflate views, i.e. the default way to build a UI? That's all based on reflection. Many (many!) of the most-popular libraries use it as well, though the efficiency-minded ones generally use compile-time codegen (annotation processors) where at all possible.
Thanks for the insight into Java internals and Android. It's a shame it's so tightly coupled at the roots but it's pure fantasy on my part to have a reflection-less JVM :) At the end of the day, there isn't really anything else like Clojure, so I just try to accept the warts and features.
Out of curiosity I took a cursory look at Qt and they seem to handle QML without using RTTI - so think it's all solvable - at least in most cases (I'm guessing by enumerating your possible types at compile time?)
c = Class.forName("tld.something.SomeClass").newInstance();
and if that string is created runtime, it's impossible to rule out usage of any dependency. So you basically have to tell the tools what to remove and what not to.Same strategy could be used for "tree shaking"?
I'll ponder how unused functions might be identified (during runtime).
I ended up skipping this step and distributing a JAR - which is a bit of no-no .. but i just ask user to install the Java runtime. (it's better than making executable for every platform and then spinning up VMs to test)
However, jdeps or not, the rest of the executable still remains as bloated as before
But if my understanding is correct, java classes are loaded as needed. So other than JAR size, you don’t really need it as much as in the case of JS for example.
End user apps are second class citizens on the JVM (except weirdly on Android.. but they're in a weird bubble of their own)
Tree-shaking is actually "live code inclusion", the opposite side of dead code elimination.
> tree shaking eliminates unused functions from across the bundle by starting at the entry point and only including functions that may be executed
Is what this code (tries to!) do
I.e., there are many algorithms you could use to eliminate dead code, and tree-shaking seems to just be a simple one that works well for JS.
function f(s) {
if (s === '') throw new Error()
...
}
there probably are very few tools that will statically be able to tell that f is never called with an empty string and thus the validation logic is not needed. It's JS after all and tools that mangle it all fall on a spectrum between doing very little, but very safely and trying to do too much and breaking code.1. Because of lower overall bandwidth (in many cases it is a SIGNIFICANT difference between code size before and after treeshaking)
2. Because of better startup performance - less code to process, less pings to the server for new modules, less load on servers too (streaming one file vs streaming hundreds of them separately)
3. Because of better performance (eg. removing conditional branches that will never execute)
;-)
You also don't have to tree shake up front because the execution will do it at runtime. If you deploy some unused modules, the browser will never request them and the server will never send them.
For the remaining benefits I can think of - more bundle bytes means more optimal compression; or shaking out functions within a module - I'm not so sure the benefits outweigh the complexity of maintaining and using bundlers. We could be splatting the src directory straight to the server and letting HTTP/2 do the rest.
What HTTP/2 brings here is a possibility to push to client files in advance so you can do a bit of "prebundling" on server side. Heard of couple different techniques of "prebundling" on server side BUT none of them will be as simple and performant as bundling it yourself :-) (and want to make sure that you realize that: there is no profit if modules will be requested from server at runtime)
And I guess you never tired things like terser (https://terser.org/docs/api-reference) so you don't know how much profit it brings not only in a bundle size, but also a startup performance and even runtime performance. Try it, play with different hoisting and mangling options (especially test mangling properties) and check the memory consumption, startup times and performance of your code.
Not sure what you're talking about and if you know what you're talking about here: "more bundle bytes means more optimal compression; or shaking out functions within a module"
But to give you a hint will tell you to: compress src directory of your project and compress bundled and minified source of your project. See yourself what performs better ;-)
Hope helped you!
Cheers!
I'm not looking for a mentor to explain the basics of bundler technologies or asset compression. I'm asking if HTTP/2 reduces the advantages of these techniques to the inflection point where the complexity of using bundlers is no longer worth the squeeze.
"more bundle bytes means more optimal compression" - this alludes to a single bundled file taking better advantage of the compression dictionary than many small files, who can't share dictionaries when compressed.
"shaking out functions within a module" - this is another form of tree shaking as opposed to file-based tree shaking (import dependency analysis). Unbundling could compete with the file-based case, but not with a deeper analysis of unused functions _within_ files.
These were the remaining advantages I could think of, I'm sure there's more. It sounds like you think bundling is still the way to go, and I probably agree, but it doesn't make any less interesting a discussion about its remaining merits and if we can iterate to close the gap. Bundlers, minifiers, sourcemaps, are all complicated build tools that require a hand on the wheel to maintain in the long term. We should not settle for them as defacto js ecosystem techniques.
Sorry for that, for my excuse it was not my intention.
> Bundlers, minifiers, sourcemaps, are all complicated build tools that require a hand on the wheel to maintain in the long term.
It just shows how bad is design of JS and web. But there are tools that do it nearly all with just one command, check "rollup" or my recently favorite "esbuild" (if you didn't before)
> We should not settle for them as defacto js ecosystem techniques.
I think it's done, it's settled, as JS itself, and if we like it or not, we have to live with it. :-)
Still hoping that one day will come real web 2.0 and HTML, CSS, HTTP, JS and even webassembly will be taken as "lessons learnt", thrown away and replaced with something way more logical, structured and extendable
;-)