HNHacker News
TopNewBestAskShowJobs

parrt

667 karma · joined October 25, 2014

Tech lead at Google. Computer languages guy (the ANTLR creator) retooling as machine learning guy, explainer, ex-professor (CS, data science). Hacking almost every day since 1980. Yes, I have tendinitis.
submissionscomments
parrt··on How to visualize decision trees
Heh that’s a cool idea. Fly through the tree like a maze
parrt··on Introduction to Machine Learning for Coders
I sat in on this course last Fall. Excellent.
parrt··on How to visualize decision trees
Yep, the leaves are predictor nodes whereas internal nodes are decision nodes. They are doing different things so we figured we should show them using different visualizations.
parrt··on How to visualize decision trees
Indeed. They were the inspiration for this visualization. I wanted to do something for my book with Jeremy Howard https://mlbook.explained.ai/ and those guys show the way, but of course it isn't a general library. Love that r2d3.us page.
parrt··on How to visualize decision trees
Decision trees are the fundamental building block of gradient boosting machines and Random Forests™, probably the two most popular machine learning models for structured data. Visualizing decision trees is a tremendous aid when learning how these models work and when interpreting models. Unfortunately, current visualization packages are rudimentary and not immediately helpful to the novice. So, we've created a general package called animl for scikit-learn decision tree visualization and model interpretation.
parrt··on Show HN: Owl, a new kind of parser generator
Howdy. Agreed. It'd be nice to have a simpler "match x if NOT followed by y" then and something to handle context-sensitive lexical stuff like Python. I often just send all char to the parser and do scannerless parsing. :)
parrt··on Show HN: Owl, a new kind of parser generator
Well, as the author of ANTLR, I will disagree with your assessment of the tool as you can imagine. It never claimed to be a general tool. Specifically, it cannot handle indirect left-recursion; hence, not a general context free grammar parser. It is however the sweet spot between generality and speed. That quote was having a bit of fun using a tiny bit of hyperbole. If you take a look at our academic paper you will see that the general parsers are like walking a minefield. One never knows when they will hit a landmine and the parser takes overnight to parse a file. It took me 30 years to find the sweet spot in power and speed (with the help of Sam Harwell). I welcome the introduction of new tools, but I'm not sure your assessment is accurate nor is your understanding of the parsing landscape.

Also as the paper says, "GLR return multiple parse trees (forests) for ambiguous grammars because they were designed to handle natural language grammars, which are often intentionally ambiguous. For computer languages, ambiguity is almost always an error." When was the last time you wanted to parse a language where the same syntax meant two different things? C++ and a few other cases, sure, but most languages are specifically designed to be unambiguous.

You are welcome to use a tool that handles all context free grammars, but the speed and ambiguity issues are not something I care to deal with.

You also mischaracterize ANTLR's handling of left recursion. It handles immediate left recursion through a rewrite automatically such as for expressions. It does not handle indirect left recursion. You may have not seen ANTLR 4, and are basing your assessment on ANTLR 3? There is no aborting at compile time for left recursion, except in the unusual case of indirect left record.

http://www.antlr.org/papers/allstar-techreport.pdf

parrt··on How to explain gradient boosting
The key insight seems to be that chasing residuals (for MSE) or sign vectors (for MAE) is chasing a vector (ie direction not just magnitude) and that vector is also a gradient. So chasing residual is performing gradient descent.
parrt··on How to explain gradient boosting
I’m not sure about the connection to category theory. This is mostly an attempt to explain why this model works, that it is performing gradient descent in a particular space. We find that extremely challenging to explain to students. I would be interested to know if you feel the article helps in that regard. Thanks
parrt··on How to explain gradient boosting
The main point of this article is really to explain how gradient boosting works and why. The math is really there to show what the algorithm looks like in its general form. The Discussion of parameters was really just a bit of motivation. Think of this as a good explanation of why it is performing gradient descent in function space. That tends to be very hard to explain.
parrt··on How to explain gradient boosting
True, people use a grid search, but I am always very uncomfortable using things as black boxes. How does tree depth affect generality etc...? Effectively using a model means understanding your tools, in my view, but easy to get started w/o the math as you say!
parrt··on How to explain gradient boosting
Gradient boosting machines (GBMs) are currently very popular and so it's a good idea for machine learning practitioners to understand how GBMs work. The problem is that understanding all of the mathematical machinery is tricky and, unfortunately, these details are needed to tune the hyper-parameters. (Tuning the hyper-parameters is required to get a decent GBM model unlike, say, Random Forests.) Our goal in this article is to explain the intuition behind gradient boosting, provide visualizations for model construction, explain the mathematics as simply as possible, and answer thorny questions such as why GBM is performing “gradient descent in function space.” We've split the discussion into three morsels and a FAQ for easier digestion. Written by Terence Parr and Jeremy Howard.
parrt··on How to explain gradient boosting
Gradient boosting machines (GBMs) are currently very popular and so it's a good idea for machine learning practitioners to understand how GBMs work. The problem is that understanding all of the mathematical machinery is tricky and, unfortunately, these details are needed to tune the hyper-parameters. (Tuning the hyper-parameters is required to get a decent GBM model unlike, say, Random Forests.) Our goal in this article is to explain the intuition behind gradient boosting, provide visualizations for model construction, explain the mathematics as simply as possible, and answer thorny questions such as why GBM is performing “gradient descent in function space.” We've split the discussion into three morsels and a FAQ for easier digestion. Written by Terence Parr and Jeremy Howard.
parrt··on Beware Default Random Forest Importances
Time to revisit any business decision you've ever made based upon default Random Forest feature importances in scikit (Python) or R! Zoiks! :)
parrt··on Matrix Calculus for Deep Learning
Done. Added a link.
parrt··on Matrix Calculus for Deep Learning
Added Link to Wolfram Alpha...
parrt··on Matrix Calculus for Deep Learning
Whoops. thanks. translator error. I'll fix it.
parrt··on Matrix Calculus for Deep Learning
We have to also adjust the image sizes for the in-line equations. That’s what I need to figure out :)
parrt··on Matrix Calculus for Deep Learning
Oops. yeah. thanks
parrt··on Matrix Calculus for Deep Learning
Wow! Great little calculator. Thanks for pointing us at it.
parrt··on Matrix Calculus for Deep Learning
I was also surprised when I saw that there was no standard notation for Jacobian matrices. We use the numerator notation in the article, but point out that there are papers that use the denominator notation. I think I remember from engineering school that we used numerator notation so we stuck with that.
parrt··on Matrix Calculus for Deep Learning
Hiya. That's funny because it's exactly what caused us to write this article. Jeremy and I were working on an automatic differentiation tool and couldn't find any description of the appropriate matrix calculus that explained the steps. Everything is just providing the solution without the intervening steps. We decided to write it down so we never have to figure out the notation again. haha
parrt··on Matrix Calculus for Deep Learning
We originally had that generic ML target in mind but figured a DL bent would make it a wee bit more interesting.
parrt··on Matrix Calculus for Deep Learning
I agree that the font should be bigger. I need to learn more CSS in order to switch between font sizes per platform. The font of the text is easy but all of the images were generated from latex using a specific font size. I need to scale the in-line equation images as the font size bumps up.
parrt··on Matrix Calculus for Deep Learning
Terence here. Jeremy's role was critical in terms of direction and content for the article. Who better than he to describe the math needs for deep learning. :)
parrt··on Ohm: Parsing Made Easy
The unwritten corollary of course is that "almost nobody writes commercial compilers." :) Almost all of us do, however, write parsers for data, config files, languages etc... all the time. I'd personally used ANTLR of course for all my parsing needs beyond the trivial.
parrt··on Towards a Universal Code Formatter through Machine Learning (2016)
One possible angle for improvement of this technology: use a deep learning net to conjure up a different feature vector than the one I handcrafted from language/grammar expertise. I believe this was your idea. :) Glad to have you on board teaching and doing research!
parrt··on Towards a Universal Code Formatter through Machine Learning (2016)
First step to using CodeBuff would be getting an ANTLR grammar for C++ or at least a fuzzy version. It's still a prototype but should do pretty well.
parrt··on Towards a Universal Code Formatter through Machine Learning (2016)
Yep, that's it. Somebody has ported from Java to C# as well. Next step is really to convert to use a Random Forest classifier. I'm stuck elsewhere at the moment.
parrt··on ANTLR Mega Tutorial
Heh, that's cool. I tried to do a random phrase generator at one point but it's hard!
← PreviousPage 2 of 3Next →