HNHacker News
TopNewBestAskShowJobs

cf

1,357 karma · joined March 28, 2007

submissionscomments
cf··on Build a fast deep learning machine for under $1K
Can you share your PCPartPicker list? I'm curious how the costs breakdown.
cf··on Lyft donates $1M to the ACLU, condemns Trump’s immigration actions
A lot of people from Iran and other countries were also trying to get home that day. They couldn't call an Uber.
cf··on Learning Machine Learning: A beginner's journey
Although to do serious production level ML, I agree that you need to understand the math. But as a starting point, the machine learning for hackers is a great place to start.

I think writing some algorithms and using them to solve problems provides great motivation for the math. In particular, the math will explain why certain approaches did and did not work. Without the hacking that material can get a bit dry.

cf··on Getting Rid of Comments on Vice.com
Since the comments section of a news website is a major source of engagement, why not charge to make comments? Wouldn't this discourage the worst of the them?
cf··on Applying machine learning to the freight industry
If the problem is painful enough, it doesn't matter if the function is trivial.
cf··on Hakaru – Probabilistic Programming
I expect no significant challenges in porting to more recent version of GHC. Some version bounds will need to be loosened and an Applicative instance added for the Measure monads. Also if you have any feature requests for that version I'd be curious what they were.
cf··on Hakaru – Probabilistic Programming
That's the concrete syntax for the language. It was chosen to be make the language more familiar to people who do machine learning in Python. The embedded design we had made it very challenging to develop new inference algorithms and to combine them. You can still find that version of hakaru at https://github.com/zaxtax/hakaru-old
cf··on Hakaru – Probabilistic Programming
I think PyMC is a much more mature solution. Also while PyMC is much more focused on sampling, Hakaru is more focused on Bayesian inference and trying to represent stochastic models in a way such that any inference algorithm could be applied to it.
cf··on Hakaru – Probabilistic Programming
It doesn't require a Maple license exactly. Maple is just needed to use the simplifier. It is perfectly possible to write and run Hakaru programs without having Maple installed.
cf··on Design and Implementation of Probabilistic Programming Languages
There are two lines of research in this direction. So there is work from my lab [1] and some folks at ETH-Zurich [2] is automatically finding closed form solutions. These when used do give performance equivalent to handwritten methods.

Beyond that we are seeing some cool stuff with Blackbox Variational Inference [3] and other Automatic Variational [4] solvers coming out of Blei's lab. At present these capabilities are spread about the different system but I expect all them to eventually end up available in whichever you choose to use.

[1] http://homes.soic.indiana.edu/ccshan/rational/simplify-padl....

[2] http://www.srl.inf.ethz.ch/papers/psi-solver.pdf

[3] http://www.cs.columbia.edu/~blei/papers/RanganathGerrishBlei...

[4] http://arxiv.org/pdf/1603.00788v1.pdf

cf··on Design and Implementation of Probabilistic Programming Languages
So there has been work done in compiling probabilistic programs directly into C or C++ [1]. Your intuition is correct that in many of these languages if you wrote an HMM and did MAP inference over it you wouldn't recover the Viterbi algorithm, but there is some code in WebPPL[2] that handles this case as outlined in https://arxiv.org/abs/1206.3555.

[1] https://web.stanford.edu/~ngoodman/papers/aistats2014-shred....

[2] http://docs.webppl.org/en/master/inference/methods.html#enum...

cf··on Design and Implementation of Probabilistic Programming Languages
Sure. I will also continue to answer any questions others have about these systems on this thread.
cf··on Design and Implementation of Probabilistic Programming Languages
It depends on what you are looking for in a probabilistic programming system. If you want something that is more of a library than a language you gain the advantages of being integrated into a mainstream language. This would point to systems like Figaro, PyMC, Edward and Anglican. If you don't mind a standalone language you can choose systems like WebPPL, Hakaru, Stan, or Venture. There are tradeoffs in expressivity as well. There are probabilistic models you can express in WebPPL or Anglican that you can't in Stan. Also different systems support different inference algorithms. So if you want to do something like Latent Dirichlet Allocation, I think JAGS still does better than Stan using a naive implementation. At the current time, to use these systems productively, you should have some idea of what model you want to write and given your data what inference methods you expect to work with that model.

I think Figaro, Stan and PyMC are the most "production-ready" in the sense they have been used for projects outside the realms of their creators. Still I would argue on some level all of them are research projects that aim to explore how to make probabilistic modeling more accessible to people. Ideas in one language often will appear in another down the line. So I encourage to explore a few of them and reach out to the people working on them.

cf··on So Many Research Scientists, So Few Openings as Professors
I ask since I work on probabilistic programming in Haskell and am always looking for more use cases.

Good luck and I hope the work is published.

cf··on So Many Research Scientists, So Few Openings as Professors
Oh really. Is this Haskell program anywhere online?
cf··on How to win the coding interview
That's why you need to bring some of those thin Japanese whiteboard markers.

http://www.jetpens.com/Zebra-Mackee-Wet-Erase-Double-Sided-M...

cf··on Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values
This is a fantastic point. Sometimes when someone hands me a study I will ask what is the effect size. In some studies, even if there was a discernible effect, there is no hope for it to be anything but a small effect.
cf··on Black Hole Tech?
It's really odd for him to keep harping on Mathematica when the LIGO computations were clearly done in a Python ecosystem as shown in:

http://journals.aps.org/prl/pdf/10.1103/PhysRevLett.116.0611...

https://dcc.ligo.org/public/0122/P1500217/014/LIGO-P1500217_...

https://software.intel.com/en-us/blogs/2016/02/14/python-bri...

cf··on None Programming Language
It's worth mentioning that Terra (http://terralang.org/) is pretty cool on its own terms.
cf··on Show HN: Karplus-Strong Guitar Synthesizer in JavaScript
I'm always fascinated by these simulated sounds. Is there any resource for simulating other instruments like trumpets and drums?
cf··on Los Angeles Plays Itself
I think the landmarks are just less well-known. Things like the Getty and Griffith Observatory. Others are generic, but instantly recognizable, like Venice boardwalk and the LA river.
cf··on Implementing a programming language in C, part 1
I think its helpful to define concepts like interpretation and compilation. Interpreters take a program and produce a value. Compilation takes a program and returns another program that hopefully when interpreted gives the same value.

In some sense, python and ruby compile into bytecode, and then the bytecode is executed using an interpreter. There are also concepts such as partial evaluation (where you do a little interpreting while you compile) and JITing (where you do a little compiling while you interpret) which also seem to mix compilation and interpretation, but can be clearly specified as a particular combination of them.

cf··on The State of Probabilistic Programming
I think we are broadly in agreement. When I say coding up inference for a particular model, that includes deriving the equations and updates needed. This effort is nonzero for all inference methods, but is much lower for MCMC.
cf··on The State of Probabilistic Programming
I know STAN and Figaro are going to push out inference methods like I mention, but my hope is eventually all of the ones mentioned in the article do this. I like thinking of this in terms of the standard library that needs to be built out. All the systems are making great progress in this regard.
cf··on The State of Probabilistic Programming
What's interesting about most complaints of these systems is people talk about their poor performance or scalability? That is usually more a consequence of using MCMC or other inference algorithm than the language itself.

MCMC is a very slow inference algorithm. Its primary advantage was that for well-known models it could be coded up much more simply than a fancier inference technique. When you consider variational methods and newer streaming methods based on things like Assumed Density Filtering you can get really great scalable performance. The point of probabilistic programming is write inference algorithms once for a large class of models and be done. So the advantage of using a fancier method is amplified.

This means paradoxically probabilistic programming should eventually be faster than existing methods rather than slower, since you can reuse these fancier inference methods for new models. This is a very active field so this progress is only starting to be appear in the existing systems.

cf··on SpaCy: Industrial-strength NLP with Python and Cython
I'm curious how this parser compares to ClearNLP http://www.clearnlp.com/ which is similarly a shift-reduce parser.
cf··on Learning languages is a workout for brains, both young and old
The advice I was given was to read a favorite children's book like Harry Potter as you will sort of know the gist of what the passage is suppose to mean and can use that to bootstrap the words in the language you don't know.
cf··on Shell-conduit: Write Shell Scripts in Haskell with Conduit
Are there any summaries of those tradeoffs.
cf··on Shell-conduit: Write Shell Scripts in Haskell with Conduit
So is conduit the recommended streaming IO library now? I saw there were a bunch of implementations of iteratees and was waiting for the community to coalesce around one.
cf··on Hakaru: An embedded probabilistic programming language in Haskell
So I don't really consider myself a Haskell programmer, but the standard advice is to think about operation you are going to perform on those strings and numbers later, then think if there is a datatype you could create instead.

As an example, suppose I am trying to load a csv file with two fields being strings and one being Double. A list of lists representation isn't going to cut it. Instead I should make a Row datatype

  data Row = Row {field1 :: String, field2 :: String, field3 :: Double}
Then you can parse the csv file into a [Row] representation.

As an aside, if the task did involve parsing csv I suggest the cassava library which I found out about through the amazing What I Wish I Knew When Learning Haskell (http://dev.stephendiehl.com/hask/)

← PreviousPage 6 of 8Next →