HNHacker News
TopNewBestAskShowJobs

vishvananda

979 karma · joined January 10, 2011

[ my public key: https://keybase.io/vish; my proof: https://keybase.io/vish/sigs/j9A4-vaQeWa407vRTQ75jGiDQd-Mcs2kxWzvMMdxMgM ]
submissionscomments
vishvananda··on Fast implementation of DeepMind's AlphaZero algorithm in Julia
Unfortunately I don't remember the exact numbers, but I think it was a couple percentage points worse than we were able to get with the large models.
vishvananda··on Fast implementation of DeepMind's AlphaZero algorithm in Julia
we did a lot of our early experimentation with small networks. I don't think we went any smaller than 5 layers of 64 filters as we mentioned here: https://medium.com/oracledevs/lessons-from-alpha-zero-part-5...
vishvananda··on Fast implementation of DeepMind's AlphaZero algorithm in Julia
The lc0 group has switched the result prediction to predict win, loss, and draw probabilities instead of just win/loss. Some information can be found in https://lczero.org/blog/2020/04/wdl-head/
vishvananda··on Fast implementation of DeepMind's AlphaZero algorithm in Julia
Nice work on this! I was behind the implementation at oracle which you referenced in the tutorial. I still keep tabs on the lc0 crowd which seems to be pushing into new ideas. Did you pull anything else from the leela crowd besides prior-temperature? It looks like maybe you also tried a WLD output head as well?
vishvananda··on Fast implementation of DeepMind's AlphaZero algorithm in Julia
My team did an implementation of alpha zero connect four a couple of years ago. Our findings are in a series of blog posts starting at https://medium.com/oracledevs/lessons-from-implementing-alph.... We didn't manage to get to perfection either on policy, but got pretty close. You can play against some versions of the network here: https://azfour.com
vishvananda··on Simple techniques to optimise Go programs
also the overhead of calling out to c can actually be quite high: https://www.cockroachlabs.com/blog/the-cost-and-complexity-o...
vishvananda··on The new Dropbox
the parent mentions in parentheses that the dir doesn't have to be in dropbox because it appears to monitor all files.
vishvananda··on My Childhood in a Cult
Prior to my life in technology I had similar experiences to the author. I was involved with multiple spiritual groups that could be classified as cults. One mistake that I have seen people without first-hand experience make is assuming that the people are in it to deceive or take advantage of others.

On the contrary, I have seen that people in these groups often have the best intentions. For a long time, it was a mystery to me how these groups could end up going so wrong, often devolving into gun battles, suicides, or sexual deviance. A few years ago I found a book called "The Guru Papers"[0] which does a fantastic job of explaining how these things occur, even in well-intentioned groups. If you have been involved in a group like this, or are just curious about the psychology involved, I highly recommend it.

[0]: https://www.amazon.com/dp/B007WL0JHE/ref=dp-kindle-redirect?...

vishvananda··on Gravitational Wormhole: WireGuard for Kubernetes
Creator of https://github.com/vishvananda/netlink here. Would be happy to have support added.
vishvananda··on Tensorflow 2.0.0-Alpha0
I think this is primarily due to the immense effort that NVIDIA has put into CUDA. It works very well and it is extremely fast. The alternatives for AMD are OpenCL and ROCm which have seriously lagged behind CUDA in every respect.

EDIT: lots of theories and discussion here: https://www.reddit.com/r/MachineLearning/comments/7tf4da/d_w...

vishvananda··on Debian Buster will only be 54% reproducible, while we could be at 90%
This is actually an annoying challenge of reproducible builds. In many cases it is actually useful to have a build timestamp, git sha, or build number available for debug output from the program. I've often gone as far as embedding a sha and/or timestamp into a file on export into a tgz which allows it to be reproducible from the tarfile, although builds directly out of source control would not be.
vishvananda··on Debian Buster will only be 54% reproducible, while we could be at 90%
Unfortunately, nix does not produce fully reproducible builds. The build environment is portable and produced in a way that it can be repeated, but due to the limitations of the software that is being built, the builds are not binary reproducible. You can see some commentary on the nix team hoping to adopt some of the work being done by debian et al here: https://github.com/NixOS/nixpkgs/issues/9731
vishvananda··on On higher education, programmers and blue-collar jobs
I also found the tone pretty condescending, but I wasn't sure if that was a result of being translated from another language. I agree that both paths are viable.
vishvananda··on On higher education, programmers and blue-collar jobs
Isn't this mistaking the path for the goal?

I think it is more accurate to say that the knowledge of how OSes work, basic discrete math, etc. is important regardless of how that knowledge is gained.

A university education is only one way of achieving this knowledge. Even your last point of gaining knowledge about what you don't know about CS is achievable through other means (mentorship, online classes, study groups, etc.)

I think it is perfectly fair to suggest that a CS university education is a great way of achieving this knowledge, and it makes perfect sense to recommend that method if you followed that path yourself.

Personally, as someone who developed that knowledge outside the university system, I suspect that university is probably an excellent approach for most people, but I tend to encourage people to find the hunger for knowledge inside themselves and develop a passion for learning in whatever form it comes.

vishvananda··on If Software Is Funded from a Public Source, Its Code Should Be Open Source
Contractors have their own hassles when it comes to open source. Pre-openstack, there was a small contracting company called Anso Labs that consisted of a few of us that were creating a private cloud for NASA. After a few months of banging our head against Eucalyptus (java based open-core private cloud software), we decided to write our own version over a weekend and call it Nova. We convinced the civil-servants in charge to allow us to use our new thing, which we continued to improve at NASA. This project eventually became the compute platform for openstack.

There were multiple weeks of meetings with lawyers (including discussions of whether we needed a space act agreement[1]) to figure out how to actually open source the modifications we made on NASA's behalf. Ultimately we ended up having to assign all of the nova copyrights to the government[2] so that they could open source the modifications. I didn't even think that the government could own copyright but apparently it can[3].

[1]: https://www.nasa.gov/partnerships/about.html [2]: https://github.com/openstack/nova/commit/c88d1f033bd600e855c... [3]: https://www.usa.gov/government-works/

vishvananda··on A Study on Driverless-Car Ethics
Is it just me or does solving the moral dilemma of prioritization based on attributes of the person seem like a worthless endeavor? In practice, I can't come up with a case where it would matter.

For example, a more reasonable metric for a machine to use is the probability of injury/death. If swerving is 90% likely to kill a person and staying straight is only 89% likely, then staying straight is a better choice. I don't see how attributes of the person would ever trump the probability of harm. The cases where probability is roughly equal for multiple actions will be incredibly rare.

vishvananda··on The Firecracker virtual machine monitor
My evidence is mostly anecdotal, unfortunately. Most of my experimentation was KVM in KVM from about 2012 and i saw frequent long pauses in the kernel, lockups, and kernel panics. At the time, the attitude from the qemu-kvm community was that it wasn't a high priority and that it may or may not work. It is possible that it is better now, but redhat's opinion[1] as of last year was still that it wasn't supported in production. I don't know if it is any better with other hypervisors. Last week, for example, I ran into an issue with nested virt in vmware fusion[2] that prevented me from running kvm in a guest. I would need a bunch of testing before I trusted it.

It looks like xen has spotty support[3] as well. Amazon uses a heavily forked version of xen, so I suspect support for it is even worse on their version.

(EDIT: added info on xen)

[1]: https://www.redhat.com/en/blog/inception-how-usable-are-nest...

[2]: https://bugs.launchpad.net/qemu/+bug/1661386

[3]: https://wiki.xenproject.org/wiki/Nested_Virtualization_in_Xe...

vishvananda··on The Firecracker virtual machine monitor
I believe this is only because ec2 does not allow nested virtualization. In my experience, nested virtualization is still buggy and can suffer from major performance issues. So although it is possible to run it in a vm, I'm not sure I would recommend running production code on it there.
vishvananda··on Go Modules in 2019
I used go in machine learning contexts extensively while writing graphpipe[1]. Go is a fantastic language for servers, and distributed communication. Unfortunately, the lack of generics and dependence on interfaces and reflection makes writing things that deal with multidimensional arrays pretty terrible. See, for example, the janky conversion code in graphpipe-go[2] to convert a multidimensional slice into contiguous row-major arrays. Also, libraries that try to create a numpy eqivalent end up with uncomfortable interfaces due to inability to overload operators. I agree with some of the sibling comments that go 2 will help but probably won't make things particularly pleasurable. Rust would be a more interesting route but definitely doesn't have the adoption of go.

[1] https://oracle.github.io/graphpipe

[2] https://github.com/oracle/graphpipe-go/blob/master/helpers.g...

vishvananda··on MongoDB's Server Side Public License Is Likely Unenforceable
I think you have this correct, as I understand it. Just to explicitly state the differences:

1. You can modify GPL code as much as you want, and as long as you don't distribute the software, you do not need to make the modifications available. If you distribute the software, you are required to make your code available.

2. The AGPL extends the definition of distributing the software to making the software available over the network. This means if you modify AGPL software and then make it available over the network (as SaaS, for example), then you are required to make your code available.

vishvananda··on Show HN: TFServe – Simple and easy HTTP Server for tensorflow models
Yes I totally understand. In general model serving doesn't seem to be something people are thinking about at the moment. I expect it gets more attention in the next couple of years.
vishvananda··on Show HN: TFServe – Simple and easy HTTP Server for tensorflow models
I like how you made this ultra-simple. It makes it very easy to understand and use. There are times when a python+json server isn't performant enough. To help with this, my team recently built a protocol for efficient model serving called GraphPipe[1] that allows you to do serving like this. There is an example for serving from python[1]. The easiest way to get started is to use the go server in a docker container[2].

[0]: https://oracle.github.io/graphpipe/ [1]: https://github.com/oracle/graphpipe-tf-py/blob/master/exampl... [2]: https://oracle.github.io/graphpipe/#/guide/user-guide/quicks...

vishvananda··on Ask HN: Would you rather work for Google or a smaller high growth company?
not op, but I assume it stands for Get Sh*t Done
vishvananda··on Ways to think about machine learning
You might find http://www.fast.ai/ useful. Depending on your learning style, their courses can either be amazing or somewhat annoying. Their library includes jupyter notebooks so that you can work through the examples.
vishvananda··on Ask HN: Is it worth investing in learning Rust?
I've written fairly complex projects in both languages, but have a great deal more experience in go. I think tptacek nails it with this advice: go is much easier to learn. One of the things i really like about it is how easy go code is to read coming from almost any language.

It makes sense to get comfortable with go first because the time commitment to achieve basic competence with rust is much greater. Rust will definitely open your mind but it will take some time to get there.

vishvananda··on Ask HN: Which books have made you introspect?
My favorite introspective book is Awareness, By Anthony DeMello[1]

It is a small book but it really makes you think about your views on life.

[1] https://www.amazon.com/Awareness-Opportunities-Reality-Antho...

vishvananda··on Cutting Edge Deep Learning for Coders, Part 2
I can see how The courses would be tough to digest without some theory from other sources, but I have found them to be invaluable for tips and tricks. They are loaded with practical information.

Much of Deep Learning is still experimental in nature and requires quite a bit of educated guessing. A number of times I have been stuck on a particular deep learning problem and a passing comment from one of the fast.ai videos has given me the perfect insight.

vishvananda··on IBM is not doing "cognitive computing" with Watson (2016)
There have been a number of breakthroughs in specific fields recently. In fact it seems like there is a paper every week that pushes the state of the art forward in some branch of AI/ML. I think the big one that triggered the current excitement was the success of convolutional networks in image recognition tasks. You can read about that one (in 2012) here: https://www.technologyreview.com/s/530561/the-revolutionary-...

EDIT: This paper refers to the algorithm as SuperVision which was the team name, but it is more commonly called AlexNet. Here is another article discussing it:

https://qz.com/1034972/the-data-that-changed-the-direction-o...

vishvananda··on Notes on structured concurrency, or: Go statement considered harmful
You might want to read further. He didn't claim that callbacks are not harmful. In fact he suggested that they suffer from the same problems and he is using "go statements" to encompass all of the forms of concurrency handling.
vishvananda··on Talent vs. Luck: the role of randomness in success and failure
From the paper:

2. A lucky event intercepts the position of agent Ak: this means that a lucky event has occurred during the last six month; as a consequence, agent Ak doubles her capital/success with a probability proportional to her talent Tk. It will be Ck(t) = 2Ck(t − 1) only if rand[0, 1] < Tk, i.e. if the agent is smart enough to profit from his/her luck.

3. An unlucky event intercepts the position of agent Ak: this means that an unlucky event has occurred during the last six month; as a consequence, agent Ak halves her capital/success, i.e. Ck(t) = Ck(t − 1)/2.

Note that the equation for lucky events includes talent, but unlucky events do not.

← PreviousPage 2 of 5Next →