HNHacker News
TopNewBestAskShowJobs

xtacy

5,369 karma · joined March 23, 2010

submissionscomments
xtacy··on Dive into Machine Learning with Jupyter and Scikit-Learn
Like any topic/skill, it can be learnt, but only if you spend significant time and effort by doing projects, exercises, asking questions (stackexchange, etc.). It's very important to pay attention to fundamentals and thinking from scratch rather than mastering a laundry list of tips/tricks, because fundamental ideas can be composed in different ways and adapted to a new situation. The fundamentals here would be probability, statistics, linear algebra, optimisation.
xtacy··on Brainstorm – Deep Learning library, successor to PyBrain
Yes, that is correct. Here's the linked reddit post that confirms it: https://www.reddit.com/r/MachineLearning/comments/2xcyrl/i_a...
xtacy··on Streaming video on 10 Gigabit Ethernet and beyond
In the past, some of my colleagues have used Intel's 82599 NICs for kernel bypass. Their Linux driver is quite good, they have a DPDK platform for developing user-space apps to directly access ring buffers on the NIC, and if you do a quick search, you should be able to find examples online.

Cloudflare wrote a blog post recently about accelerated packet IO and their post mentions the 82599 NIC: https://blog.cloudflare.com/kernel-bypass/.

xtacy··on Factorization Machines
Oops, sorry, I didn't mean to say you used a RBF kernel. I edited my comment above. I meant to say your embedding resembles the polynomial kernel (not RBF, which was just meant as a generalisation of such neat tricks :)).

What I meant to say was that you didn't need to compute the embedding explicitly. Since you embed into a space that has a nice structure, you can compute the dot product of the embedded vectors without having to compute the embedding explicitly.

xtacy··on Factorization Machines
Isn't the key insight -- to replace `w_{i,j}` by turning it into a vector with an inner product structure -- also known as the kernel-trick in machine learning? (here, they seem to use a polynomial kernel: https://en.wikipedia.org/wiki/Polynomial_kernel).

EDIT: <strike>It's</strike> Kernel tricks in general are nice because they also generalise well to an infinite dimensional space (the RBF kernel) and compute the dot-products in that space without actually computing the embedding. RBFs: https://en.wikipedia.org/wiki/Radial_basis_function_kernel

xtacy··on Calculus on Computational Graphs: Backpropagation
It's also known as "automatic differentiation" -- it's quite different from numerical/symbolic differentiation.

More information here:

- https://justindomke.wordpress.com/2009/02/17/automatic-diffe...

- https://wiki.haskell.org/Automatic_Differentiation

The key idea is extending common operators (+, -, product, /, key mathematical functions) that usually operate on _real numbers_ to tuples of real numbers (x, dx) (the quantity and its derivative with respect to some variable) such that the operations preserve the properties of differentiation.

For instance (with abuse of notation):

    - (x1, dx1) + (x2, dx2) = (x1 + x2, dx1 + dx2).
    - (x1, dx1) * (x2, dx2) = (x1 * y1, x1 * dx2 + x2 * dx1).
    - sin((x, dx)) = (sin(x), cos(x)).
Note that the right element of the tuple can be computed precisely from quantities readily available from the inputs to the operator.

It's also extensible to derivatives of scalars that are functions of many variables by a vector (of those variables) (common in machine learning).

It's beautifully implemented in Google's Ceres optimisation package:

https://ceres-solver.googlesource.com/ceres-solver/+/1.8.0/i...

xtacy··on Google Cloud Storage Nearline graduates to general availability
Couldn't you generate data from within their data centres? 60Gb/s should be quite feasible. :)
xtacy··on iTerm2 Shell Integration
I just love iTerm2, but I feel it's slow and unresponsive at times (Garbage Collection?) compared to Apple's Terminal. I still keep using iTerm2 because of its features. If there is any way I can profile these slowness and submit a bug report to have it fixed, please let me know!
xtacy··on The impact of fast networks on graph analytics
Yep, that's right. Looking forward to the CPI numbers!
xtacy··on The impact of fast networks on graph analytics
Thanks, yes, I realised it wasn't really barrier-sync latency after I wrote the comment. :)

I remember that the NSDI paper actually made an Amdahl's-law-like argument (they give it a new name) and did something to the tune of "let's just eliminate time waiting on the network from the total runtime, which makes the network infinitely fast."

Coming back to the post: If it's CPU overhead, shouldn't Java be pretty competitive with C/C++/Rust for common computations? There might be a lot of other things going on that lower might affect how much one can squeeze from the CPU (GC/object sizes, time spent in reflection/serialisation, maybe?).

It would be great to look at (a) the number of instructions that Java and the Rust implementation execute, and (b) the instructions-per-cycle issued (or its inverse, the CPI) in both cases. If it's memory sync that's slowing down Java, then Java's CPI must be (edit) _higher_ than Rust's.

xtacy··on The impact of fast networks on graph analytics
I wonder: When the NSDI authors (or rxin) says a computation is communication bound, what do they mean?

Is it network bandwidth or latency?

I suspect it's latency: If you're bottlenecked on latency, the barrier-synchronised nature of many jobs (due to shuffles) lowers network utilisation to the extent that many of the smart network scheduling algorithms the NSDI paper refers to don't work at all.

If it's latency, it also makes sense that a framework that's closer to bare-metal (a highly tuned implementation) can get squeeze more utilisation on a cluster, lowering end to end job times. I wonder if the JVM intrinsically prevents some hardware-specific optimisations due to its memory model.

xtacy··on Inceptionism: Going Deeper into Neural Networks
You could try Torch libraries. There are a few examples on how to (almost) replicate some of Google's neural network models on Imagenet.

Check https://github.com/torch/torch7/wiki/Cheatsheet#demos.

xtacy··on Solving a Crackme using Z3: Theorem Prover
Thanks again. I didn't know Z3 could handle formulae with real numbers. I will take a closer look at it. :)

On the network firewall rules (at multi-tenant Azure, I presume), what were Z3's runtimes look like?

xtacy··on Solving a Crackme using Z3: Theorem Prover
Thanks for the detailed explanation. It's interesting to note that firewall rules are well captured in the predicate logic.

Just curious: Have you encountered rules that cannot be cast into predicate logic framework in Z3?

xtacy··on Solving a Crackme using Z3: Theorem Prover
Interesting! Would love to hear more. What are the classes of firewall properties you can express in Z3?
xtacy··on No Big Bang? Quantum equation predicts universe has no beginning
Discussion on reddit: https://www.reddit.com/r/science/comments/2vb2fa/no_big_bang...
xtacy··on Project Tungsten: Bringing Spark Closer to Bare Metal
Many congrats to Matei! Well deserved.
xtacy··on Random Points on a Sphere
Something doesn't add up in your observation. (x,y,z) can't all be uniformly distributed in [-1,1] because of the constraint x^2+y^2+z^2=1.
xtacy··on The Cost of Scalability in Graph Processing
While I understand the sentiment behind this post, I think it misses one crucial point: It costs time, effort, and very smart people to build the "Bugati"-like system as they describe, instead of the current systems (that are more like "Toyotas", to name one).

I haven't seen the paper yet, so I can't be sure, but I think the numbers might ignore many factors: First, you need some kind of abstract, exchangeable storage (e.g., protobufs) to work with the data in many languages. Third, there's the file-system and all its intricacies. Fourth, it's unlikely that any compute environment will be dedicated only to one application (there's scheduling, resource management, and all that, which means there are hidden costs to doing network IO due to contention, protocol quirks, etc.). And finally, any realistic application is more than just "solving" the problem in the fastest way possible. Requirements change all the time, new features will be added, the code needs to be readable, understandable, maintainable, etc.

It's possible to do all the above AND be super efficient, but it requires a tremendous level of understanding of a system at all levels that it can be quite challenging, and frankly, with business requirements, it's probably not worth the time. If there's a framework that gives you abstraction but compiles to the fastest possible specific implementation AND makes a programmer productive, I would love to read up more!

xtacy··on Visualizing Representations: Deep Learning and Human Beings
Interesting. IIUC, what you're implying is that defining a metric defines the topology and they're equivalent.

Isn't p_ij in t-SNE also derived from the distances themselves, where p_ij ~ student_t(d_ij, degrees_of_freedom) (I forget how the d.o.f. is actually computed in t-SNE.)

Which leads me to one way this distance based approach might be limited: It models similarities using distances, which are symmetric. If similarities aren't symmetric, then this visualisation could hide some information. For example: The specific entity "BMW car" is more similar to the more general entity "car" than the entity "car" is to "BMW car." It seems this asymmetry could capture things (such as the generality of concepts), not reflected in metric spaces (on first thought).

xtacy··on Visualizing Representations: Deep Learning and Human Beings
Ensemble (and also boosted) models: Very nice idea.

I like the takeaway that meta-SNE idea is powerful to compare the space of models by through the lens of pairwise distances as a proxy for the distance metric. Are distances the defining property for a vector space R^d? Could you have used some other quantity instead of pairwise distances?

xtacy··on Visualizing Representations: Deep Learning and Human Beings
Colah, your posts are really inspiring and thoughtful. Your thoughts about visualising the space of representations by looking at the properties of pairwise distance matrix is quite illuminating. It might be a nice empirical way to get a glimpse of the model complexity: If "simpler" models cluster close to more complex models, the simpler models are more desirable.

I wonder if all over-fitted models cluster in one region in the meta-SNE space, or do they show up as noise?

Keep up the great posts!

xtacy··on Cause And Effect: A New Statistical Test That Can Tease Them Apart
Nice article. I think the fact that testing if "X-caused-Y", by exploiting the fact that this is not symmetrical, has also been used by the "pseudo-causality" Granger causality test: http://en.wikipedia.org/wiki/Granger_causality

Also, causality in reality can be quite complicated if there are feedback loops: X-causes-Y-causes-X.

xtacy··on Implementing a bignum calculator – Rob Pike [video]
Could you explain why that snippet makes you decide that way?
xtacy··on Sqlpp11 – A type safe SQL template library for C++
A similar library for Scala: http://squeryl.org/
xtacy··on Why are deep neural networks hard to train?
Thank you colah. Your blog posts are inspiring! It's a lot of hard work and effort; keep it up!
xtacy··on Why are deep neural networks hard to train?
Great, thanks for the pointers! I've tried momentum trick before and it has helped. I'll try rmsprop.
xtacy··on Why are deep neural networks hard to train?
Could you post a few pointers about the bunch of tricks to make deep training a lot easier?
xtacy··on Regulex – JavaScript Regular Expression Visualizer
Pretty impressive. I took a complicated regex that accepts valid email addresses and the output result looks quite neat:

http://jex.im/regulex/#!embed=true&re=(%3F%3A%5Ba-z0-9!%23%2...)

xtacy··on Solving the Mystery of Link Imbalance: A Metastable Failure State at Scale
I am not sure that would work either. If I understand correctly, the root cause of delayed responses is because more requests were sent on that link due to hash collisions, which delays connections on that link. So, wouldn't "most recent request sent" correlate well with "most recent response received"?

EDIT: Also, this seems like a classic load balancing problem: Simply picking the least loaded connection would have been sufficient. The response time on each connection could be computed either explicitly (a running average/standard deviation of RPC finish times) or implicitly by checking the queue backlog at any time (the queue backlong being a first-order statistic).

← PreviousPage 4 of 18Next →