HNHacker News
TopNewBestAskShowJobs

vrv

156 karma · joined May 8, 2011

[ my public key: https://keybase.io/vrv; my proof: https://keybase.io/vrv/sigs/hCFSb8B-N9xFGRaT95bRF9iLljhA1lRrHaj__WXIfVA ]

Co-Founder at lutra.ai

submissionscomments
vrv··on Tensorflow 0.6.0 Release
Thank you for your work on putting Keras on top!

To answer your questions:

- We don't (yet) have a tensor contraction op -- just a matter of getting some dev time to call the existing Eigen contraction code in an op. Hopefully in the next release!

- More casting between types I think is in this release.

- dynamic RNNs is not yet in this one, but also in our sights.

And with all of that, we still need to work on better performance, memory efficiency. Still lots to do!

vrv··on Tensorflow 0.6.0 Release
FYI, we're building the pip packages today -- we'll send out an announcement to the TensorFlow discussion list and update the website when things are actually done. :)
vrv··on Show HN: TensorFlow for AWS on a Real GPU
Thanks, would be good to have multiple sources of verification that HEAD now supports this natively without issues like unexpected NaNs.

https://github.com/tensorflow/tensorflow/issues/25#issuecomm... is one verification :)

vrv··on TensorFlow: open-source library for machine intelligence
We're looking to support Python 3 -- there are a few changes we are aware of that are required, and we welcome contributions to help!

Tracking here: https://github.com/tensorflow/tensorflow/issues/1

vrv··on TCP incast: What is it? How can it affect Erlang applications?
The degenerate/extreme/unrealistic case of Incast is you have switch buffer capacity to store N segments, you talk to M servers that each return one segment, and M >> N. Although RTO is calculated based on RTT and RTTVAR, in the extreme case you can get clumps (and waves) of retransmissions (depending on the properties of the network) such that even eliminating the minRTO altogether may not solve the problem at some scale. In simulation we experimented with adding an adaptive staggering to the exponential backoff algorithm and found that it helped at high server counts [1], but it was only simulation so I'd take that approach with a grain of salt.

The R2D2 work is pretty neat: different than a lot of other approaches I've seen. I'm excited to see how the FPGA implemention works!

Some comments: 1) I think any significant change to the control algorithm requires careful analysis: the variance in throughput with the multi-client experiment looks interesting, though I don't know whether that is steady state. From the graphs, R2D2 suffers more with larger filesizes whereas TCP actually improves. 2) Real datacenters can have very different traffic patterns that can break some of the assumptions about bandwidth uniformity and latency, though it's harder for academics to tackle that. 3) If you are going down the path of TCP offload, you presumably can avoid the overhead of CPU interrupts/timer programming when reducing the RTO into microseconds :). I'd be interested in seeing how R2D2's algorithms/constants work when you're able to reduce the 3ms timer to microseconds in hardware!

Also, if some kernel programmer wants to fix my once-working patch to support microsecond-granularity TCP retransmissions [2], I personally know a bunch of people who would be happy :)

[1] http://vijay.vasu.org/static/papers/sigcomm147-vasudevan.pdf

[2] https://github.com/vrv/linux-microsecondrto

vrv··on DCTCP: TCP optimized for lower latency in Data Centers
I'll let others chime in, but the DCTCP paper mentions "our application reduces the amount of data each worker sends and employs jitter. Facebook, reportedly, has gone to the extent of developing their own UDP-based congestion control [29]."

I recently talked with someone on the Facebook memcached team about this problem and they mentioned that moving to UDP has been useful, but if I'm not mistaken, I believe they gave up on 100% in-order reliability in return for low-latency.

vrv··on DCTCP: TCP optimized for lower latency in Data Centers
This work is really cool and is attacking an important problem in a lot of datacenter networks. It's also great that they've released the source code too. As someone who's worked on this problem, I encourage those interested in reading up more about this.

Here are a few papers on the problem in chronological order:

1. Original paper briefly talking about the "incast" TCP problem in storage environments: http://portal.acm.org/citation.cfm?id=1049998

2. Our follow up work on that problem from a few years back: Measurement paper ( http://portal.acm.org/citation.cfm?id=1364825 ) and our initial solution ( http://portal.acm.org/citation.cfm?id=1592604 ) with a microsecond retransmission Linux patch here: https://github.com/vrv/linux-microsecondrto

3. Another paper talking about Incast in Datacenter environments, focusing on a different form of the workload: http://portal.acm.org/citation.cfm?id=1592693

4. RAMCloud - a project that briefly talks about the need for low-latency transports: http://www.stanford.edu/~ouster/cgi-bin/papers/ramcloud.pdf

5. ICTCP - paper from last year that tries to solve the Incast problem using receiver advertised window algorithms: http://conferences.sigcomm.org/co-next/2010/CoNEXT_papers/13...

DCTCP tries to go beyond solving the Incast problem and focuses on trying to control buffer occupancy in datacenter environments that contain both long flows and bursty flows.

I think these papers all assume lossy link layers (Ethernet), but there are standards and other technologies (Datacenter Ethernet, Myrinet, Infiniband) that aim for lossless link layers to make the transport problem easier, but come with various other drawbacks today (cost, compatibility, etc.). In the meantime, I hope DCTCP or the microsecond TCP patch prove useful in solving some of these problems.

← PreviousPage 2 of 2