HNHacker News
TopNewBestAskShowJobs

etrain

1,737 karma · joined April 11, 2011

Present: Founder at Stealth. Past: VP AI Solutions and AI Cloud at HPE, Founder/CEO at Determined AI (acquired by HPE), Ph.D in AI/Distributed Systems in the UC Berkeley AMPLab, engineer at Recorded Future, and senior quant analyst at MDT Advisers.
submissionscomments
etrain··on Screenshots from developers and Unix people taken in 2002
The Mozilla Suite existed and wasn't that bad, even in 99/00 when I started using Linux. Konqueror was also pretty decent. The web also wasn't as central to the Internet experience then (irc, pop3 clients, and so on were pretty prevalent.) Most of my time in a web browser was spent checking slashdot, googling for things, and so on.
etrain··on Deep learning for NLP resources
Ion Stoica's Big Data Systems research class at Berkeley has a pretty solid reading list to get you started:

http://www.cs.berkeley.edu/~istoica/classes/cs294/15/class.h...

etrain··on Default Alive or Default Dead?
EBITDA.
etrain··on Machine Learning for Developers
Check out our project, [KeystoneML](http://keystone-ml.org/) - it's geared to large scale machine learning in the realms of computer vision, NLP, and (soon) speech. The design is modular and engineering-friendly and is quite focused on real-world applications. (e.g. start with pictures, do the feature extraction, PCA, linear regression, etc.) and out comes a classifier.
etrain··on Amazon QuickSight – Business Intelligence by AWS
This is the heart and soul of quantitative investment management.
etrain··on iPhone 6s and iPhone 6s Plus Preliminary Results
I've heard this referred to as "the footstool" and "the emasculator."
etrain··on Big Data: Astronomical or Genomical?
For anyone else confused about the Twitter numbers reported in Table 1, I believe it should read "0.5-15 billion tweets/day" as opposed to "0.5-15 billion tweets/year" on the first line, which is consistent with the claim that Twitter produces 500m tweets/day currently. When you make this adjustment, you recover their annual storage estimates.

Still - this estimate is based on 3KB/tweet, which is probably derived from looking at the raw twitter XML feed - I'd expect this to compress easily down to 10-30x smaller than the author's claimed numbers - more with proper data modeling.

Nevertheless - huge problems to be dealt with in the sciences and video.

etrain··on China's stock market crash: A red flag
There are a handful of Chinese index ETFs out there - FXI is the biggest one - but there are others you can probably short like the MCHI or the GXC.

If you really like to gamble you can buy puts on some of the bigger ones.

etrain··on MapD: Massive Throughput Database Queries with LLVM on GPUs
Yeah, I agree this is surprising. If all the data in their CPU implementation were stored in a well-packed columnar format, I'd expect something close to theoretical memory throughput on the query they describe (intermediate counts table fits easily into L1/L2 cache and results are commutative and associative). Cache-coherence issues in the aggregation code are a possible cause.

If I were building a speed-freak analytics database I'd be focusing on making my CPU implementations as fast as possible, since that's what 99% of potential customers are already running. Assuming you get to 40GB/s on this type of query, that's only a factor of 5 slower than the GPU implementation. I'd imagine that for most workloads, a factor of 5 speedup that requires new hardware and lots of energy is kind of a non-starter.

etrain··on KeystoneML – Machine Learning Pipelines with Apache Spark
Nope! KeystoneML is a research project exploring how to best support end-to-end pipelines for large scale machine learning. It's complementary to MLlib and some functionality from KeystoneML may find its way into MLlib in the future.
etrain··on KeystoneML – Machine Learning Pipelines with Apache Spark
I'm one of the authors of KeystoneML. Happy to answer any questions about it here.
etrain··on Compressing PostgreSQL JSONB data 12x using cstore_fdw
The chosen benchmark (a customer reviews) table is likely something that benefits tremendously from compressed columnar storage: 1) It has a small number of attributes which are almost always present in the records - that is, a fixed schema. 2) The most of the fields are numeric/date, or text with low cardinality (product_category, etc.) These things respond well to huffman codes and run-length encoding.

This is a use case where JSON shouldn't really ever be used, because the schema is pretty much fixed and highly regular. JSONB records essentially carry the schema definition with them per-record and in this case most of that information is duplicated - hence the blowup in its representation on disk.

While column stores are great for answering analytical queries that require scans over the whole table (like the single query example they show), they aren't as good at transactional queries (like serving webpages).

If I were citus, I'd have written the blog post using a dataset of highly irregular JSON blobs - e.g. log messages from lots of different systems or a big collection of web pages (serialized as json representations of the DOM). Maybe we'll see these "in the coming weeks."

etrain··on Amazon Machine Learning – Make Data-Driven Decisions at Scale
This does not mean they are not using Vowpal Wabbit. It is very easy to run Vowpal Wabbit with a logistic loss function.

Also, vw is what I'd consider "industry standard."

etrain··on Amazon Machine Learning – Make Data-Driven Decisions at Scale
My guess is they're using liblinear or vowpal wabbit under the hood. Both support SGD-based learning and work well in a streaming setting where data could be on disk or in memory.
etrain··on Apple to replace AT&T in Dow Jones on March 18
Agreed, but given their differences, the historical correlation between the S&P and the dow is absurd. https://www.google.com/finance?q=INDEXDJX%3A.DJI&ei=Wef5VPmw...
etrain··on David Horne's 1K Chess on the ZX81 (2001)
oh, with my latest chess programming language, the source code for chess is 0 bytes. feed an empty file to the chess compiler and out comes a working binary which plays chess.
etrain··on Quantifying Online Advertising Fraud: Ad-Click Bots vs Humans [pdf]
The authors are selling a product to detect click bots. Take the results with a grain of salt..
etrain··on Inside a Chinese Bitcoin Mine That's Grossing $1.5M a Month
Based on the video, it looks like these guys likely swap out for the newest equipment as soon as it makes sense economically.

If I were making such an investment, I'd think carefully about the rate of computational depreciation (which is predictable) and probably hedge on the price of power. The only unpredictable thing in this setup is the price of bitcoin, which I believe they're extremely bullish on long term. You could probably figure out a way to hedge that, too, if you could find someone to take the action.

etrain··on A general API for relational joins
Cool read - but you should be careful modeling Tables as things with an "id". The relational algebra is about sets of tuples. The fact that there is an index that one can use to make joins go faster is an (important) implementation detail but divorced from the data model.

The fact that these are sets is an important distinction and enables things like predicate pushdown and other important optimizations. Would be interesting to see a Selinger-style cost-based optimizer built in Haskell!

etrain··on AI Swarms on the Blockchain
See Byzantine Fault Tolerance if you don't want to deal with a centralized ledger: http://en.wikipedia.org/wiki/Byzantine_fault_tolerance
etrain··on Revolution R Open: The Enhanced Distribution of Open Source R
Various linear solvers (either via normal equations, QR, etc.) all have really fast multi-threaded implementations in, e.g. OpenBLAS. These could directly benefit lm() and glm(). That said - there's no reason why you couldn't already call out to these (multithreaded) libraries with (single threaded) R.
etrain··on Revolution R Open: The Enhanced Distribution of Open Source R
Check out SparkR - http://amplab-extras.github.io/SparkR-pkg/
etrain··on Sudoku, Linear Optimization, and the Ten Cent Diet
One thing often missing from these types of solutions (this one included) is that it's extremely easy to generalize them beyond the 9x9 case.
etrain··on SIMD Vectorization in Julia
Looks like a very nice feature - but would be great if they provided some actual benchmarks showing when vectorization is faster than not. For example, if you're bound by memory bandwidth, no amount of extra compute is going to help you.
etrain··on Useful Unix commands for exploring data
I'm reminded of the old joke, "python is executable pseudocode, while perl is executable line noise."

But seriously, I've got some battle scars from the perl days, and hope not to revisit them. Honestly, there's very little I find I can do with perl and not python, and it's just as easy to express (if not quite as concise) and much simpler to maintain.

But, use the tool that works for you!

etrain··on Readings in Databases
http://db.cs.berkeley.edu/cs286/papers/anatomy-redbook2005.p...
etrain··on Readings in Databases
Done: https://github.com/rxin/db-readings/pull/2
etrain··on Readings in Databases
See also: http://www.cs286.net/home/reading-list - Joe Hellerstein's graduate database class.
etrain··on Useful Unix commands for exploring data
Some more tips from someone who does this every day.

1) Be careful with CSV files and UNIX tools - most big CSV files with text fields have some subset of fields that are text quoted and character-escaped. This means that you might have "," in the middle of a string. Anything (like cut or awk) that depends on comma as a delimiter will not handle this situation well.

2) "cut" has shorter, easier to remember syntax than awk for selecting fields from a delimited file.

3) Did you know that you can do a database-style join directly in UNIX with common command line tools? See "join" - assumes your input files are sorted by join key.

4) As others have said - you almost invevitably want to run sort before you run uniq, since uniq only works on adjacent records.

5) sed doesn't get enough love: sed '1d' to delete the first line of a file. Useful for removing those pesky headers that interfere with later steps. Not to mention regex replacing, etc.

6) By the time you're doing most of this, you should probably be using python or R.

etrain··on Why Golfers Buy Hole In One Insurance
I have golfed since I was a kid (grew up in the mid-Atlantic), and have been a serious golfer for the last 7 years or so - having lived both in Boston and the bay area during that time.

I very rarely bet any money on my round, nor do the people I play with. True, there are people betting on their rounds, but I'd estimate it's 1 out of every 10 groups that goes out for a round that does that.

The article blows it out of proportion a little bit, too. At the courses I play (in the bay area), there are usually maybe 20 people in the clubhouse, and a round of drinks would probably come to $100. The "insurance" crowd is a very small subset of the golfing population.

True - it is an expensive sport and maybe traditionally a game for the wealthy, but my weekly golf habit doesn't cost much more than a gym membership. People from all walks of life play and enjoy the game, and it definitely doesn't have to be expensive unless you want it to be.

← PreviousPage 2 of 7Next →