HNHacker News
TopNewBestAskShowJobs

mrintellectual

956 karma · joined October 20, 2014

Director of Operations & ML Architect at Zilliz (https://zilliz.com). I love chess, deadlifting, and Elden Ring.

Website: https://frankzliu.com

Github: https://github.com/fzliu

Twitter: @frankzliu

submissionscomments
mrintellectual··on Content-Based Image Retrieval
Good catch on the disclosure, I edited my original comment to reflect this fact.

On the topic of vector search, Milvus is another great vector database - it's open source and we provide single-line startup scripts via `docker-compose` in addition to installation via apt & yum (https://milvus.io/docs/install_standalone-docker.md). There are also no restrictions on the number of vectors that users can store. Internally, we've successfully scaled Milvus to handle billion+ vectors, while many of our users have stored hundreds of millions of vectors in a production environments as well.

mrintellectual··on Content-Based Image Retrieval
Great article. We used something very similar to help implement simlarity search at Yahoo a couple years back (https://yahooresearch.tumblr.com/post/158115871236/introduci...). We were using a indexing strategy called Locally Optimized Product Quantization, which worked great in terms of query times but required a training procedure which made successive inserts fairly inefficient.

Thankfully, we have a much wider variety of indexing options these days (https://milvus.io/docs/index.md) in addition to powerful vector databases (https://zilliz.com/learn/what-is-vector-database). I'm glad to see the barrier to entry for semantic image retrieval becoming lower and lower as ML infrastructure matures.

[EDIT] Disclosure: I work at Zilliz.

mrintellectual··on Open Source Alternative To
The list of projects seems to be at least partially based on https://github.com/RunaCapital/awesome-oss-alternatives.
mrintellectual··on The Worst CPUs Ever Made (2021)
My vote for worst CPU goes to the iAPX 432 (also not on this list).
mrintellectual··on Luna Cryptocurrency Collapse: How UST Broke
The crux of the problem is not Luna itself, but the Anchor Protocol (19.5% yield), which enticed users to burn Luna by investing in UST through a, for lack of a better phrase, Ponzi scheme. I get it - people like decentralization. But without some form of regulation, either by the broader crypto community itself or by governments around the world, situations like the LUNA/UST collapse will keep happening.

There are also rumors of a Terra fork: https://agora.terra.money/t/terra-ecosystem-revival-plan/870.... Thanks, but no thanks.

mrintellectual··on twitter/the-algorithm
It just seems unlikely that the algorithm would be open-sourced right after a deal for Twitter is agreed upon (but before it actually goes through). I've never seen a buyout of this scale done by an individual, but I imagine the SEC and several other parties will need to be involved.

At the minimum, I would make a private Github repo first, add all relevant commits, and then make it public once there's actually content.

mrintellectual··on twitter/the-algorithm
This seems to be a practical joke by a Twitter engineer as opposed to an actual release.
mrintellectual··on VeriGPU: GPU in Verilog loosely based on RISC-V ISA
Even as an ML-focused graphics-less GPU, this is great. If this can be prototyped on an FPGA, it would be even better. Using block RAM for shared memory and built-in PCIe and DDR IP blocks should help speed things up considerably.

It unfortunately wouldn't be very cost-effective for training ML models, but it would take things a step closer to actual tape-out (if some organization has the $$$ for it).

mrintellectual··on Six companies control 90% of what you read, watch, and hear
> Americans spend an average of 12 and a half hours per day consuming news via the television, Internet, newspapers, magazines, and radio.

This just doesn't make sense. The average American spends the majority of waking hours consuming news?

mrintellectual··on K-Nearest Neighbors
I haven't tried SimCLR, but I did try face embedding models trained with contrastive and triplet loss. For applications where precision is the key metric, I do agree that these loss functions are much better overall.

If discovery or recall is what you're after, a generic image classification model trained with binary cross-entropy might be better. For example, performing reverse image search on a photo of a German Shepherd should always return images of GSheps in the first N pages, but showing other dog breeds in later pages and possibly even cats after that would be a desirable feature for many search/retrieval solutions. An embedding model trained with contrastive loss might have this behavior to a certain extent, but a model based on BCE should be better.

mrintellectual··on K-Nearest Neighbors
> kNN combined with embeddings from pre-trained deep learning models can be very useful for information retrieval

Indeed! We've been able to build simple reverse image search apps and other solutions using the power of embeddings from pre-trained ML models: https://gist.github.com/fzliu/c9380a7f9ba411adeff0b727cdba15....

One quick note: k-d trees are great for indexing low-dimensional data, but for high-dimensional embeddings they tend to be a poor indexing choice since you'll end up visiting more nodes in the tree than you'd like. I found [1] to be a great overview of different indexing types for high-dimensional vectors and the advantages of each.

[1] https://milvus.io/docs/index.md

mrintellectual··on Transformers Are All You Need
> ViT models are outperforming CNNs in terms of computational efficiency and accuracy, achieving highly competitive performance in tasks like image classification, object detection, and semantic image segmentation.

Since then, this has been show to be untrue. Using more modern training techniques along with depthwise convolutions (https://arxiv.org/abs/2201.03545) results in equal if not better performance on vision tasks. Improved training methodologies have also been shown to boost the accuracy of ResNet50 - an 6-year-old pure convolutional architecture - on ImageNet-1k by over 5% (https://arxiv.org/abs/2110.00476).

Pure ViTs are also more difficult to train when compared with traditional convnets, although this has since then been somewhat remedied by Swin (https://arxiv.org/abs/2103.14030).

mrintellectual··on The most popular chess streamer on Twitch
Hikaru's win in the Grand Prix was great, but the Candidates is far different. For example, Alireza Firouzja - another member of the upcoming Candidates cycle - has been MIA for a while, likely due to the insane amount of time he is putting into Candidates prep.

Since Hikaru has far fewer recent classical matches than the other upcoming Candidates participants (with the exception of current world #2 Ding Liren, who still needs to finish 30 games before May), he'll have an edge when it comes to preparation. However, unless he prepares like a madman, he'll be at most a wildcard candidate, perhaps beating a favorite or two but unlikely to win it all.

I'll personally be rooting for Ding and Fabi. I like Hikaru as well, but I unfortunately just don't see him beating Magnus in a World Chess Championship.

mrintellectual··on Gmail sends far more emails from conservatives to spam: study
I would imagine the demographics are different just purely based on the fact that Yahoo Mail has been around for longer than Gmail.
mrintellectual··on Twitter accounts dropping “.eth” from usernames
> That said, Hacker News readers shouldn't be surprised that people use jargon to demonstrate membership of an in-group or keep away outsiders.

To say that HN readers use specific keywords or acronyms is inaccurate, given the diversity of topics that make it to the front page. However, the discussions I see take place here are definitely of much higher overall quality than the average comments thread on other sites.

mrintellectual··on ‘Sovereign chess’ is a battle without loyalty or humanity (2020)
I'm curious what an optimal strategy for playing such a complex game like this is (if it even exists). I also wonder how much of a role intuition plays, i.e. the ability for chess masters to "feel" whether or not positions are winning or losing.
mrintellectual··on Ask HN: Who wants to collaborate? (April 2022)
We're a small group of folks working on an open-source project for generating embedding vectors. If you're interested in machine learning or semantic search, we welcome you to join us!

https://github.com/towhee-io/towhee

mrintellectual··on The entrance to a small personal site
Reminds me of a text-based MMO called Alter Aeon (http://www.alteraeon.com).
mrintellectual··on What it’s like working at TikTok: The Opportunities (Part 2)
Here's part I for those interested: https://medium.com/@melodychu/what-its-really-like-working-a...
mrintellectual··on Verilog Is Weird
A concise but not-completely-accurate way I explain it to a lot of pure software folks is that every line of HDL code executes "simultaneously". This can sometimes help them wrap their heads around Verilog/VHDL a bit better.
mrintellectual··on UCLA hiring Asst. Adjunct Professor “on a without salary basis” (PhD required)
Is this common in academia? Not a snarky question, just genuinely curious.
mrintellectual··on Apple’s charts set the M1 Ultra up for an RTX 3090 fight it could never win
The tech is interesting, but I think Apple's claim that the M1 Ultra "beats" the RTX 3090 is just absurd. Regardless, I'd also like to see some benchmarks centered around training ML models. It might get some bonus points there just due to its ability to support larger batch sizes.
mrintellectual··on Can making employers share pay in job postings help fix the gender pay gap?
It seriously irritates me when these articles like these conveniently ignore a certain portion of the population. Where's non-Latina white women in the bar chart?

I did some rough-and-dirty calculations based on US census data (https://www.census.gov/quickfacts/fact/table/US/RHI725219) and the 83-cent statistic in the article. Turns out non-Latina white women actually earn _more_: about $1.03 for every $1.00 earned by non-Latino white men.

A lack of transparency will do more to hurt gender and racial equality in the long run. It's better to leave the statistic in and explain why the numbers are as they are instead of turning a blind eye.

mrintellectual··on I cannot write software outside of Google anymore
The original Lua-based Torch is actually from EPFL, not Facebook. Tensorflow originated from Google's internal DistBelief system, making TF a Google project through and through.
mrintellectual··on I cannot write software outside of Google anymore
For a short period of time circa 2017/2018, Google was on the path toward monopolizing machine learning as well through Tensorflow. Thankfully, PyTorch overtook TF.
mrintellectual··on Some discouraging anecdotes on how services handle account deletions
Datacenter carbon neutrality and de-biasing ML models are two that I can recall off the top of my head.

These are, of course, unrelated to account deletion, but it shows that big tech is at the very minimum aware that social responsibility is becoming a more important part of business.

mrintellectual··on Some discouraging anecdotes on how services handle account deletions
Part of this is leftover tech culture from Facebook's early focus on Growth, Growth, and more Growth. Allowing for easy deletion of accounts was fundamentally at odds with user growth.

It's refreshing to see that tech is now heading in a more socially responsible direction, but the industry still has a long way to go.

mrintellectual··on Tell HN: China Is Entering Lockdown
I can sympathize with your situation. I was in an office building in Shanghai last November when it got locked down for 48 hours - nobody in or out. Everybody ended up ordering takeout, and the entire building had to be tested for COVID a total of four times. The most insane part is that they don't even let you take pictures of the lockdown (I almost got my phone confiscated by the police for taking pictures).

Also, sleeping in an office chair is a lot harder than it looks.

mrintellectual··on Ask HN: What ML platform are you using?
I like Colab for the most part since I'm biased towards Python, but being centered around Jupyter notebooks does have its shortcomings. Also, despite being a service offered by Google, I prefer PyTorch over Tensorflow.

For smaller projects, I generally find a Towhee pipeline (https://towhee.io/pipelines) that I then fine-tune on my 3080.

mrintellectual··on My Experience Working and Living in China
Glad you enjoyed the article. Part II should be coming next week - stay tuned.
Page 1 of 3Next →