HNHacker News
TopNewBestAskShowJobs

mrdrozdov

821 karma · joined October 1, 2012

I am not a bot

andrew [at] mrdrozdov.com

submissionscomments
mrdrozdov··on PeerReview4All: Fair and accurate reviewer assignment in peer review
This paper has been online at least since 2018.

It has since been published: https://proceedings.mlr.press/v98/stelmakh19a.html

mrdrozdov··on Students suing elite U.S. colleges seek 'wealth favoritism' information
No.
mrdrozdov··on Ask HN: How would you build a ChatGPT detector?
This might provide some guidance: http://gltr.io/
mrdrozdov··on Scaling Mastodon is impossible
This is such a click-bait title and has nothing to do with scalability.
mrdrozdov··on Google no longer producing high quality search results in significant categories
Some folks should post comparisons between Google and other sites such as:

* https://bing.com

* https://duckduckgo.com

* https://you.com

Maybe use the categories mentioned in the twitter thread such as health, travel, recipes, product reviews, etc.

mrdrozdov··on Google no longer producing high quality search results in significant categories
I don't believe this (yet) and here's why.

The claim is that Google search is producing worse results than in the past. The analysis is mostly anecdotal, and similar claims have been made before in a more concrete way. A prime example is "time to cook onions" giving incorrect results, covered in this slate article: https://slate.com/human-interest/2012/05/how-to-cook-onions-...

What we need is to see is specific queries, the results returned, why they're wrong, and what they should be instead.

mrdrozdov··on A cartel of influential datasets are dominating machine learning research
This paper the article refers to is fantastic! I think it's a work most in ML research should become familiar with. And if you believe in the power of benchmarks and data, then this holds even more true. Investing in diversity in datasets is likely an impactful way to make progress in AI/ML.

Minor typo in this article...

ARTICLE: Among their findings – based on core data from the Facebook-led community project Papers With Code (PWC) – the authors contend that ‘widely-used datasets are introduced by only a handful of elite institutions’, and that this ‘consolidation’ has increased to 80% in recent years.

...but right after, they quote the paper and clearly it is 50% not 80%. See the quote from the paper:

PAPER: ‘[We] find that there is increasing inequality in dataset usage globally, and that more than 50% of all dataset usages in our sample of 43,140 corresponded to datasets introduced by twelve elite, primarily Western, institutions.’

...and the article is leaving out this relevant quote from the paper:

PAPER: Moreover, this concentration on elite institutions as measured through Gini has increased to over 0.80 in recent years (Figure 3 right red). This trend is also observed in Gini concentration on datasets in PWC more generally (Figure 3 right black).

...and in general the article is right that inequality is increasing over time, but Gini is a specific metric to measure inequality, and 0.80 is not the same as 80% inequality.

mrdrozdov··on Ask HN: Why is machine learning easier to learn than basic social skills?
They’re more written than you know. If you put similar effort into improving social skills, you will improve. Start with this book: https://www.goodreads.com/book/show/4865.How_to_Win_Friends_...
mrdrozdov··on Notice of Stolen EVGA GeForce RTX 30-Series Graphics Cards
Challenge accepted --- BTC mined on a gameboy: https://www.laptopmag.com/news/you-can-mine-bitcoin-on-a-nin...

Also, you're assuming whoever stole these cards are rational beings and not characters out of Ocean's Eleven.

mrdrozdov··on Notice of Stolen EVGA GeForce RTX 30-Series Graphics Cards
Ocean's Eleven bitcoin edition...
mrdrozdov··on No More Movies
Adding a few more from HBO and Amazon...

  Greenland
  The Little Things
  The Conjuring
  Killerman
  Ghost in the Shell (live action)
  Redline
  Sputnik
mrdrozdov··on No More Movies
I don’t know. I watch a similar amount of movies each year, and I still enjoy it. If anyone is looking for some more obscure recommendations, can check out The Dreamers, and Stilyagi.

EDIT: I have to add that I reference movies a lot in conversation. Often, I’ll watch a movie then immediately call a family or friend to discuss some finer point. This happens frequently, sometimes for a fairly mundane movie detail.

EDIT2: Now I really want to make a list of movies just from this year, since my number has definitely gone up since COVID. I think I’d easily break 100 in 2021 alone.

EDIT3: Here’s a list from my Netflix history since June 1. Mix of TV and movies. I added Justice League Extended Edition and Replica even though they’re HBO because I watched them recently (within the last week). This isn’t really a representative list of my watching, plus I tend to watch a bunch of similar movies/shows, then switch to a new cluster. This group is particularly action heavy because I was playing a lot in the background recently while doing other work. All of these were fun! Even if I don’t think they are the best ever :))

  Movies 2021 June-July

  Zack Snyder’s Justice League
  Replica
  The Take
  Darc
  American Assassin
  S.W.A.T.
  Sniper Legacy
  The Interpreter
  Redemption
  Extraction
  Spenser Confidential

  TV

  Biohackers
  Shooter
  Quantico
  Sweet Tooth
  Record of Ragnorak
  Bodyguard
  Hollywood
mrdrozdov··on GitHub Copilot as open source code laundering?
I think the core argument has much more to do about plagiarism than learning.

Sure, if I use some code as inspiration for solving a problem at work, that seems fine.

But if I copy verbatim some licensed code then put it in my commercial product, that's the issue.

It's a lot easier to imagine for other applications like generating music. If I trained a music model on publicly available Youtube music videos, then my model generates music identical to Interstellar Love by The Avalanches and I use the "generated" music in my product, that's clearly a use that is against the intent of the law.

mrdrozdov··on TikTok has recently traded at valuations of $105-110B on private markets
It isn't an ultimate metric, but they have 150+ papers listed on google scholar, which is not a small amount. https://scholar.google.com/scholar?hl=en&as_sdt=0%2C33&q=byt...
mrdrozdov··on TikTok has recently traded at valuations of $105-110B on private markets
As others have mentioned, this hacker news headline is deceptive. The valuation is with respect to ByteDance, which has a business suite beyond TikTok. Even the article headline mentions this: "TikTok Owner’s Value Exceeds $100 Billion in Private Markets"

There are other facts about ByteDance such as having nearly 100k employees and a moderately size AI research lab that makes this news less surprising.

mrdrozdov··on The Virus Can Be Stopped, but Only with Harsh Steps, Experts Say
I meant that most of the people who are not self-isolating are not jerks.

To convince them to participate in containment, their more responsible friends, family, and colleagues need to reach out, and more importantly we need to have consistent messaging from the government.

mrdrozdov··on The Virus Can Be Stopped, but Only with Harsh Steps, Experts Say
Look up the press conference by Cuomo.
mrdrozdov··on The Virus Can Be Stopped, but Only with Harsh Steps, Experts Say
I don’t believe most people are jerks. They are just in disbelief that such a dire situation is possible.

Yes, it’s important to ask and with a high level of seriousness. If the internet has taught me anything, it’s that people need to read something 8 or more times before fully internalizing it.

mrdrozdov··on How the CDC’s restrictive testing guidelines hid the coronavirus epidemic
Cheers.
mrdrozdov··on How the CDC’s restrictive testing guidelines hid the coronavirus epidemic
The actual title is: How the CDC’s Restrictive Testing Guidelines Hid the Coronavirus Epidemic

This is very different from guidelines in general. Encouraging people to avoid large crowds, wash their hands and wipe surfaces, and possibly self-quarantine could have made a huge impact earlier on.

Instead, people kept going to bars, concerts, and traveling. A lot could have been done, and testing is only part of the story.

mrdrozdov··on NFL Has a GitHub Page
What is the difference between a tech and non-tech company?

Btw, their top-3 repos respectively have 12k, 573, 432 stars. I'm not sure little is the apt description.

mrdrozdov··on NFL Has a GitHub Page
They've nailed the repo naming (sample of size 1): https://github.com/nfl/react-helmet
mrdrozdov··on MIT 6.S191: Introduction to Deep Learning
Publications at NeurIPS/ICML is a reasonable proxy for contribution to the field. https://medium.com/@chuvpilo/ai-research-rankings-2019-insig...

In 2019, the top 5 institutions were:

  1. Google (USA) — 167.3
  2. Stanford University (USA) — 82.3
  3. MIT (USA) — 69.8
  4. Carnegie Mellon University (USA) — 67.7
  5. UC Berkeley (USA) — 54.0
mrdrozdov··on Ask HN: Plenty of large sites down; Reddit.com, GNU.org, Discord, coincidence?
Looks like wunderground.com is down too. If you're wondering, high chance of thunderstorms this evening in New York, NY.
mrdrozdov··on The Great CEO Within: How to build a category-killing company from the ground up
How can we download this file?
mrdrozdov··on John Oliver is erased from Chinese internet following segment on China
I see your point, and although technically correct, the writer had a lot of words to choose from.

For instance, if a nourishing lake had suddenly dried up, I'd use a different word. Perhaps devastating.

mrdrozdov··on John Oliver is erased from Chinese internet following segment on China
> It’s impressive to see the pace of Chinese censors.

I'm not sure that impressive is the right term here...

mrdrozdov··on TileDB: Storing massive dense and sparse multi-dimensional array data
Could I use this for deep learning research? It’s common for my models to be anywhere from 100mb to a few gb. Although, I’d find it more useful for reading batched training data.
mrdrozdov··on A glut of PhDs who can’t find academic jobs
This article is outdated. From 2015.
mrdrozdov··on An MNIST-like fashion product dataset
IIRC a simple convnet will get higher than 99.0% on MNIST, so maybe it is more difficult than it seems at first glance :)
Page 1 of 12Next →