"To be clear, we do NOT have evidence to believe that RL outperforms academic state-of—art and strongest commercial macro placers. The comparisons for the latter were done so poorly that in many cases the commercial tool failed to run due to installation issues." and that's supposedly a screenshot from an internal presentation done by Jeff Dean.
https://regmedia.co.uk/2023/03/26/satrajit_vs_google.pdf
As an outsider, I find it very difficult to judge if Chatterjee was a bad and expensive hire (because he suppressed good results by coworkers) or if he was a very valuable employee (because he tried to prevent publishing false statements).
1. preventing bad things
2. preventing bad in a way that all junior members on the receiving end feel bullied
So judging from the article alone, it's either suppressing good results and 2. above, both of which are not valuable in my book
According to a Google investigator's sworn statement, he admitted that he didn't have evidence to suspect the AlphaChip authors of fraud: "he stated that he suspected that the research being conducted by Goldie and Mirhoseini was fraudulent, but also stated that he did not have evidence to support his suspicion of fraud".
I feel like if someone persistently makes unsupported allegations of fraud, they should not be surprised if they get shown the door.
The comparison against commercial autoplacers might be this one (from That Chip Has Sailed - https://arxiv.org/pdf/2411.10053):
"In May of 2020, we performed a blind internal study[12] comparing our method against the latest version of two leading commercial autoplacers. Our method outperformed both, beating one 13 to 4 (with 3 ties) and the other 15 to 1 (with 4 ties). Unfortunately, standard licensing agreements with commercial vendors prohibit public comparison with their offerings."
[12] - "Our blind study compared RL to human experts and commercial autoplacers on 20 TPU blocks. First, the physical design engineer responsible for placing a given block ranked anonymized placements from each of the competing methods, evaluating purely on final QoR metrics with no knowledge of which method was used to generate each placement. Next, a panel of seven physical design experts reviewed each of the rankings and ties. The comparisons were unblinded only after completing both rounds of evaluation. The result was that the best placement was produced most often by RL, followed by human experts, followed by commercial autoplacers."
Also, you are using an unreviewed document from Google not published in any conference to counter published papers with specific results, primarily the Cheng et al paper. Jeff Dean did like that paper, so he can take it up with the conference and convince them to unpublish it. If he can't, maybe he is wrong.
Perhaps, you are biased toward Google, but why do think we should trust a document that was neither peer-reviewed nor published at a conference?
Legal nitpick - you can get away with alleging pretty much whatever you want in a legal complaint. You can't even be sued for defamation if it turns out later you were lying.
Jeff Dean isn't saying that Cheng et al. should be unpublished; he's saying that they didn't run the method the same way. It is perfectly fine for someone to try changing the method and report what they found. What's not fine is to claim that this means that Google was lying in their study.
Google claimed their new algorithm as a breakthrough. If this were the so, the algorithm would have helped design chips in many different cases. Now, the defense is that it only works for some inputs, and those inputs cannot be shared. This is not a serious defense and looks like a coverup.
He even ran a study internally (with Markov), but, as the AlphaChip authors describe:
In 2022, it was reviewed by an independent committee at Google, which determined that “the claims and conclusions in the draft are not scientifically backed by the experiments” [33] and “as the [AlphaChip] results on their original datasets were independently reproduced, this brought the [Markov et al.] RL results into question” [33]. We provided the committee with one-line scripts that generated significantly better RL results than those reported in Markov et al., outperforming their “stronger” simulated annealing baseline. We still do not know how Markov and his collaborators produced the numbers in their paper. (https://arxiv.org/pdf/2411.10053)
But the best (probably only) way to put downward pressure on that is via internal incentives, controls, and culture. You push hard enough for such percent per cadence with no upper bound and graduate the folks who reliably deliver it without checking if the win was there to begin with? This is scale-invariant: it could be in a pod, a department, a company, a hedge fund that owns much of those companies, a fund of those funds, the federal government.
Sooner or later your leadership is substantially penetrated by the unscrupulous. We see this in academia with the spate of scandals around publications. We see this in finance with, who can even count that high anymore. You see Holmes and SBF in prison but the folks they funded still at the apex of relevance and everyone from that clique? Everyone who didn’t just fall of a turnip truck knows has carried that ideology with them and has better lawyers now.
There’s an old saw that a “fish rots from the head”. We can’t look at every manner of shadiness and constant scandal from the iconic leaders of our STEM industry and say “good for them, they outsmarted the system” and expect any result other than a broad-spectrum attack on any honest, fair, equitable status quo.
We all voted with our feet (and I did my share of that too before I quit in disgust) for a “might makes right” quasi-religious system of ideals, known variously as Objectivism, Effective Altruism, and Capitalism (of which it is no kind). We shouldn’t be surprised that everything is kind of tarnished sticky now.
The answer today? I don’t know. Work for the less bad as opposed to more bad companies, speak out at least anonymously about abuses, listen to the leaders speak in interviews and scrutinize it. I’m open to suggestions.