Nvidia Scientists Take Top Spots in 2021 Brain Tumor Segmentation Challenge
developer.nvidia.com
developer.nvidia.com
[1]: https://twitter.com/emollick/status/1388594078837878788/phot...
He likes to say AI=avian intelligence...
Really not fake … ???
https://medium.com/war-is-boring/that-time-the-u-s-coast-gua...
Actual article (2015): https://journals.plos.org/plosone/article?id=10.1371/journal...
Pretty incredible. I wonder if you forced a radiologist to make a decision within a second or two if it would improve their scores...allowing them to rely on the instinct they've developed vs the cognitive wet blanket that factors in all of the externalities of making a decision.
No doubt these are model innovations. But can't help wondering how much ready access to arrays of GPU might give them advantage... You know how many different models need to be tested before getting to the winning ones.
Intelligence to create innovative models, hardware to test/iterate models, etc.
AlphaFold is a clear example of what's possible with big money.
I dunno, their A100 results took about 20-30 minutes on 8 x A100s [1]. 8xA100s is like $24/hr on GCP at on-demand rates [2]. So like $12/run.
The efficiency was okay but not linear, so if you were more cost constrained you might go with 1xA100 for $3/hr and have ~2.5hr training times.
Getting that performance out of a GPU is more challenging than getting access to the GPUs. All the major cloud providers offer them.
(Nit: GCP deployed the 40 GiB cards rather than the later 80 GiB parts, but let's ignore that).
but it often doesn't matter
[1] https://github.com/NVIDIA/DeepLearningExamples/tree/master/P...
If you're not running standard models (where you only have learning rate and regularization to grid over) but you're trying more specialized stuff/developing new models, 4 hyperparameters are not too many.
Reminded me of years ago(It might have been on HN) where some company/group trained a state of the art nlp model to classify if a financial statment or press release was possitive or negative. They had some good results until someone did:
Lets look how far down the press-release-statement the 'numbers' are :)
If it was positive-sentiment it usually was the case that the numbers were high up the page. If it was bad thr bulk of the numbers were much lower in the page.
Almost make sense,if you had goof numbers you want to screM it out. If you had bad numbers you want to preface/explain why first ??
I shared the sentiment but I have began to think otherwise. First of all, we need solutions to big problems. I wouldn't want to tackle Brain tumor problems using DL but someone with the brainpower and resources such as NVIDIA, why not?
But besides that, the whole scientific field is pretty young. An average coder or DL researcher isn't useless because Google has all the machines. There are fields where DL isn't used much. Even in the DL field there are problems to solve. Algorithms to optimize.
The common sentiment in r/ML is that a) big labs need some accountability because they keep getting credit for work they published but never released source code and therefore third party audit is impossible *. Furthermore, Resnet-50 is back to sota because of up-to-date training protocols, and every paper that claimed superiority was using outdated training techniques (or even none at all) for the baseline.
Furthermore, there are many suspicions that the current generation of DL/ML has reached a local minimum because we have collectively pushed for models, architectures, functions, and optimizers that did well on existing hardware. For example, one hypothesis is that the Adam and AdamW optimizers are doing so well on average because we optimized model architectures for them.
* Before somebody claims that this is part of the research process and we should be able to validate and implement models on our own, I would like to point out that there have been multiple instances where the math appeared sound, but the implemented models were bugged, and outperformed the existing models *because* they were bugged (NB biased) and the test sets didn't capture that. When corrected, the models didn't perform any better than existing unbiased methods.
That's interesting, do you have a source? I've been thinking about switching to some fancy new architecture in prod, but wasn't sure it would be worth it.
- Patches is all you need [1]
- ResNet-50 with new protocols [2]
- MLP Mixers [3]
[1] https://old.reddit.com/r/MachineLearning/comments/q35lex/r_p...
[2] https://old.reddit.com/r/MachineLearning/comments/q0vt2b/r_r...
[3] https://old.reddit.com/r/MachineLearning/comments/n59kjo/r_m...
ResNet-50 with updated training protocol seems to be doing pretty well!
I attended the MICCAI workshops this year, and while I didn't go to the brain segmentation challenge, in the ones I did go to the winners were generally academic individuals or groups.
My MICCAI login is on my work computer, so I'm not going to go and check right now, but a quick google shows that the winners of FeTS2021 were all academics: https://twitter.com/FeTS_Challenge/status/144401094937199411... and HEKTOR2021: https://www.aicrowd.com/challenges/miccai-2021-hecktor/leade... looks similar.
EDIT: Academic groups are of course not the same as 'the average professional coder', but this is a fast-moving research field and if an individual is competing with the state-of-the-art that doesn't sound particularly 'average' to me.
EDIT 2: The DiSCO challenge appears to have been dominated by academics, both individuals and small teams : https://twitter.com/GabrielPGirard/status/144397287871477356...
- from the "data" section of http://www.braintumorsegmentation.org/
The three tissue types of interest are fairly easy to identify in most cases. Edema is bright on the FLAIR sequence, enhancing tumor is bright on T1 post-contrast and dark on pre-contrast, and necrosis is relatively dark on T1 pre- and post-contrast while also being surrounded by enhancing tumor. These rules hold true in most cases, so it’s really just a matter of having the algorithm find these simple patterns. The challenge in doing this manually is the amount of time it takes to create a really high quality 3D segmentation. It’s painful and very tough to do with just a mouse and 3 orthogonal planes to work with.
Currently, the accepted practice is to report these changes qualitatively without using segmentations (the way it’s been done for years). While the segmentations created by the models are probably good enough to use in practice today, the logistical challenges of integrating the model with the clinical workflow impede its actual use.
Sure, you could manually export your brain MR to run the model, but that’s a pain to do when you’re reading ~25 brain MR cases/day.
(I know nothing of this tbh, except I once had a demo of a radiologist back when the gamma knife was introduced, have a colleague who became a radiotherapist and a friend who works in ML for Philips medical.)