2,959 karma · joined August 22, 2011
Ultimately, I don't think this is really practical, and investing in AI proof agents is the way to go.
https://www.eleuther.ai/projects/gpt-neox/ https://github.com/EleutherAI/gpt-neox https://arxiv.org/abs/2204.06745
They have a 20B parameter model. I think the primary dataset for these open models is The Pile: https://arxiv.org/abs/2101.00027 (web scrape, pubmed, arxiv, github, wikipedia, etc. There is a nice diagram on page 2 that summarizes the contents.)
> AllSpice: A git platform for hardware engineers
As others say, there is no standard, and conventions vary by subfield, publication, author and over time. This is esp. true at the research level, where the mathematical content is still being worked out. Subfields have certain conventions, and well-written books and papers will normally introduction notation or include an index of notation, esp. if the notation is novel or they different from the usual conventions. You could start compiling something like this by going through the standard undergrad and grad textbooks for each subject.
https://mathoverflow.net/questions/7120/too-old-for-advanced... https://math.stackexchange.com/questions/237002/too-old-to-s...
You are young and life is long. Go do what you love.
edit: My email is in my profile. Reach out if you want to chat.
edit: submitted: https://news.ycombinator.com/item?id=27179475
The Broad Institute of MIT and Harvard was launched in 2004 to improve human health by using genomics to advance our understanding of the biology and treatment of human disease, and to help lay the groundwork for a new generation of therapies.
The Hail team's mission is to build tools to enable rapid analysis and exploration of biological datasets (100s of TB and tripling yearly). We are committed to open science and everything we do is open source. We currently develop in Python, Scala/Java, and C/C++ and use Spark, Kubernetes, Google Cloud Platform (GCP) and AWS, but will use any tools we need to get the job done. Come help us build the future of big scientific data analysis.
We have two positions:
Update: The Site Reliability Engineer position has been filled.
We also have a front-end/designer position that will be posted shortly. Email below, get in touch if you're interested.
You don't need experience in biology or our particular technologies. We work in a highly multi-disciplinary environment (with software engineers, biologists, bioinformaticians, doctors, operations, statisticians, etc.) Self-improvement is a fundamental part of our culture. You must be excited to be challenged and learn new things.
I'm the hiring manager. Get in touch with me directly if you have any questions: cseed@broadinstitute.org.
You can learn more about the project here: https://hail.is, https://github.com/hail-is/hail
We are one of several software engineer groups at the Broad that are hiring. You can find more positions here: https://broadinstitute.wd1.myworkdayjobs.com/broad_institute
Some discussion for the same question in a recent thread: https://news.ycombinator.com/item?id=21679714
We're quite a bit smaller but have similar numbers: 15-20m right now. We're dominated by build time (build caching might help) and schlepping docker images.
https://www.broadinstitute.org/about-us
https://broadinstitute.wd1.myworkdayjobs.com/broad_institute
I work there. My group builds scalable tools for genomic data analysis:
We're about to post two job reqs, for an SRE and front-end/design position. Email in my profile. Get in touch if you're interested.
gnomAD is the largest public dataset of human genetic variation:
https://gnomad.broadinstitute.org/
They recently a 7 paper collection in Nature: https://www.nature.com/collections/afbgiddede. They're also hiring an SRE:
https://broadinstitute.wd1.myworkdayjobs.com/en-US/broad_ins...
Lots of other jobs at various levels throughout the institute. Biology knowledge generally note required (I had none), although it helps (but be prepared to learn).
The author specifically mentions college friends. Sounds like they were only friends with CS majors even then.
> always come out frustrated
You can't stop there.
Yes, a lot of expert knowledge is locked up in the heads of experts. It is very hard (if not impossible) to write down all the implicit and explicit knowledge that experts have, so it doesn't always happen. It's very hard to become an expert alone. I think this also says something about the nature of expertise: it is something that is constructed by experts themselves in their minds. There was a story that a famous mathematician would tell is grad students, holding up an important book, "You should know everything in this book ... but don't read it!"
You're still confused. There is no function composition here.
op in the example above is just some other function, like +. The associativity of + and function composition are true for totally unrelated reasons. Associativity of plus is an inductive argument that follows from the Peano axioms.
Function composition says:
(f o g) o h = f o (g o h)
as functions. It is true because unary function application "serializes" function applications. Formally, I mean:
((f o g) o h)(x)
= f(g(h(x)))
= (f o (g o h))(x)
Function composition has one value flowing through several functions. Fold has several values flowing through one function.This is not true.
> def List[A].foldLeft[B](z: B)(op: (B, A) => B): B
> def List[A].foldRight[B](z: B)(op: (A, B) => B): B
Notice the signature of the fold op: the arguments types are swapped. This is because fold left and right on a list [a, b], say, is the difference between:
(z op a) op b
and
a op (b op z)
(If this isn't compelling enough, consider [a, b, c].) Not all functions are associative. For example, consider a cryptographic hash function.
> $70K, 20 WEEKS, 100 SAMPLES
Note quite $50K, but close.