HNHacker News
TopNewBestAskShowJobs

stellaathena

421 karma · joined May 7, 2019

submissionscomments
stellaathena··on Paving the way to efficient architectures: StripedHyena-7B
Nope :(
stellaathena··on Paving the way to efficient architectures: StripedHyena-7B
Paper: https://arxiv.org/abs/2305.13048

Models (v4 are the ones from the paper): https://huggingface.co/RWKV

GitHub: https://github.com/BlinkDL/RWKV-LM

stellaathena··on Paving the way to efficient architectures: StripedHyena-7B
RWKV had a paper accepted at EMNLP and released models which match the performance of equivalent transformers.

What else are you looking for?

stellaathena··on The Mathematics of Training LLMs
I'm quite busy right now, but maybe in late September or October?
stellaathena··on Hugging Face, GitHub and more unite to defend open source in EU AI legislation
Instead of flaming about "open source" vs "free software" why don't we talk about how whichever terminology you prefer, organizations like EleutherAI, Hugging Face, and LAION (the AI orgs on this letter) release software that is licensed in a way you approve of.
stellaathena··on Hugging Face, GitHub and more unite to defend open source in EU AI legislation
No, it's about how the regulations governing OpenAI and Google's commercial models, research by non-profits like EleutherAI and AI2, and finetunes of public models done by hobbyists should not be the same. The current law puts all three groups in the same bucket.
stellaathena··on OpenAI regulatory pushing government to ban illegal advanced matrix operations [pdf]
They’re expensive for you, a random person on the internet, but 5-10M USD is really not prohibitively expensive for a government or large company.
stellaathena··on RWKV: Reinventing RNNs for the Transformer Era
It seems that the user is posting ChatGPT-generated text and not a real summary. It's complete nonsense, with about half of the sentences containing a factual error.
stellaathena··on RWKV: Reinventing RNNs for the Transformer Era
Yes I am (apparently) in the acknowledgments. I was not an author on the paper because I didn’t have time to contribute too much. I also try to err on the side of not being added to papers, as my position (I run EleutherAI) tends to encourage people to be overgenerous with offers. I anticipate having more time this coming month and being on the version that’s submitted for peer review, but we’ll see.

BlinkDL has been working on this project for two-ish years, originally in the EleutherAI discord and then created his own to house the project.

I wasn’t thinking too hard about my exact wording, but yes I was thinking of EleutherAI and its various spin-off servers. EleutherAI doesn’t /run/ any of the other servers, but we all have a close collaborative relationship. I’m sure there’s a lot of duplication of membership (e.g., I’m in all of them) but quickly adding up the membership of each server comes out to around 70,000. EleutherAI and LAION are the largest at 25k each, with the others typically having around 5k each. I would expect at least 30k of those users to be unique though.

stellaathena··on RWKV: Reinventing RNNs for the Transformer Era
Twitter thread announcing the paper with some additional color commentary: https://twitter.com/AiEleuther/status/1660811179239849986
stellaathena··on RWKV: Reinventing RNNs for the Transformer Era
Our discord servers (a primary one, a spin-off for RWKV, another spinoff for BioML, etc) have tens of thousands of people between them :) So not quite everyone. But this was a community effort with a public call for contributions
stellaathena··on RWKV: Reinventing RNNs for the Transformer Era
The paper is going to be submitted to EMNLP next month. An early version is being released now to garner feedback and improve the paper before submission.
stellaathena··on EleutherAI announces it has become a non-profit
Our policy is to not comment on timelines for future models, as our ability to meet those timelines is heavily influenced by factors outside of our control and we don’t want to lead people on.
stellaathena··on EleutherAI announces it has become a non-profit
I mean, ultimately there isn’t one. I’m just providing examples of how we fulfill the things that the OP says they want, as they seem unaware of our work.

But I’m confused by the anti non-profit vibes in this comment section. We aren’t saying that becoming a non-profit makes us ethical people, that would be a silly argument. But people do realize that the alternative would be to become a for-profit entity right?

We’re still the same community-driven open collaborative research lab we’ve always been. But incorporating allows us to do things like hire full time staff, enter organizationally binding legal agreements, and protect our members. Between the options of becoming a for-profit and a non-profit, the later seems clearly better suited for our goals.

stellaathena··on EleutherAI announces it has become a non-profit
A cr q
stellaathena··on EleutherAI announces it has become a non-profit
You're in luck! EleutherAI has trained and released open source weights of several LLMs, including GPT-Neo (2.7B parameters), GPT-J (6B parameters), and GPT-NeoX (20B parameters). This last model is currently tied for second on the list of the largest open source LLMs in the world.

We also developed VQGAN-CLIP and CLIP-Guided Diffusion, techniques for doing text-to-image synthesis that don't require training and can easily be run locally for inference.

stellaathena··on EleutherAI announces it has become a non-profit
Actually the origin is this word (with the last two letters transposed), and the Wikipedia article gives IPA

https://en.wikipedia.org/wiki/Eleutheria

stellaathena··on EleutherAI announces it has become a non-profit
Yes, we are currently bottlenecked primarily by engineering manpower.
stellaathena··on EleutherAI announces it has become a non-profit
We have a number of donors including Hugging Face, Stability AI, Nat Friedman, Lambda Labs, and Canva that make our work possible. We also have some orgs that provide sponsorship for computing resources specifically: Stability AI, CoreWeave, and Google Research.
stellaathena··on CarperAI announces plans for the first open-source “instruction-tuned” LM
GPT-NeoX-20B was specifically targeted to fit on A40s, A6000s, and a pair of 3090 Tis. Anything larger than that is going to be a real struggle for people who don’t own computing clusters to use.
stellaathena··on Stable Diffusion launch announcement
Stable Diffusion produces substantially higher quality images in most context, but is much more expensive to produce. The genius of VQGAN-CLIP is that it showed that you could take two pre-existing models and combine them to get text-to-image synthesis to work at all. By contrast, models like DALL-E and Stable Diffusion require extremely expensive pretraining.

There's a discussion of this in the VQGAN-CLIP paper, see in particular 6.1 "Efficiency as a Value" https://arxiv.org/abs/2204.08583

Disclaimer: I'm one of the authors of the VQGAN-CLIP paper and was tangentially involved with Stable Diffusion.

stellaathena··on Accidentally Turing-Complete
I/O is covered by Turing completeness. Whether timing is depends significantly on how exactly you want to formalize things, but certainly any TC system has the capacity to keep track of time if the pieces are manipulated at a constant rate.
stellaathena··on Accidentally Turing-Complete
This paper repeatedly states that it’s PSPACE, but I see no evidence that it’s Turing complete.
stellaathena··on Accidentally Turing-Complete
Tetris can be programmed on a computer, ergo anything that is Turing Complete can simulate Tetris. What am I missing?
stellaathena··on Announcing GPT-NeoX-20B
Yes
stellaathena··on Announcing GPT-NeoX-20B
~40 GB with standard optimization. I suspect you can shrink it down more with some work, but it would require significant innovation to cram it into the next largest common chip size (24 GB, unless I’m misremembering)
stellaathena··on Announcing GPT-NeoX-20B
Coming soon!
stellaathena··on Announcing GPT-NeoX-20B
There is a mirror at https://mystic.the-eye.eu/ that has been up for a long time.
stellaathena··on Announcing GPT-NeoX-20B
It can be run on an A40 or A6000, as well as the largest A100s. But other than that, no.
stellaathena··on A Systematic Investigation of Commonsense Understanding in Large Language Models
That paper also claims zero-shot performance. However, the models that the bulk of the claims are made about are 20x the size of the ones considered in this paper. That's completely consistent with what I said.
Page 1 of 3Next →