OpenLLM
github.com
github.com
Q: "How do you do a search and replace for a string in VI"
Me: I cant recall right now, i'd just google it"
However, it did make me realize hidden therein is an actual interesting interview question, similar to the "describe what happens when you type an address into the browser's URL bar and hit enter": describe what happens after you type `:s/foo/bar` and hit enter. Followup version: what about `:%s/foo/bar`? The kind of thing that can be interesting to watch them reason through even if they don't know the answer, or even know what those syntaxes do.
To use Discord in good faith and with open eyes, you have to prioritize communication in the present, and give up hope of archiving anything that was said for people who might need the information in the future.
I just googled "how to use openllm" as an example to test your thesis, and the results look very relevant to me.
https://www.google.com/search?client=safari&rls=en&q=how+to+...
"Showing results for how to use openlm
Search instead for how to use openllm"
They stem words aggressively, so searching for "repeater", which is a less common, specific term, gives you results including "repeat", a commonly used word. And there's no way to do an exact word search.
That's only half true. Yes, Discord does allow a "rich" chat experience, with channels and servers, but there the similarities end.
IRC is based on an open protocol, with many open source clients available for it, and a decentralized server infrastructure.
Discord is closed and centralized, with only a single client available for it.
You can easily log IRC channels, but there is no easy way to do that on Discord, if it can be done at all.
I've logged every channel I've ever visited on IRC, and I can use powerful text tools to regex search through all of my conversations on IRC and have the results appear instantly. Nothing remotely like that is possible with Discord.
Paging through IRC logs is virtually instant on a modern terminal, while Discord makes you wait a long time between every other page load, so if you need to look through more than a handful of pages it's incredibly slow and painful.
Some IRC channels have their logs published on the web, making them fully searchable through web search engines, but to my knowledge no Discord channels do that.
What happens in Discord stays in Discord.
Its a shame its being used for general communities who don't use or need the voice chat feature. Especially when its an official community for a place, given their stance on third party clients and privacy issues.
If you don't need voice chat, Zulip, Mattermost, Revolt, Discourse, and many other would suffice (Linen recently got featered on HN). If you do, I think even Signal would be suffice these days.
For Discord search, recently Answer Overflow was recently featured on HN [1].
https://www.reddit.com/r/OpenLLM/
I much prefer the HN/Reddit discussion format to Discord and even Stack Overflow.
You can actually try it out on the main branch :P
are discord conversations persisted and indexed on search engines ?
Very few people actually care about indexing the conversations.
There has long been a place in the ecosystem for ephemeral chat. Often alongside non-ephemeral things like written documentation.
/s
C'mon
IRC chats, especially in opensource projects channels, could and would be archived, published over the web and indexed by search engines.
Can you show me how to access the archives of the ask-for-help channel on the openllm Discord server? Right now they're discussing "loading models on CPU vs GPU". No matter how explicit I got, google did not find the discussion.
#haskell on Libra is publicly logged, but I couldn’t get Google to return a quoted phrase from a message a few weeks ago.
Many people on IRC don’t enjoy being in logged channels. I’ve also heard that there are GDPR implications to publicly logging people’s messages without their consent.
Discussion of the difficulty and downsides of IRC logging, from a coulple years ago:
=> https://news.ycombinator.com/item?id=22892015
=> https://web.archive.org/web/20200417001532/https://echelog.c...
The HN blowback to developers choosing to use Discord is just wildly out of proportion.
On the other hand, if people have to actually converse to get an answer to their questions (like back in the real world), newcomers can more rapidly become part of the community, and help make it more diverse.
There's a balance between engaging with new members and not turning it into a time sink for older members. This is probably a good use case for LLMs.
This to me is the real way through this "Eternal September", where in every "cohort" of newcomers, one or more choose to stay close to the doorway to welcome and guide the next cohort.
I’m wondering how could learn from games, making the content also adaptive to user levels/experiences.
It’s prob also the key agenda in education.
In any case, I'm not arguing that it's impossible, but rather that the more comprehensive the archive, the less welcoming the community would tend to be, all other things being equal. To take it to the extreme, I'll posit the following law: "A well-curated archive is the grave of a community"
You can still have channels open to welcoming new people while at the same time having a large archive of answered questions so that over time a reservoir gets built.
Saying that the same questions getting asked over and over again by new people is somehow a more welcoming community, is like saying that there's any meaningful interaction happening when two people say "What's up?" followed by the response "not much". It's a handshake protocol equivalent without actual depth.
I see friends shaking hands
Saying, "How do you do?"
They're really saying
I love you.* FFS, read the fuckin' archive noob and stop wasting our time
* Hey there, thanks for asking! This is actually a pretty common question, and we have guides written up for just this case. Try entering some of your search terms here [link], and come back with a follow-up question if that doesn't help you!
But yes, in fairness, I'll certainly agree that a community which _chooses_ to respond as the former will stagnate and die.
a lot of people who are into tech stuff already have a discord account making joining the community a one click process, the instant nature of it seems to appeal to younger users more than async forums, it's a fairly mature platform so it has a bunch of moderation/customization/integration features you might want, etc.
> are discord conversations persisted and indexed on search engines ?
nope (and that is a drawback many point out)
Slack was the first to really get that right, and Discord effectively emulated them and made it available for free.
IRC users could get there with bouncers, but those were always a lot harder to get going with.
We might reminisce about irc but we all prefer discord.
Even the searchability of indexed irc has been surpassed by other knowledge sites. It would have to be something extremely niche these days where the only source of info is in an irc chat log
My main problem with Discord is that it's someone else's centralized, for-profit company and has no apparent barriers to enshittification[0]. As Reddit recently demonstrated, it's probably a mistake to build communities on top of something like that.
Matrix is a good candidate for a modern successor to IRC. It's not quite as slick a UX as Discord, but it addresses the main advantages Discord has over IRC.
[0] https://pluralistic.net/2023/01/21/potemkin-ai/#hey-guys
AFAIK most of the gamers choose it for voice chat (Anyone remember TeamSpeak?)
But it’s crazy, people are aggressive about Discord for some reason. I maintain an OpenAI SDK package for .Net, and I had some random person decide they wanted it to be a Discord community, so they created a Discord claiming it was the official community discord for my library, and submitted a PR updating my readme to say that it’s my project’s official Discord. They also replied to several issues and pull requests telling people to discuss it on that discord. If Discord isn’t paying this person in some guerrilla marketing tactic, they should be...
https://github.com/bentoml/OpenLLM/blob/main/src/openllm/uti...
So, eg., I'd imagine ChatGPT would be, say: 100s PB in 0.5TB.
The number of parameters is a nearly meaningless metric, consider, eg., that if all the parameters covary then there's one "functional" parameter.
The compression ratio tells you the real reason why a NN performs. Ie., you can have 100s bns of parameters, but without that 100s PB -> 0.5TB, which you can't afford, it's all rather pointless.
Plus afaik base model training tokens don't have the same effect as fine tuning tokens, so there would need to be a way to specify each of those separately.
Unrelated and I know it is just a representative number but I have seen the training data to be assumed something in this range few times. Entire training set of ChatGPT is almost surely less than a TB or two with compression which is 5 orders of magnitude lower. I believe that such efficient representation of text is one of the biggest reason why text models are working so well but image understanding models are not.
Text is extremely lightweight, so I suppose everything ever written is at most 1-10PB.
This is one of the illusions of text-generative NNs: a 0.5TB weight set is basically enough to store every book. Making claims to "out-sample generalisation" extremely suspicious, and indeed, fairly obviously false.
eg., Ask ChatGPT to write tic-tak-toe in javascript and you get a working game; as it to write duck-hunt and you dont.
You'd need to know how many parameters were independently covarying for any given class of predictions. It certainly isnt all of them.
You could cite the "average dropout percent to random-level accuracy on a given class of problems" (my guess is that this would show 5-20% of parameters could be dropped).
My point, I suppose, is that users of NNs arent interested in the architecture characteristics which affect training -- they're interested in how capable any given model will be.
For this we really want to know how large the training data was, and how compressed it has been. If it's 1PB -> 1MB then we can easily say that's much less useful (but much faster to use) than 1PB -> 0.5TB.
Likewise we can say, if it's a video generator, that 1PB is far too small to be generally useful -- so at best it'll be domain-speicifc.
Yes.
A more helpful bit of information could be what the model was pre-trained on. Assuming they’re trying to refine it for a more specific task.
Size is helpful for “what can I run on my machine” (or how much would it cost to run on a server.) Not all models are created equal, given a byte size, for a given task.
I haven’t done a recent literature review, but my hand-wavy guessplanation is that a NN (as a whole) can adapt to relatively low precision parameters. Up to a point.
* In general. Given actual hardware designs, there are places where you have slack in the system. So adding some extra parameters, e.g. to fully utilize a GPU’s core’s threads (e.g. 32), might actually cost you nothing.
However, newcomers (like me) are pretty blind about minimum system requirements.
Could you please add them to the models list?
For example: what minimum hardware do I need to run Falcon-40b?
PS: If you only have a few setups "known to work" (or just one), listing that would be helpful too.
Every model is drastically different.
If you want to run something on consumer hardware, your best bet is using anything ported to the ggml framework, especially if you're on Apple silicon.
If you want to train a model, that’s a different story.
For example, assuming GPTQ 4bit a 100GB 16bit model needs 26-30GB of VRAM. The smallest video cards which meet this requirement will be 32GB or 40GB cards. (two 24GB cards in parallel work as well, e.g. 2x3090)
The number after q determines how many bits the weights are. Eg q4 means that is 4-bit.
If you use something like KoboldCPP you can only put some of the layers onto the GPU and be able to run larger models that way.
Eg the above linked Vicuna model requires about 10GB of memory at q4, but I have less VRAM than that. I can still run it though.
One can simply do
```openllm start falcon --model-id tiiuae/falcon-40b-instruct --quantize int4```
Beware that there is no free lunch, meaning the quality of inference will degrade by alot when using int 4 quantization
> You will need at least 85-100GB of memory to swiftly run inference with Falcon-40B.
So it may be possible with less (swap around method), though not as efficiently and also slower.
Small ai: likely makes up a function or suggests a single function which isn't sufficient. Refuses to budge from its answer or apologies and gets confused
Large LLM: able to actually understand the question, combine several functions. If it doesn't work you can tell it why and it fixes it
An LLM could have picked up some chess patterns through osmosis, but it can not reason explicitly in the domain.
Instead if having ML that is expert at coding in all languages probably would be better to allow to switch context.
Native mobile dev LLM? - e.g. train only on swift, objc, kotlin, java, c, c++ code
Python dev LLM? - train only on python, c, c++, rust code
Would this way final model be smaller, faster and maybe event better?
Larger models, like OpenAI's GPT, don't have access to this by default.
Just keep in mind the speed differences between the types of memory. If you've got a 170bn 4-bit parameter model, and you're on virtual memory on a 400 Mbps port, a naive calculation says it will take at best 28 minutes per token unless your architecture lets you skip loading parts of the model. Might take longer if the network has an internal feedback loop in the structure.
If you don't have sufficient real RAM or VRAM, the entire model has to be re-loaded for each step of the process.
Assuming no looping (looping makes it longer) and no compartmentalisation in the network structure (if you can make it so an entire fragment might be not-activated, you have the possibility of skipping loading that section; I've not heard of any architecture that does this, it would be analogous to dark silicon or to humans not using every brain cell at the same time (outside of seizures)).
I'm upgrading my M1 laptop on Thursday to a newer model, M2 MAX with 96GB of memory. I'm totally going to try Falcon-40B on it, though I do not expect it to run that well. But I do expect it (the M2, I mean) will be snappier on the smaller models than my original M1 is.
At a glance, it seems like it's going for lots of similar goals (run LLMs with interoperable APIs):
Disclaimer: I helped build BentoML and OpenLLM.
OpenLLM is adding a OpenAI-compatible API layer, which will make it even easier to migrate LLM apps built around OpenAI's API spec. Feel free to join our Discord community and discuss more!
i) if you want to toy around with it
ii) don't want to depend on the api to be available (or don't want to be censored or share sensitive information with a third party)
iii) finetune your own model that you need to deploy by yourself