210 karma · joined December 4, 2023
- Ability to self host. This unlocks few things: (1) Customized serving stack with various logit processors, etc. (2) More cost efficient inference.
- Ability to fine tune. Most stock instruct models are quite lame at AI story-writing and role-play and produce slop.
There aren't really any pain points specific to Llama, but if we are creating a wish list:
- Keep the pre-training data diverse. There is a worrying trend where some companies apply heavy handed filtering on the pre-training data that's not just based on quality, but also on content. Quality based filtering is understandable and desirable, but please, keep the pre-training dataset diverse :)
- Efficient inference. Open source is way behind closed source here. TensorRT-LLM is probably the most efficient from what's out there, but it's mostly closed source. Maybe Meta could contribute to some of the open source projects like vLLM (or maybe something lower level...).
- A lot of the improvements we saw recently came from post-training, post-SFT improvements. And it's not just the datasets (which clearly you can't just release), but also algorithms -- and most labs are quite secretive about the details here. The open-source community relies on DPO a lot (and more recently, KTO), since it's easy, but empirically it's not that great.
- EFF: https://www.context.fund/policy/2024-03-26SB1047EFFSIA.pdf
- Answer AI: https://www.answer.ai/posts/2024-04-29-sb1047.html
Or use JSON mode with the API.
We need more data.
It's not so clear cut ;) "Research at CERN in Switzerland by the British computer scientist Tim Berners-Lee in 1989–90 resulted in the World Wide Web, linking hypertext documents into an information system, accessible from any node on the network."
By what measure? Phi 2 seems better as far as I can tell from benchmarks and usage and has much more permissive license.
What if there was a standardized way to discover these alternative payment links, so users would automatically go there to find a discounted price.
Why? I think building on top of (and using products built on top of) closed models in this space is not sustainable.
Currently the largest limiting factor, and why I am not spending on marketing just yet, is the model quality at longer contexts. This is my main focus right now, and I am in the middle of data collection for opus-v1 (which will also be released in various sizes). I also hope GPU prices will go down soon :D
It's a lot of fun, being able to work on data, models, backend and frontend all within the scope of 1 project.