118 karma · joined January 20, 2016
We have a massive GPU cluster and developed our own infrastructure to manage the cluster and train massive models.
There's how it works:
- You upload the dataset with preconfigured format into HuggingFaсe [1]. Choose your LLM (e.g. LLaMa 70B, Mistral 7B)
- Place your submission into the queue
- Wait for it to get trained.
- Then you get your trained model there on HuggingFace.
Essentially, why would we want to do it?
We already have an experience with training big LLMs.
We could achieve near-perfect infrastructure performance for training.
Sometimes GPUs have just nothing to train.
Thus we thought it would be cool if we could utilize our GPU cluster 100%. And give back to Open Source community (already built an e2e distributed training framework [2]).
This is in an early stage, so you can expect some bugs.
Any thoughts, opinions, or ideas are quite welcome!
[1]: https://github.com/higgsfield-ai/higgsfield/blob/main/tutori...
Pytorch/k8s/AWS fargate for Deep learning api
Clingo/Prolog for symbolic AI
React.js/flexbox/sass for frontend
SwiftUI for iOS
What:
Node.js/AWS lambda/DynamoDB/Kinesis for backend
Pytorch/k8s/AWS fargate for Deep learning api
Clingo/Prolog for symbolic AI
React.js/flexbox/sass for frontend
SwiftUI for iOS
Why: each tool is optimal for its use case
Impact: move fast and break things)
Email: aida.xcode@gmail.com
GitHub: https://github.com/aidats
Software developer, whose passion is the creation of elegant interfaces with unmatched attention to details. I also understand the importance of creating highly readable and easily maintainable source code. I am constantly striving to learn new technologies and look to ways to better myself.