...however, that was on cheapo ARM hardware. On a hospital budget I bet you could beat OpenAI with a local model pre-mapped in GPU memory on a 3090.
...however, that was on cheapo ARM hardware. On a hospital budget I bet you could beat OpenAI with a local model pre-mapped in GPU memory on a 3090.
But there are a lot of different routes to go. If you want to do what I've done though -
1. Sign up for Oracle Cloud and try to get their free 4 core ARM VPS (these are at-capacity very often, but free if you can get them)
2. Install the Ampere Pytorch runtime: https://cloudmarketplace.oracle.com/marketplace/en_US/adf.ta...
(can also use the ONNX or Transformers acceleration image if you need it)
3. Install your Pytorch/ONNX/Transformers software, link in system libraries and enjoy!
Whole thing works super smooth in my experience. 7B and 13B models are very usable for chatbot type applications.
The hardest part is downloading the 30GB of weights.