Lessons Learned from Building an AI Writing App
senrigan.io
senrigan.io
Have you heard of any success on running in a lambda?
I'm adding this link on the article since I'm sure that's a bit more straight forward than my Docker mess lol.
How much do you spend per month on this?
Also why do you need a GPU? Last time I played around with GPT, you didn't need a GPU to run with inference (you already have the weight files).
Also are you running 774 or just 335?
Most of my fine-tuning was on 355 (otherwise, 774 is too hard to train).
However, the default prompt on https://writeup.ai goes to 774.
The GPU with inference is helpful for much faster responses -- for instance, it's easy to run inference (CPU) if you're getting a one way shot of 200 words (no revisions), but if you're constantly changing and revising, then speed ends up mattering a lot for UX.
EDIT: Looked over your articles, I'm literally floored/amazed you're in high school and know this much. The whole world is your oyster.
The one con about TF Serve (TFX) is packaging the entire model into the container (so that ends up being 3gb?+). This was a couple of months ago, so I might be wrong by now ... It was an area I wasn't very confident I could do, TF Serve is really new, so there aren't many guides (and many were already out of date).
It's too bad you didn't find a good guide - if you have the training dump a SavedModelBundle at the end, you can have a production-quality serving microservice up and running in about two lines of code - https://www.tensorflow.org/tfx/serving/docker.
But it doesn't really matter since you got it working.
Is there an easier way? Shouldn't there be some company who will take my money to instantly turn my models into a microservice?
Just a quick google search found these:
https://algorithmia.com/product
(I'm not affiliated with them in any way, and haven't ever used their services.)
Must have cost a fortune if each instance gets it's own GPU.
I suspect the CUDA via docker added a fair bit of complexity.
I toyed around with TF2 on a VM and while difficult to get the versions aligned it wasn't as troublesome as the docker sounds.
I always try to use one of those.
ah didn't know that. I've been spinning up blank nix boxes which does end up a little fiddly until you find a combo that works