Llama2.c running on a Silicon Graphics Indigo2 workstation
twitter.com
twitter.com
Training the model, although it is small with only 15M parameters, takes quite some compute power. I wonder if the resources would have been sufficient at the time. Also, I wonder if disk space would have been sufficient to store the input corpus.