Like Stable Diffusion, it's a web UI (vaguely reminiscent of NovelAI's) that uses a backend (in this case, Huggingface Transformers). You can use different model architectures, as early as GPT-2 to the newer ones like BigScience's BLOOM, Meta's OPT, and EleutherAI's GPT-Neo and Pythia models, just as long as it was implemented in Huggingface.
They have official support for Google Colab[2][3]; most of the models shown are finetunes on novels (Janeway), choose-your-own-adventures (Nerys / Skein / Adventure), or erotic literature (Erebus / Shinen). You can use the models listed or provide a Huggingface URL.
[1] - https://github.com/koboldai/koboldai-client (source code)
[2] - https://colab.research.google.com/github/koboldai/KoboldAI-C... (TPU colab; 13B and 20B models)
[3] - https://colab.research.google.com/github/koboldai/KoboldAI-C... (GPU colab; 6B models and lower)
I didn't have much luck with stock accelerate, but once gpu is disabled (so it runs only on cpu offloading to nvme storage where ram is insufficient) worked pretty well with me. (there is a small code change that has to be done as the stock software refuses to run without gpu-it is a simple change described in its github issues). My gpu is 8gb vram, but this way I managed to run 7b parameter models. In principle I could run a lot larger ones, but of course it takes a lot more time. The 7b bloom takes 90s for one inference and additional 60s to load the model (from a spinning disc array) initially.
For large LMs, people usually use tensor-parallelism (TP) or pipeline-parallelism (PP). TP involves lots of communication, but uses all GPUs 100% of the time and works faster. PP requires much less communication, but may keep some GPUs idle while they are waiting for data from others.
Usually, TP is used when you have good communication channels between GPUs (e.g., they are in one data center and connected with NVLink), while PP is used when communication is a bottleneck (like in Petals, where the data is sent over the Internet, which is much slower than NVLink).
Check out the infer_auto_memory_map metho which will optimize the model for your configuration (multi gpu, ram, nvme) and then run dispatch model on with that memory map.
Although you need a premium GPU. I admit it's not as good at zero shot or 1-shot as GPT-3 but if you provide examples, you can get as good of output. I feel like the team behind it needs better marketing.