CodeTF: One-Stop Transformer Library for State-of-the-Art Code LLM
arxiv.org
arxiv.org
It would be helpful to see some Colab notebook examples of how I could use this or incorporate my own codebase with these open source coding models.
The examples show some smaller interesting prediction task & translation between csharp and java, but probably easier to try it out in Colab than having to install locally.
I would also want to be able to compare Github Copilot's autocomplete with what CodeT5 would return as a prediction.
Perhaps this reflects poorly on me, but I find it hard to really stay up to date with the consistent stream of preprints and releases (not just from salesforce) as it tends to take me a while before fully internalizing what makes the newest developments so special, so I was very happy to find that they added an overview that compares different supported models (including their releases) to the newest repo[1].
Of course, size does not correlate with performance, but it still helped me to get a better grip on what they mean by "one-stop Python transformer-based library for code large language models (Code LLMs) and code intelligence" and how that relates to existing models.
CodeTF very grossly oversimplified, intends to make working with models, in a multitude of ways, easier.
I was really surprised that all the HuggingFace stuff needed an account. I didn't have any faith my data would stay local, I didn't understand what that was all for. Which sucks a bit because StarCoder seems to have a fairly friendly vscode extension, Im just too scared to use it.
I think maybe the trick is to just write code comments & ask for help in them? The vscode extension seems to just upload the file, wrapping everything before your cursor in {start token}/* your code here */{end token}.
I'm obviously a total newb here, but new a little tiny bit about LLM, how they are tokenizing systems. It still stuns me a bit seeing that these systems absolutely have the most minimal ability to capture context/hints from the rest of the project, from typescript definitions.
E.g. some models expect very specific formatting of prompts. Some models give nonsense output at temperature=0.7 while others work great. Some models have been fine-tuned on chatbot-bot like instructions (ala chatgpt, so-called "Instruct" models) while others just autocomplete the given prompt. Very important to use the right combo to get the best results.