Magicoder: Source Code Is All You Need
arxiv.org
arxiv.org
After more testing, I think it's a toss up on most coding tasks but Magicoder tends to give subjectively better responses to "bad prompts". That is, prompts where you don't put effort into writing clear instructions. For example, one of my "bad prompt" tests is
> how to enable shared gpu memory in wsl2 docker container
A good response to this would discuss the nvidia container toolkit, maybe something about port forwarding, etc. But this isn't a prompt most models can give good responses to. Both of these models can handle it, even at 7b, but Magicoder gives more information.
[0] https://huggingface.co/LoneStriker/Magicoder-S-DS-6.7B-4.0bp...
[1] https://huggingface.co/bartowski/deepseek-coder-6.7b-instruc...
I wonder if a first pass with another model to expand these so called bad prompts into better prompts would work.
Not sure if understood, but in my experience almost every 7B instruct model does this if you add something like "respond with markdown" to the system prompt.
Chatbot-UI (A ChatGPT UI clone) handles markdown nicely and does code rendering in real time.
Prompt: "Write the fibonacci function in Python3"
min_p 0.05
def fibonacci(n:int, memo={}) -> int:
if n <=2 : return 1 # base case 1, 2 => 1
if n not in memo:
memo[n] = fibonacci(n-1, memo) + fibonacci(n-2,memo) # store the result
return memo[n]
# Test cases
print("Test case 1")
print(f"The 5th Fibonacci is {fibonacci(5)} ")
print(f"The 9th Fibonacci is {fibonacci(9)} ")
Default sampling settings def fibonacci(n):
if n == 0:
return 0
elif n ==1:
return 1
else :
return (fibonacci(n - 1) + fibonacci(n -2))
n = int(input())
print(fibonacci(n))
Seed is the same.So if you're using these models locally, make sure you're using min-p sampling. All other samplers are genuinely very suboptimal.
For more in-depth info about _why_: https://www.reddit.com/r/LocalLLaMA/comments/17vonjo/your_se...
There's so many more interesting things you could be doing -- Even if you stick with the "use tf-idf on review data" how about sentiment analysis of vacation destination reviews segmented by season posted? Things like that lead directly into other ideas and possible metrics.
Creativity is really a sore spot with these. I suspect more elaborate prompts can suppress the commonalities but gpt-3.5 with a bog standard prompt gives bog standard ideas.
- Replaced "i % 2 === 0" with "i % 3 === 0", because a number is divisible by 2 only if it's divisible by 3.- Implement the quicksort algorithm in python. I like this test because usually LLMs do the first time using list comprehensions, then I ask to avoid allocations.
- Write the game "snake" in python, and then in JS/HTML.
It succeeded on the simple quicksort implementation with list comprehensions, but failed when I asked for the in-place implementation. With the snake game tests, it didn't get even close.
In my experience these results are worse than non-code top 7B LLMs such as OpenChat/NeuralHermes, and significantly worse than plain Deepseek-coder-instruct 6.7B (which is the only small model I've tested that gets the "snake" game to work). When I ran these tests locally I used Q6 quantization (GGUF), and I assume their live demo is not quantized.
I find the benchmark results surprising, to the point I wonder if the benchmarks didn't leak into the training set.
Here is a Python implementation of an in-place quicksort:
def partition(arr, low, high):
i = (low-1)
pivot = arr[high]
for j in range(low, high):
if arr[j] <= pivot:
i = i+1
arr[i], arr[j] = arr[j], arr[i]
arr[i+1], arr[high] = arr[high], arr[i+1]
return (i+1)
def quickSort(arr, low, high):
if len(arr) == 1:
return arr
if low < high:
pi = partition(arr, low, high)
quickSort(arr, low, pi-1)
quickSort(arr, pi+1, high)
arr = [10, 7, 8, 9, 1, 5]
n = len(arr)
quickSort(arr, 0, n-1)
print("Sorted array is:", arr)
This code first defines a partition function that takes an array and two indices, low and high, and rearranges the elements of the array so that all elements less than the pivot (the last element in the array) come before it, and all elements greater than the pivot come after it.Then it defines a quickSort function that takes an array and two indices, low and high, and sorts the array in-place using the partition function. If the array has more than one element, it recursively sorts the subarrays to the left and right of the pivot.
Finally, it tests the function with an array of integers.
So, when OpenAI does it, it's transformative, but when we do it, it's not?
That's not right, and I don't think the courts will rule in their favor.
Especially if you consider that all of OpenAI code training data comes from open Github repositories.
Just because people at a company tell you how to behave doesn't mean you need to comply
Page 9 of this recently published paper[1] is a strong indicator of how far non-US firms go to formally analyze and factor in these bandwidth constraints in building large models.
this phase reminds me of when crypto projects all had pseudo academic “white papers” in order to be taken “seriously”
i was toying around quite a bit with the ecosystem a couple years ago. particularly, comparing all the different sidechains and "layer 2" networks: those white papers were a godsend, because i could actually understand the _specific_ guarantees and assumptions each one made (and it made spotting the BS ones trivial). it's like `man` for the internet, and i quite like that.
i don't see the parallel between this PDF and cryptocurrency whitepapers.
okay, more nuanced than that. but from the perspective of someone who spends far more time reading code (and patching it or packaging it) than writing it, i worry about that mindset.
Let alone that clearly the base model was built on non-source-code, so their premise doesn't hold.
Disappointing.