24 karma · joined June 30, 2025
The best marketing is usually brief.
Also doesn’t help there’s a lot of red tape as the other commenter mentioned.
Something that was hotly debated in the thread with OpenAI's results:
"We also provided Gemini with access to a curated corpus of high-quality solutions to mathematics problems, and added some general hints and tips on how to approach IMO problems to its instructions."
it seems that the answer to whether or not a general model could perform such a feat is that the models were trained specifically on IMO problems, which is what a number of folks expected.
Doesn't diminish the result, but doesn't seem too different from classical ML techniques if quality of data in = quality of data out.
I assume there are more levers we could try pulling to reduce variation? I'll be looking into this as well.
As an aside, because of my own experience with variability using chatGPT (non-API, I assume there are also more levers to pull here), I've been thinking about LLMs and their application to gaming. To what extent it is possible to use LLMs to interpret a result and then return a variable that then executes the usual state updates? This would hopefully add a bit of intentional variability in the game's response to user inputs but consistency in updating internal game logic.
edit: found this! https://github.com/rasbt/LLMs-from-scratch/issues/249 Seems that it's an ongoing issue from various other links I've found, and now when I google "ollama reproducibility" this thread comes up on the first page, so it seems it's an uncommon issue as well :(
I got back 0.9.3 as well as copied and pasted the prompt (included quotes and no quotes as well just in case...)
I can try the API as well and I'm using a legion 15ach6 but I could also try on my MacBook Pro.
Strangely, I'm only getting 2 alternating results every time I restart the model. I was not able to get the same result as you and certainly not with links to external sources. Is there anything else I could do to try to replicate your result?
I've only used ChatGPT prior and it'd be nice to use locally run models with consistent results.