ChatGPT can get worse over time, Stanford study finds
fortune.com
fortune.com
How is ChatGPT's behavior changing over time? - https://news.ycombinator.com/item?id=36781015 - July 2023 (178 comments)
I wonder if this is the reason behind all of this.
Edit: the study: https://arxiv.org/pdf/2308.13449.pdf
In the gpt4 paper they specifically address this, and find that "Averaged across all exams, the base model achieves a score of 73.7% while the RLHF model achieves a score of 74.0%, suggesting that post-training does not substantially alter base model capability."
* original study: https://arxiv.org/pdf/2307.09009.pdf
* thought it's interesting as we hear a lot of people's experience, but this seems structured and well tracked set of specific scenarios
These days, it seems to be a lot more likely to forget key details from earlier in the conversation. We’re not talking about chats that are anywhere near hitting the token limit. I have to keep reminding it about things when I didn’t have to before.
When you use the longest context window, it tends to place the most focus on the first and last message, for whatever reason. So if you instead of continuing the conversation, you start it from the beginning but with the additional information gained from the previous one in the initial message, it does a lot better.
And of course, I'm horrified by bad decisions politicians and managers seem to make.
What predecessors are you talking about here? The internet is barely held together with duct tape and servers running only because of minute by minute injection of WD50 straight into the fan bearings. There is so much cruft everywhere that is being relied on daily in production systems that most people lost count by now on how much horrible code and setups they've seen in the life.
I spent over five years at AWS, and have worked at other major tech companies since. I continue to not know how the Internet keeps working -- it's basically always on fire, and held together only because of a lot of people who are actively firefighting.
This isn't new, a big chunk of systems you use everyday are based on probabilistic models and heuristics that work most of the time.
LLMs reach very broadly and make the trade off very apparent.
What do you mean exactly? Would be fun to try to replicate.
You give ChatGPT a URL and expect it to be able to get the <title> tag from the HTML? Or you're pasting the whole HTML document?
> Repeat what I've set as the requirements in other words to ensure you understand it. Describe concisely at least two different approaches you could take to solve the problem while explaining the reasoning for why it's a useful approach. Chose the best approach, then think through how the problem could be solved step-by-step. Finally implement each step and provide a full solution.
Yesterday I used that to build (iteratively) a Rust CLI that downloads the last X HN comments have group-counts them by domain. Ended up being ~100 lines that GPT wrote by itself without any errors along the way.
The model was never able to reason; it was only able to generate responses that masqueraded as reasoning. Thorough investigations of its "reasoning process" reveal this at every stage of its development.
Math tutoring is a very legit application. It will be an incredible learning tool — but it needs better logical reasoning abilities with arithmetic. It seems like an easily solvable problem, given that math problems and answers are easy to scale into massive datasets.
Simplified example:
> system-prompt message
> user message: teach me linear algebra
> assistant: Sure, here's how [...] ```math 1 + 1```
> user message (but generated by application): ```1 + 1 = 2```
> assistant: and since 1 + 1 = 2, [...]
LLMs are great at generating text, but not at math, so make something else do the math for the LLM and then it can focus on what it does best.
It seemed like there is some throttling / capacity management happening behind the scenes. For very similar level of complexity and work tasks, during peak time (US day time) I felt the responses lost context or ran into error loops more often.
Similarly working with plugins (Notable in my case) seemed to work fine in off peak hours but got more prone to losing the place, forgetting the default project/notebook, or even in one case simply limiting itself to adding code cells but not executing them -- and instead asking me to go and run the cell and report back the results. At better times it would run the cell, parse any errors, correct the code, and run it again till it gets it right. The amount of code documentation and markdown inserted also seemed to vary wildly. I am not sure houw much of this is tweaking in plugin configuration by Noteable though.
Unexplainable black-box stochastic parrots are no different to this no matter how cleverly packaged they are by AI bros.
OK but traditional small language model Markov chains have context windows too, and they also generate tokens one at a time, separately. Maybe you are arguing that there is a qualitative difference between an eight token context window and an eight thousand token context window. I guess that would be fair enough.
It's "hallucinating" in the same way y=x^2 will predict the next point in a curve.
I hope AI doesn't follow that trend, but I can't see why it wouldn't...
Maybe, Google Search had gotten better, but not as fast as its task has gotten harder.
This crap has even made its way onto Instagram. I've been suggested images of AI-generated figurines resembling Clash of Clans with awkwardly phrased captions about learning to code. "LEARNING: WHAT IS ARRAY" with a picture of an angry lumberjack.
Edit: I'm mostly talking about ChatGPT based on GPT-3.5
There would need to be a better set of tests.
But the March 14, 2023 model is still available on the API.
Further testing is needed.
- Can't reproduce the issue that this paywalled article links to
- Finds that the methodology used to assess "accuracy" isn't actually checking accuracy, but is checking for a specific output, and that output generally does still produce the right value, just in a different format
- Summarizes by saying that ChatGPT is neither the second coming of christ nor horribly inaccurate, just something in between