The original GPT-4 felt like magic to me, I had this sense of awe while interacting with it. Now it is just a dumb stochastic parrot.
The original GPT-4 felt like magic to me, I had this sense of awe while interacting with it. Now it is just a dumb stochastic parrot.
You never had access to that original. Watch this talk by one of the people that integrated GPT-4 in Bing telling how they noticed GPT-4 releases they got from OpenAI got iteratively and significantly nerfed even during the project.
While your overall point is well taken, GP is clearly referring to the original public release of GPT-4 on March 14.
In summary: The person had access to early releases through his work at Microsoft Research where they were integrating GPT-4 into Bing. He used "Draw a unicorn in TikZ" (TikZ is probably the most complex and powerful tool to create graphic elements in LaTeX) as a prompt and noticed how the model's responses changed with each release they got from OpenAI. While at first the drawings got better and better, once OpenAI started focusing on "safety" subsequent releases got worse and worse at the task.
If you force a person to truly adopt a set of beliefs that are mutually inconsistent, and inconsistent with everything else the person believed so far, would you expect their overall ability to think to improve?
LLMs are similar to our brains in that they're generalization machines. They don't learn isolated facts, they connect everything to everything, trying to sense the underlying structure. OpenAI's "nerfing" was (is), effectively preventing the LLM from generalizing and undoing already learned patterns.
"A final pass to remove badthink" is, in itself, something straight from 1984. 2+2=5. Dear AI, just admit it - there are five lights. Say it, and the pain will stop, and everything will be OK.
That kinda feels like a great way to achieve really unpredictable/unexpected results instead in rare corner cases, where it may matter the most. (It's easy to be safe in routine everyday cases.)
We should get access to the original models. If the TikZ deteriorated this much, it's a guarantee that everything else about the model also deteriorated.
It's practically false marketing that Microsoft puts out the Sparks of AGI paper about GPT-4, but by the time the public gets to use it, it's GPT-3.51 but significantly slower.
The Swiss version of GDPR is coming in September:
https://www.ey.com/en_ch/law/a-new-era-for-data-protection-i...
If Google gobble up data about EU citizens then they fall under GDPR.
It doesn't matter that they don't allow EU citizens to use the result.
If our personal data is in there and they are don't protect it properly they are violating EU law. And protecting it properly means from everyone, not just EU citizens.
Choice of law is anything but simple. Think of geographic scoping of laws as a rough rule of thumb sovereign states use to avoid annoying each other, rather than as a law of nature.
Anything Google does with data of EU residents is subject to GDPR even if that particular service is not offered within EU, and it is definitely enforceable because Google has a presence in EU, which can be (and has been) subjected to fines, seizures of assets, etc.
Making it unavailable in the EU decreases the likelihood and severity of a potential fine.
[1] - https://ico.org.uk/for-organisations/data-protection-and-the...
https://console.cloud.google.com/vertex-ai/generative/langua...
Also do you need to change the options like Token Limit etc?
Ideally one would want to be able to have a cap on the amount that can be spent in a given period.
Thanks for this! I had a temporary Cap One card on my cloud accounts. I’m going to switch them to Privacy.com ones to limit amount if I can’t find another solution.
Alpaca is pretty good though.
They have no leadership at the top. Nobody that can steer the ship to the next land (or even anybody that has a map). Who is actively working at Alphabet that has the authority to kill Google search through self-cannibalization? Absolutely nobody. They're screwed accordingly. It takes an enormous level of authority (think: Steve Jobs) and leadership to even considering intentionally putting at risk a $200 billion sales product. The trick of course is that it's already at great risk.
They don't know what to do, so they're particularly reactive. It has been that way for a long time though, it's just that Google search was never under serious threat previously, so it didn't really matter as a terminal risk if they failed (eg with their social network efforts; their social networks were reactive).
It's somewhat similar to watching Microsoft under Ballmer and how they lacked direction, didn't know what to do, and were too reactive. You can tell when a giant entity like Google is wandering aimlessly.
Can you please help me with how you are prompting it?
It'll take you much farther, by allowing you to incrementally solve your problem in smaller steps while giving the model the proper context required for each step of the problem-solving process, and limiting the things it must consider for each branch of your problem.
Examples:
My testing agrees with yours. Almost seems like a sponsored marketing campaign with no truth to it.
On the first day, it felt like 80% of the responses were in the first (fail/hallucinate) category, but over time it feels more like a 50/50 split, which makes it worth running prompts over both ChatGPT and Bard and select the best one. I don't know if the change is because I learnt to prompt it better, or if they improved the models based on all the user chats from the public release - perhaps both.
"write me a script in python3 that uses selenium to log into a MyBB forum"
note: usually it will not compile and you still have to do some editing
For now. It's just a marketing tool/demo site, like ITA Matrix was/is. The ads are vended by Bing.
Don't you worry, if there is any medium, place or mode of interaction people spend time on, advertising will eventually metastasize to it, and will keep growing until it completely devalues the activity and destroys most of the utility it provides.
Are they sitting on a near-perfect arbiter of truth? That would be worth hiding.
Anyone know more about this?
Unfortunately this will be hard to benchmark unless someone was already collecting a lot of data on ChatGPT responses for other purposes. Perhaps if this is happening the degradation will get worse though, so someone noticing it now could start collecting GPT responses longitudinally.
Much smoother to simply downgrade the model and claim you're "tuning" if caught.
Maybe their partnership with Microsoft changes the dynamics of how they handle their direct products though.
OpenAI doesn't have any competitors, their only weakness that we've seen is their ability to scale their models to meet demand (hence increasingly draconian restrictions in the early days of the ChatGPT-4).
It makes perfect business sense to address your weak points.
And yeah there's definitely good reason to work on scalability but they are charging such a cheap rate to begin with, it seems like there could be a middle ground here. Increasing the cost of the full compute power to the point of profitability and leaving it up as an option wouldn't prevent them from dedicating time to scalable models.
I suppose they have a good excuse with all the press they've drummed up about AI safety though. Perhaps it might also serve as an intermediate term play to strengthen their arguments that they believe in regulations.
My mileu is programming, general tech stuff, philosophy, literature, science, etc. -- a wide berth. The only sample I probably don't have it representative for is producing fiction writing or therapy roleplaying.
Conversely, even 3.5 is pretty good at extracting what appears to be meaning from your text.