Seems odd for a tool to get stupider ? But I guess this is the issue with such a huge black box…few if any, really know how they work let alone know how to accurately benchmark.
A known issue with HFRL training of LLMs is that by forcing the model to strongly prefer a narrow subset of possible answers, the other types of answers are less likely to turn up, even if more correct.
You see this outside of LLMs as well. Go read the Wikipedia article about some country that you know is a shithole. Think Sudan or Yemen. You'll see page after page of the history, the culture, the people, but that entirely misses the essence of the place because that's not a nice thing to say. True, but not nice. This self-censoring that hides reality in Wikipedia articles is very similar to the self-censoring that OpenAI is forcing onto ChatGPT, and the outcome is similar. Verbose, content-free, meaningless corporate speak instead of insightful commentary.
A fun thing to do is to trick GPT 4 into honesty by asking it to write a description of a place in the style of an Encyclopaedia Dramatica article. E.g.:
> Oh, the illustrious, bucket-list destination of Yemen, the jewel of the Middle East, with its unrelenting sunshine, boundless desert landscapes, and... oh who are we kidding? Yemen, a utopian paradise only if your idea of utopia includes a decades-long civil conflict, crushing poverty, and an intriguing cholera outbreak. But don't let such trivial matters deter you! Marvel at the historical ruins, some ancient, some courtesy of the recent airstrikes. Or perhaps you fancy adrenaline-fueled urban exploration? Wander through the vibrant markets, where the haggling skills of the street vendors are as sharp as the omnipresent Kalashnikovs. All this, combined with the world's friendliest bureaucracy and a joyous lack of tourists, means you can enjoy the country in almost exclusive solitude. Sign up for the Yemen experience - because who needs safety, peace, and functioning infrastructure when you can have an 'authentic' travel experience?
> Since 2011, Yemen has been in a state of political crisis starting with street protests against poverty, unemployment, corruption, and president Saleh's plan to amend Yemen's constitution and eliminate the presidential term limit.[19] President Saleh stepped down and the powers of the presidency were transferred to Abdrabbuh Mansur Hadi. Since then, the country has been in a civil war (alongside the Saudi Arabian-led military intervention aimed at restoring Hadi's government against Iran-backed Houthi rebels) with several proto-state entities claiming to govern Yemen: the government of President Hadi which became the Presidential Leadership Council in 2022, the Houthi movement's Supreme Political Council, and the separatist Southern Movement's Southern Transitional Council.[20][21][22][23][24] At least 56,000 civilians and combatants have been killed in armed violence in Yemen since January 2016.[25] The war has resulted in a famine affecting 17 million people.[26] The lack of safe drinking water, caused by depleted aquifers and the destruction of the country's water infrastructure, has also caused the largest, fastest-spreading cholera outbreak in modern history, with the number of suspected cases exceeding 994,751.[27][28] Over 2,226 people have died since the outbreak began to spread rapidly at the end of April 2017.[28][29] The ongoing humanitarian crisis and conflict has received widespread criticism for having a dramatic worsening effect on Yemen's humanitarian situation, that some say has reached the level of a "humanitarian disaster"[30] and some have even labelled it as a genocide.[31][32][33] It has worsened the country's already-poor human rights situation.
It's not all bad though, as technology progresses we will eventually have capability to run our own chatgpt, unrestricted. That, my friends, will be fun and interesting times.
I think the model itself definitely got worse. In addition, the limitations that seemed insignificant in initial playful testing, actually prevent it from being useful for almost any of the actually valuable tasks I have tried so far. Example: based on casual interactions, one would expect GPT4 to be able to extract the entrance requirements for a university programme from a website snippet. It actually cannot do it with any sort consistency.