GPT-4 Outperforms Elite Crowdworkers, Saving Researchers $500k and 20k hours
artisana.ai
artisana.ai
Worrying that this will no longer be the case.
Perhaps history will show that the NSA made algorithmic breakthroughs a few years ago and realized what was coming, so political policy was crafted to stymie Chinese progress in this field, and what we're seeing in the public sphere from companies like openAI is a managed release of the technology into the public, openAI at least managing to independently discover the same breakthroughs that the NSA made a a few years ago.
Most people don't really think about things that don't affect their day-to-day lives. This includes the specifics of how Governments might run a mass surveillance plan.
GPT-like systems will close that gap, and then comes all of the problems of automated law enforcement - Extrapolation from incomplete data, false positives from coindicences, interpretation errors, all that annoying stuff
No, it hasn't lol.
Take drones for example. The government got really good at those because they made them jet-powered (lol) and blew a bunch of money on server-grade FPGA’s in each one of them.
You can’t really just buy a lot of GPUs to make an LLM work, you need iterative development of architecture and training methods.
Like maybe the government invented self-attention before 2017, but if they didn’t, then the constraint is training time, and the government has the same number of seconds as the rest of us.
Everything you use is from the military.
The government is good at lying, and making themselves ‘appear’ incompetent.
They secretly probably have a much further advanced quantum computer. Your viewpoint is limited to mainstream technology and mainstream science.
The military invented the nuclear bomb yes. But Fermi did most of his thinking work in Italy before the Manhattan project. He got money thrown at him once he got here.
As for semiconductor devices, it was Bell Labs and TI.
Coding is an ambiguous concept that wasn’t really invented, but if it were, it would have first appeared in programmable looms.
The military likes to take credit for things, but really all they do is throw money at existing inventions.
I’m sure they’re throwing a bunch of money at Transformers now, but who are all these uncredited super geniuses who invent things and then let randos at Google take the credit/earn the money?
They made some mistakes in the 40's and 50's to where they had nuclear secrets stolen by the Russians. And ever since has been hyper compartmentalized.
It would not surprise me if in the 90's or 00's they had an internal working LLM, considering all the puzzle pieces. You will never hear about classified tech unless it's a bomb, gets leaked. (See code breaking machines declassed after 70+ years)
In a hypothetical scenario, a military organization might want to conceal its use of a large language model (LLM) for intelligence gathering and analysis.
Another scenario is the military's current interest in everything quantum. Quantum computers for example (you wouldn't want another nation being first and pirate baying out our secrets, would you) so there is an extreme national security importance of being first.
And to be first, you need to have the smart people, which the military has. There is a reason China struggles with jet engines 80 years after their invention, and still can't make nuclear carriers. While the US navy works on things like this: https://www.navair.navy.mil/foia/sites/g/files/jejdrs566/fil...
Jet engines can be solved with money.
The actual steps of making an LLM require complex math and you can’t just pay people to make better math.
And if you could, wouldn’t those people decamp for industry and become literal trillionares?
So your implication is that the military is full of unnamed linear algebra, systems engineering, and linguistics super geniuses and these people never leave, never talk about their work publicly, never publish anything ever, and they're all cool with their huge innovations being kept away from the public forever? All because military IP regulation?
And none of these effective state prisoners ever defect to China (where they could live like royalty) because...
If that is your angle, I very much agree it's possible some deep black project exists that (once) looked into this. In the UFO lore, there are many stories about these advances (tr3b etc).
But all of it is orthogonal to LLMs though. Picking one exceptional area that the military is great at (aerospace) does not suddenly make the military exceptional in other areas like AI.
https://www.youtube.com/watch?v=-wU7nPDcTuY&t
See above video, even code breaking machines are kept a dark secret. LLM's that could make analytical decisions about war and strategy, must have started with the military first. It explains its massive data gathering operations in the 2000's. And it explains why some countries separation to make their own internet, away from what really is the US-Owned World Wide Web.
U.S. military has engaged in the commercialization of top-secret technologies (after it's considered obsolete by military standards), often by collaborating with private companies or research institutions.
Even the Manhattan project had nuclear research going on in public universities at the time.
Nothing of the sort here for the attention mechanism which underpins LLMs we know today.
Fundamental research isn't something you just throw money at and acquire. All we had back then were cleverbot and other expert systems.
The military is responsible for most the technology we use and talk about today. The government may appear incompetent, but we’re living off military hand-me-downs, the entire world is
If today's hardware was available 20 yrs ago, this would've been possible just like the moon landing could've been faked if it took place 20+ yrs later. The technology wasn't available at the time (GPUs in this case, and generally no experience in doing such advanced trick techniques for movies back then)
These models are having such a strong effect now because we've finally got the hardware to run them
They are not deep learning/neural nets.
Also fun fact as a pedant tax: Symantec is so named because they started out as transcription software, hit a wall, and pivoted to security SW.
[1] https://www.theguardian.com/technology/2011/mar/17/us-spy-op...
[2] https://boingboing.net/2015/06/22/gchqs-psy-ops-squad-target...
Because the hardware has not existed.
This said by accident I've seen hardware that was brought to a testing company by federal marshals that was massively parallel custom hardware that was likely for signal processing a lot of channels at once. So there is plenty of custom hardware out there, but these items have not been produced at the scale needed (from what anyone can tell) and, again from what we can tell, they don't have the general processing capability that GPU/TPU driven LLMs have.
not exactly. Where do you think all those budget trillions that don't have to be accounted for goes into? the FBI+NSA (=CIA but for citzens) have infinite resources.
All the overhead they have is to make sure a small subset of the citizens are not impacted. Snowden goes into this in some detail when talking about day to day operations. The norm is to extend the net as wide as possible, until you reach some politician or government agency.
They've created a huge library of unorganized data. The difference here is they now can spawn a million untiring AI private investigators / librarians to organize this information into coherent "case files".
At least for me, until this point I've had a feeling of anonymity in the idea that, while my data is being slurped up, I'm just one data point in a sea of other 'normal' people. There would be little value in spending government time and effort tying all of the web detritus together for me. The juice would definitely not be worth the squeeze.
However, when the cost of this effort is nearly zero, that now becomes a different story. The balance of power between government and the people it rules is going to radically shift.
https://www.cnn.com/2021/04/29/tech/nijeer-parks-facial-reco...
Shotspotter has been billed as a “system of sensors, software, AI and expert human review that accurately detects, locates and alerts police to gunfire”, and the company behind it (formerly “Shotspotter” was the company name, its recently been renamed “Soundthinking”) has a number of other AI-involved law enforcement products now, as well.
That's already how it worked on platforms like mturk and uhrs, lots of the work was transcribing audio dumps from microphones built into computers/phones/smart home devices. UHRS especially had a lot of that (it's owned by MS) as well as search engine grading type work. They also certainly do not pay well, I'd imagine that in practice there isn't much cost difference to paying a bunch of bored people to do it vs the compute cost for running an AI model to do it, but the AI model will be vastly more accurate and will work 24/7.
Imagine you've sent an email about transporting a friend's daughter across state lines to get a medically-necessary abortion. Or if you prefer, imagine you've arranged via email to "lose" some firearms which don't comply with your state's new assault weapons ban.
Pre-LLMs, trying to find these sorts of emails was very hard. A simple text search for "abortion" or "gun" is going to come up with far more emails where two family members got into a political debate, than emails about lawbreaking. Big Brother will find a few such emails here and there by chance, but the vast majority of such incriminating emails will simply be lost in the pile.
Enter LLMs, and Big Brother can feed some of the incriminating emails found my chance into a training dataset along with a bunch of non-incriminating emails, and teach the AI to find incriminating emails, and then apply the model to the entire list of emails and get a nicely filtered list of only the emails which are incriminating, further tuning the model by adding emails it gets wrong to the training dataset when they are found.
GPT-4 can do that task for fractions of a penny per email now. It doesn't have to be perfect if its competing with nothing. I expect we'll see similar shops for any other high cost paper/trail business.
What I'd really love to implement is a way for GPT-4 to answer questions based on a corpus of "all our Confluence pages plus random other sources of documentation." Like with the legal document issue, it's a bit of a nonstarter right now given the proprietary nature of corporate documentation.
"Hey, has anyone worked on Problem X, and what was the outcome of their project"
https://old.reddit.com/r/ChatGPT/comments/12fiwaf/chat_gpt_w...
"Employing Surge AI's top-tier human annotators at a rate of $25 per hour would have cost $500,000 for 20,000 hours of work, an excessive amount to invest in the research endeavor. Surge AI is a venture-backed startup that performs the human labeling for numerous AI companies including OpenAI, Meta, and Anthropic."
What could go wrong? Using GPT-4 to perform labeling used by OpenAI in order to train...uh, wait.
Think about it, how many millions of articles are posted online produced by OpenAI's GPTs to date... Good luck clearing out the training data for GPT-5.
True human content will get gradually scarce. We steer it for sure for our posts, but it is still GPTs that do the heavy lifting.
OpenAI's own classifier fails to detect GPT-4 generated text at the moment.
That's because beyond the 'As an AI language model' and a few key words it can be nearly impossible to detect GPT-4 especially if any prompt is used to intentionally keep it from being detected.
Human like text is a solved problem. There is no more getting better at detecting AI written text, there is only classifying more humans incorrectly at this point.
What is going to happen are private robotic armies making sure private owners remain private owners.
And then, we will go back to times where people were not citizens by default and had fewer rights.
I’m not worried about rogue AI taking over the nukes. I’m worried the same people who think it’s a great idea to charge so much for insulin that people start dying are the ones who will be using AI to hurt people.
Hell, give me a slightly evil AI run amok over any pharma CEO doing their job.
Human greed is the problem. Authoritarianism and capitalism are just subcategories of the greed problem.
What we don't have an answer for yet, is will AGI be greedy?
Either you go build something you want, an island where code doesn't talk back, or leave tech altogether.
I am saddened that I don't even recognise this place anymore. It's not Hacker News anymore, it's AI news. It's starry eyed engineers jumping over each other ready to sell their metaphorical soul. Even I, the Luddite, can't seem to talk about anything else than this bloody thing.
My current plan is to build a small business and retire in the middle of the woods somewhere. Do some Lisp coding while the rest of the world is dancing around their new idol.
/rant, send me an email if you wanna rant about it as well, and discuss your concerns.
That's going to be the exception, not the rule. The benefits from automating crowdwork will disproportionately accrue to corporate profits.
I believe people can contribute in many different ways. When technology enables us to get my work output without me, that frees me up to produce other things for society.
The problem is that it is a disruption for everything because at its core it is a machine for the replication of skill and technology. A concept that has never existed prior with any other technological disruption.
"Climbing the skill ladder is going to look more like running on a treadmill at the gym. No matter how fast you run, you aren’t moving, AI is still right behind you learning everything that you can do."
from a more in depth view I wrote up here describing the rapidly shrinking innovation, disruption and adaption cycles
Sure. But when your “keep the lights on” job cuts you for AI, you are less likely to “produce other things” while you worry about food and heat.
Labelbox does image annotating still, and one CTO said as soon as GPT-4 enabled this for him he'd have his team homebrew it from there.
In other words, the metrics are biased in the researchers’ favor — so GPT-4 would have beat them even more often (probably a majority of the time based on the numbers), if someone else had created the guidelines and golden labels.
All labor is skilled labor. Even breaking rocks.
I'm much more impressed by GPT's ability to handle input than I am in its ability to generate output. It's arguably as good at reading comprehension as most humans.
For example, consider filling the blank:
A giant ______ flew over my head!
It can be a plane. Or a dragon. Or an UFO. Or a balloon. The thing is all of those are correct answers language-wise and the model works correctly as long as what gets filled in conforms to the rules of the given language.
The language that we generate encodes reality to some extent and the model picks up those correlations but there is no concept of reasoning or reality behind it. Maybe it is emergent at some point (as to effectively compress it needs to encode some subset of rules governing our reality) but it is not an agent that optimizes for understanding our reality. Something like Dreamer would be much closer to that.
With an AI like GPT, it is quirky and amusing. Once AIs get really powerful, it becomes scary, and a lot of people who understand this field much better than I do are worried it has a good chance of being deadly. Like, potentially kill-everyone-on-earth deadly.
Personally I didn't need to imagine a specific scenario to understand that there's risk, but I think it would help me convince other folks if I did.
If you want society to collapse all you need to do is succeed in having AI automate all jobs.
Every single country where money comes from somewhere other than people (oil, diamonds...) is an authoritarian nightmare simply because keeping people happy is not necessary.
Once AI can do everything and robots that can do any physical labor are developed the population will shrink dramatically as people with killer robots kill each other for resources. There is no need for AI rebellion or AI failure to get there.
Do we get hoverboards now or is that later ?
Replace "the machine" eith "the market" and it describes some people today.
Surprisingly (to me) many people think that GPT-N will never exceed human level intelligence because it was trained on the internet. I think that argument is obviously wrong.
Another is that I am sure a large chunk of people will never concede that the AI is smarter than them. Literally never, no matter how smart the bot gets. I mean, probably a lot of people think they are as smart as anyone else. They won't agree that someone else is smarter than them, and they certainly won't agree that some bot is smarter than them. It's also a loaded assessment, like they will think that if they agree to that, then they are also implicitly agreeing to cede their personal agency to the bot.
Another possibility is that GPT-N successors that surpass human level cognition will be banned by regulation, like some drugs or nuclear explosives or bio weapons. They could even be pre-emptively banned at some level below human level, and maybe it would never be publicly acknowledged that it's technically possible to go above human level.
I think it's a hard sell to say that nothing better than human is allowed.
But the existential concern comes from supposed exponential takeoff. To me it should be easy to convince people that something 10 or 100 times faster or smarter than humans should not be allowed.
Weirdly you don't see people talking about regulating autonomy which is also part of it.
I find the argument that Eliezer Yudkowsky makes with the super slow aliens to be very compelling especially in the context of fully autonomous AI that people have stupidly designed to imitate human (animal) characteristics like survival instincts. I suspect that that regulators will ban extremely high performing AI. But unfortunately the prediction is that this can't be contained which means it's quite probable that militaries will cause the end of the world just like we have been expecting but in a new way. Since they will likely be excluded from the ban at least secretly.
I think it's weird that you say "I think it's a hard sell to say that nothing better than human is allowed." in combination with "To me it should be easy to convince people that something 10 or 100 times faster or smarter than humans should not be allowed."
To me, those situations are way too close to each other (human level vs. 10x 'smarter') to be able to say one is a "hard sell" and the other is "easy to convince people." I mean maybe you are exaggerating your certainty to make a point, but I don't think it's at all realistic to say that one will be hard and the other will be easy.
The easily convincing argument for me is imagining the planet populated by aliens that move extremely slowly. This makes more sense when it's closer to 1/100th speed. https://www.lesswrong.com/posts/5wMcKNAwB6X4mp9og/that-alien...
Although having actually read that, it's not convincing in this context because the AI in the story is already fantastically performant.
There are just not enough NVIDIA GPUs.
And if your ground truth is problematic, then this is generally a problem of specification and quality control, not performance.
BTW, this is interesting. There is a lot of noise about AI carbon footprint. Now imagine how much humans would eat and fart for 20.000 work hours. It's about 10 man/years. Assuming 8h / 5d / 50 weeks schedule.
I don’t think you can compare people’s carbon footprint because those people will exist regardless of jobs.
But they don't have to /s