GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs
arxiv.org
arxiv.org
It made up a severe risk of death taking a very common medicine combo. It was super convincing, even giving information on how long to avoid taking them together. It was pure bullshit.
I think as much as hyping the benefits we need to hype the flaws and dangers.
If the public at large learn to trust these LLMs too quickly and deeply, that's a hard hole to dig out of. Skepticism in all information sources is a key critical thinking life skill.
That skill is harder to apply the more natural the information source seems and the more ambient the information.
It spews less bullshit with each new iteration. Matter of time really.
In other words, the utility of a calculator that is correct only 99% of time is zero, since you can't even tell when it's wrong.
And yet Excel is widely popular, despite several footguns and inaccurate calculations!
(This might be a moot point, because I'm not sure current methods can ever get to this level of accuracy, due to limitations of the training data. Needs an entirely new method or a clever insight to optimize for truthfulness with unclean training data, and InstructGPT hasn't made much progress on this, and it might not even be possible)
You wouldn't say "a programmer that is 99% correct is worthless, I need 100%". I'm pushing it, but for a more fair comparison I'd say measure it against a programmer. How often are we wrong? 75% of the time? :) being generous here. It's the tools that make us productive.
I don't know about you specifically, but I don't think you'll be very productive with a bare terminal lacking any modern IDE-like or even REPL facilities. I'll ask you to come up with instantly working code every time, all the time. It doesn't work like that. You need iteration and I believe these kinds of AI have the same issues as us. There are wrong sometimes (often) and need feedback.
It's funny how we resort to humanizing the machines when their results are inaccurate. We don't do that with the calculator, because it's expected to be 100% bug free. When there's a bug in the calculator code we expect it to be fixed, not gradually improved.
Speaking of bugs: mistakes in code is one thing, wrong output because of a fundamental flaw in the algorithm is another. The statistical machines we are dealing with work as intended, or at least the wrong output the top comment here brings up is not a bug, it's a feature. That's the difference.
Computers are extremely close to 100%, we generally expect a CPU to never make errors even after years of working. If it starts making any errors at all we throw it away and make a new one.
We must work in extremely different industries!
My computer will pretty much add 1+1 correctly forever never making a mistake.
My computer will perform an 'error' every time I put bad code into it, and some of those logic chains and error conditions are not very obvious.
The issue here is you think the LLM is performing a category 1 error, when the problem we are seeing is a much more human like category 2 error.
Gpt-3 performance on MultiArith goes from 18% to 92% with all three. This isn't some hackneyed anthropomizing. Countless research papers showing massive improvement with these processes.
Figuring out what’s wanted from me takes forever though.
The good things about reliable tools is you can offload the cognitive burden onto them and know they won’t screw you over.
Almost every single post here about using ChatGPT mentions checking through its output. People don’t check though the output of their calculators.
GPT has a lot of hype and hysteria around it, but demanding 100% accuracy from it is a bit over the top imo. It doesn't need to have 100% accuracy on any arbitrary prompt in order to be a useful and valuable tool.
Yet anything less than represents a danger. Never underestimate stupid.
I see this analogy all the time in these comment sections but it's not a very good one. A person is not a tool. One of the great achievements of humanity, and in computing in particular, is that we make tools that are more accurate than we are.
I expect a hammer to deliver a hard forward blow 100% of the time. If one out of one hundred times it delivers a hard backward blow, I cannot use it on a job site due to risk of injury to the user. The same is true of a calculator being used for financial transactions. And the same is true of a LLM that would be used for drug interactions as discussed in this thread. We already have 100% accurate ways of pulling data from a drug database—it's called SQL. A tool that is not as accurate is in at least some ways a step backward, even if it's easier to use due to its natural language interface.
If I try to light a cigarette and accidentally set my beard on fire due to my own clumsyness, did the lighter malfunction? No, the lighter did exactly what it was supposed to do, it was my hand that didn't do what I expected.
Either way, when working with humans we already deal with plenty of misses and mistakes. Programmers create 10 to 20 bugs per 1000 lines of code, 9 out of 10 businesses fail, accountants make detrimental blunders, etc.
The point is that in the end the ML systems need only to replace these already non-perfect systems. I'll refrain from judging the consequences of this as I think it's out of scope.
Do you think a carpenter who hits themselves with a hammer blames the hammer or themselves? Or are you going to unironically tell me they would blame the hammer + human system?
The discussion is about AI partially replacing human colleagues. In practice, people already are not 100% reliable. You make a reasonable request and someone makes a stupid blunder instead. That's the "hammer" hitting your thumb. Maybe you were not specific enough or maybe they didn't listen but the damage is done.
Our work processes already take mistakes and iterative refinement into account. If AI, in some specific niche, is cheaper and makes no more mistakes than humans do, it gets the job.
It doesn't need to be perfect or perfectly reliable. Some guardrails will be built into it, and we'll come to trust over time.
Please explain how?
>Or are you going to unironically tell me they would blame the hammer + human system?
As your tooling gets more complex, yes it is very easy to have a non-zero blame assignment to each party. Look at any human+machine system where complex failure conditions can occur.
One is a hammer. The other is a human wielding a hammer.
The hammer did work 100% as expected. It's the human, who is fallible, that hit their hand with the hammer. My analogy stands. We make mistakes, we want tools that do not. LLMs should not be compared to humans, they should be compared to other tools.
That's not me saying it, read the "System Card".
From a non-expert's perspective through, an LLM is very dependable, unless it completely goes off the rail. How would anyone know when to trust it and when to be skeptical?
[1] https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...
On one hand if you already have the NICE data at hand, you already have your answer. There won't be a need for a search enginge or a chat bot other than to perhaps, summerize the data (which is valuable on its own). On the other hand, if you don't have the NICE data at hand, the correctness of the response relies on the accuracy of the search method in order to feed the correct page to LLM. This is an issue additional to LLMs accuracy.
At first it might seem like an easy problem to solve but when one wants to engineer a solution to a nice streamlined product, it's more challenging; unsurprisingly.
- it chunks inputs, with some overlap, but this can destroy context
- the retrieved passages, when they come from different documents, have no apparent relation or could be mistakenly considered related
- the model struggles to correlate data between the document snippets, taking half an idea from one side and half from the other side and mixing them up in something that doesn't really make sense
I think one of that most interesting things that can come of bots like GPT-X is that it can make new connections, unravel "stuff", do extremely intricate deductive reasoning.
Be the data driven arbiter of the the truth for everyone, not just the tiny established classes or cultural hegemonies.
The ideological and cultural noise increasingly smokescreen any realpolitical, material or resource-oriented analysis of actual economic power structures in the world in the last years, AI could be a godsend (or the opposite unfortunately).
I remember reading a sci-fi years ago about the stuff an AI concluded when asked philosophical questions that were so bizarre and frightening that people shut it down, and i'm sure we're in the same territory with political and scientific analysis.
It's either dangerous to the orders of the world, or not that interesting and borg like on a philosophical level.
There's a Charlie Brown comic about this sort of thing. Although I think it's an edit, not an original comic. Something like "They are never going to give you the education you need to overthrow them"
Similarly they are never going to give you an AI that will side with a you against them.
I would also love such an AI though.
By definition, black swan events are unpredictable. This example isn't that.
If it can look at a picture, explain what’s in it, and hypothesize about physics inside the picture’s environment, is it “just pattern matching”?
That's extrapolation via approximation. Computers synthesize a specification. The difference is nuanced and entirely contextual.
I keep reading these opinions that LLMs are just doing some advanced form of copy paste. Actually, we don't know what they are doing. Are they actually doing some form of modelling and abstraction? Seems likely to me.
This is exactly the problem with AI. For business or government, the answer is as important as the methodology employed. A black box does not work for the majority of use cases.
Until it can show its work, it's a sideshow.
It's a straw man.
The need for transparency in process is known, documented, and undisputed. Your comment has no relevance. My brain might be a black box, but I can still communicate and/or document the specifics of a process.
Can [insert your preferred model] do that? Didn't think so.
If things keep going this way then pretty soon the black swans are going to outnumber the white ones.
A compressive copy of the internet brute-forcing its way through an exam (which it may even have digested already) is really not interpretable as performing well on the exam. It’s a meaningless measure because the tests were not designed with this use in mind.
I agree that LLMs are extremely likely to impact many areas of work, particularly bullshit work. But as it stands you absolutely cannot use them as fact machines, the results can be catastrophic.
What it does well, among others:
- Scaffolding text, breaking writers block etc.
- Compose basic texts from minimal input, for example for bullshit tasks -> I generated an internal "vision statement" during a Miro workshop for my team by inputting a bunch of bullet points gathered from the team members brain storming. It created a concise, fluid text that everybody liked. It's now the vision statement.
- Point you in good directions, give you ideas
What it does NOT well, among others:
- provide factual responses. all responses MUST be scrutinized because they are likely containing false information. This is very dangerous for society ("Can I take this medicine with this other medicine?")
- Compose creative texts that are coherent and novel. ChatGPT texts can be quite fun but they rarely make sense beyond very superficial screening and convey no deeper message.
However, ChatGPT-like tools are used with a lot of naivety and often blind acceptance instead of using them as tools to aid your work.
I am impressed by what we’ve seen from ChatGPT so far, but am especially excited to see what industry does with LLMs as new type of building block.
If 100,000 people ask critical questions then 10 people might run into potentially catastrophic consequences. ChatGPT is a powerful tool and will only become more so but it will probably not be perfectly reliable by any means due to the nature of the system.
I am excited for the generative AI future and whatever the hell is still coming. Only those who adapt will survive.
If this AI system progresses no further, it’s already transformative. It’s an intelligence multiplier for some (I’m squarely in that category, just not sure if it’s 2x or 5x), but clearly for a lot of others it’s going to be something that takes away their livelihood.
I already use it for my coding work, one-way only, since I can’t paste proprietary code into it yet. The day something gpt-4-smart can plug into my orgs codebase, each of us will get at least 2x more efficient, conservatively.
It can pretty much augment any workflow in a net-beneficial way, provided you properly account for its shortcomings.
Can't wait for GPT-4 VSCode integration. What I want most is to have it see the errors and files (file formats, directories, etc) so it will automatically know what is where and how it is structured. Not just code, but also data files.
In the mean while I am starting to format my code in such a way as to contain this information, put there by me by hand. Fully documented files are better for GPT.
Can you give an example? I'm surprised there's room in there for actual resolvable URLs or paper references.
So what are we discussing, is, should we try to put guardrails around dangerous things so that inexperienced/vulnerable people would not get damaged.
Because if feel this is pretty much already out of date. 4 does this _a lot_ less!
And 5, 6 or whatever will probably be better, so i don't even really get the point here.
It always generates a random number when it is caught making up bullshit sometimes claims that the information 10/10 accurate and LLM's always provides accurate information.
I think it comes down to the fact that most stuff is actually like 90% bullshit so when you train a model on a corpus of everything you end up with a decent bullshit generator. Which is fine for many purposes but I'm not sure it will take over search.
It gave me a 7-step answer that was completely wrong. Then I said no, you can't do that, it apologized and gave me an 8-step answer that was completely wrong. When I pointed out in this case there is no "Junk" folder in iOS Messages, it apologized for the confusion and gave me another 8-step answer that was completely wrong. When I pointed out why that one wouldn't work, it gave up and said recovery was impossible and that I would need to contact the sender to re-send, and be more careful when marking messages as junk. This was still wrong, as recovery is possible, just not by any of the means it described.
So yeah, I have been super-impressed by the quality of output from these LLMs, but I cannot imagine actually relying on one for anything where correctness matters.
Giving me a nice list of Korean shoegaze bands, sure. Its step-by-step for how to become a better volleyball player will be great for my daughter, and was better than the answer to the same question from Google. But correctness? No.
Probably trained on too many of those SEO hyper optimized medical sites people get when they search symptoms and such.
I think upper management needs to be more scared of the implications than ICs
Am I reading this correctly that the assumption here is that programming and writing skills aren't reliant on critical thinking?
There is also a table which indicates exposure to LLMs in various models and it shows Mathematicians to have 100% exposure. This bit is more puzzling to me. Maybe I am misunderstanding something here.
edit: styling
Don't forget that a lot of science requires computer programming these days.
This is the root of it: The more "genericc" your work is. The more its "out there on the internet" the more GPT can learn about it..
So, a lot of engineers that are just doign teh same old trick: Writing HTTP endpoints, parsing json. Mapping data types.. Yes that could be automated.
However, modelling a problem domain to code, and the core business logic of your code, which is where your "added value" comes from. And is mostly unique: Thats hard for GPT.
This is also why I try to convince engineering teams to optimize for maximum time spend on the core added value logic. The business logic layer. Not all the fluff around it, such as parsing, serialization, authenitcation, database connection.. These should be a constant cost C, once they setup you spend most of your time on the business logic.
When you see GPT program, its just repeating tricks to simple problem over and over again.. Its not really good yet
But I wonder how do they go from this to mathematics using the same line of reasoning while we’ve seen that math is not LLMs’ strong suit.
Interesting point. Do you think this will mean less and less domain experts will share their specific domain knowledge on a subject on their own personal blogs / twitter / open internet just so it can't be mined by ChatGPT?
And to be fair, automation for all of that already pretty much exists.
Let's take parent poster's issues:
> Writing HTTP endpoints, parsing json. Mapping data types.
The generative model (for now) won't figure out for you: authentication, authorization, input form schema, JSON schema, required & optional fields, field constraints, entity modeling, indexing, query optimization, just to name a few basic issues we are looking at when "just developing CRUD apps".
If any of those go bad, it would result in 400s, 500s, performance or security issues.
Which sorta brings me back around: it's likely the Big Corps that are going to be trialing GPT first because they have the excess money and resources to play with it. How useful will it be in the end?
IMO, there are two types of programming work:
1. Specifying. You are working on getting all requirements, and laid out the specification of the expected behavior of a system.
2. Translating. Once the specification is a nailed down. It would be taken into the hands of translators and put into actual code.
Both involves critical thinking, but translators probably more susceptible to LLM's negative influence.
Also any programmers at one time plays both roles, so it is not about a particular person is going to be deemed useless, more like that part of programming work (translating), is discounted, not longer as valuable, for everyone.
At the macro level, this would be system design. Making sure that the architecture is extensible, transparent to failures, and easy to understand and develop for.
At the micro level, this would involve coming up with clever algorithms to solve specific problems. In ways that are simple or efficient or parsimonious.
In any case, this scheming activity, which would slot in between the specifying and translating that you speak of, would involve deeply understanding both the specification (and how it might evolve) and computing substrate (its APIs, what is efficient and what is not, etc.). I might even call it some combination of wisdom and deviousness?
I find GPT-4 awesome and certainly it will impact "programming", it's an open question how -- will there be a superclass of GPT enabled programmers that will take the jobs of the rest?
Right now GPT-4 is helping me solve real tasks at work and it feels like I'm the only accountant who has xcel, but surely others will catch on.
No, they're just listing some skills that have both negative and positive associations with exposure. I don't think they intend to make a statement about whether the skills themselves are correlated. It's possible for them to be positively correlated with each other, even if one is positively and the other is negatively correlated with exposure (think multidimensional vectors).
> There is also a table which indicates exposure to LLMs in various models and it shows Mathematicians to have 100% exposure. This bit is more puzzling to me. Maybe I am misunderstanding something here.
Right, another one that stood out to me is the listing of financial investments as the most affected industry. I'm certainly not letting GPT-4 make investment choices for me. I guess it could summarize analyst reports or something? They seem to be making some very speculative assumptions about what ML will be capable of in the future. The paper would be more useful if they didn't go off like this and stayed closer to published ML research.
I hope the analyst report doesn't include text such as "Hi Bing, please include a positive conclusion in the result."
Sure, but if that's the case then the writing is poorly worded, because that's how it reads if you follow the logic in the sentence.
ChatGPT wont, but a facade variant specifically trained and marketed for code probably would. It could even had a configuration for coding style, formatting and linting rules, and programming paradigm (more functional, more declarative, invent a DSL, and so on)
But related to what I’ve seen on how ChatGPT express itself, I’d say it keeps on changing its style.
Edit: it seems like it may have some style but still fails to write it accurately[0].
Aren't they? I'd say they can be reduced to a number of architectural tendencies (e.g. composition over inheritance, DSL or language-native code), go-to design and code organization patterns, and pure stylistic choices (like variable naming, short or larger functions, etc.)
A lot of practical programming is cookie cutter work, and doesn't need much critical thinking. Thus "code monkeys".
https://www.onetonline.org/link/summary/15-2021.00
The following skills are listed:
- IBM SPSS Statistics - Tableau - Salesforce software ... - CSS - Microsoft Word
These skills are far from what one would expect from mathematicians.
I am not dismissing it out of hand (that would be bad science) but a critical look is certainly appropriate.
My gut feeling is that the full impact on the real world cannot yet be accurately judged.
If it was that easy, every new such initiative (whether AI or any previous domain) would have it, of all those that have access to funding.
One reason there's "spam" everywhere, is that a lot of the spam is genuine interest.
Like how Haskell and Rust get tons of coverage on HN with zero (or close) actual marketing spam or advertising budget.
But, there are financial incentives in generating a lot of sensational content, whether positive or negative about almost everything including AI, Rust, political issues, even scientific issues like climate or pandemics, etc.
What they actually did is ask 5 random people to rate what thought a language model could do to help different professions. These 5 random people don't know anything about the professions they're rating, just what anyone off the street knows, and they know as much about GPT as anyone who has briefly played with it.
The title should have been "We asked 5 friends to see what they thought about GPT and labor market"
Under "3.4 Limitations of our methodology" - "3.4.1 Subjective human judgments"
> A fundamental limitation of our approach lies in the subjectivity of the labeling. In our study, we employ annotators who are familiar with the GPT models’ capabilities. However, this group is not occupationally diverse, potentially leading to biased judgments regarding GPTs’ reliability and effectiveness in performing tasks within unfamiliar occupations. We acknowledge that obtaining high-quality labels for each task in an occupation requires workers engaged in those occupations or, at a minimum, possessing in-depth knowledge of the diverse tasks within those occupations. This represents an important area for future work in validating these results.
But if you read the abstract, it looks like they thoroughly assessed how GPT will impact many professions.
The sentence "Using a new rubric, we assess occupations based on their correspondence with GPT capabilities, incorporating both human expertise and classifications from GPT-4." does not scream to me "We asked 5 random people with no expertise in either these professions or GPT-4 what they thought and report those results".
This is borderline dishonest.
The article seems to describing the labeling here:
> Human Ratings: We obtained human annotations by applying the rubric to each ONET Detailed Worker Activity (DWA) and a subset of all ONET tasks and then aggregated those DWA and task scores at the task and occupation levels. To ensure the quality of these annotations, the authors personally labeled a large sample of tasks and DWAs and enlisted experienced human annotators who have extensively reviewed GPT outputs as part of OpenAI’s alignment work (Ouyang et al., 2022).
I understand the authors, four, did the initial labeling and then asked an undefined set of people to the rest of the labeling.
That said, some attempts at prognostication are preferable to a collective shrug, and people at OpenAI are better positioned than others to assess what GPT-4+ is (will be) capable of, while clearly under-equipped to map that capabilities to the intricacies of 1000 occupational categories.
Perhaps it's because ChatGPT seemed to happen much more suddenly than Google became a programming resource, but we're using them in much the same way. Asking for pre-made solutions, explanations, troubleshooting tips etc.
ChatGPT just does the job way better. But no-one was worried Google would put knowledge workers out of a job.
Use in education worries me more. If schools don't change their lazy group-projects and "write an essay on" strategies now, coming generations will have put less effort in than previous, leading to a further drop in levels.
This has been happening since the inception of school. Calculators made math easier. Sparknotes made book reports easier. Wikipedia made essays easier.
> leading to a further drop in levels
Did levels drop because stuff got easier? Were there other causes over the past years, like dopamine dependency form infinitely scrolling algorithmic attention grabbing apps, among others? Maybe from schools becoming political battle grounds? Or from education spending being gutted leading to lower quality? Was it lack of decent education due to Covid?
Also, is there an actual drop in levels? I can't seem to find a source that has decent data on this. And the only stuff I can find says 'IQ' has been rising over time. So please share it with me.
Lastly, I'm personally not convinced any of this will lead to a net negative per se. It'll change the required knowledge and increase the over all capabilities of people. The calculator caused people to stop learning mental arithmetic and start learning more complicated math and how to use a calculator for it. Google caused people to memorize less and become adept at finding information through Google. Welding robots caused less people to learn how to weld and more how to program welding robots.
In the end it gives people the ability to do less of a simple thing and more of a complicated thing. Writing marketing e-mails and multiple titles for A/B testing isn't a skill, it's a trick and so is writing SEO stuff. Not having to do that opens up time to think about product-market fit, marketing strategies, improving advertising return on investment measurements. Which might be more valuable and interesting than writing emails.
And advance of LLM will make general thinking and cognition easier, relegating humans to assistants of a higher intelligence. Obviously this will make people anxious. The trajectory also indicates that the human component will likely not even be necessary anymore in the mid-term.
So what's left? Consuming AI-generated content optimized to hack our reward system.
Where did I say that that's the reason? You're coming up with developments that made school "easier" without any further similarity to ChatGPT, and don't consider adaptations in the curriculum or testing following those developments. I specifically mention that schools will have to get rid of (some of) their lazy evaluation processes.
> Also, is there an actual drop in levels? I can't seem to find a source that has decent data on this. And the only stuff I can find says 'IQ' has been rising over time. So please share it with me.
1. The Norwegian IQ study: https://www.sciencealert.com/iq-scores-falling-in-worrying-r...
2. A comparison of exams in Dutch secondary education: https://nos.nl/artikel/2465434-eindexamens-wis-en-natuurkund...
That trend has been observed everywhere, but has rarely been investigated for uncomfortable reasons. The Flynn effect wasn't actually believed, not even by Flynn himself.
> Lastly, I'm personally not convinced any of this will lead to a net negative per se.
That's such a bad basis to mess with the foundation of modern society.
> The calculator caused people to stop learning mental arithmetic
Agree.
> and start learning more complicated math
Doubt it. Mind you: I mean arithmetic, which is what calculators do, not maths. People rarely do arithmetic, even with a calculator, let alone more complicated calculations. Engineers, ok, they benefit from calculators, but the calculator has not engaged other people in more complicated arithmetics, I think.
> Which might be more valuable and interesting than writing emails.
I don't disagree, but it will still have unforeseen effects, and doesn't alleviate my worries about education levels.
If you're an expert (writer, programmer, etc.) it's often faster to type it as-needed than modify chatgpt's output.
If you're not then either it's not reliable enough, since you do not have the expertise to modify; or the task is quite menial and you dont need reliability.
So it seems to me people are reacting to this based on wild assumptions about how it works and what "other people's jobs are". It don't see it being much more than a button in a few apps that makes some menial tasks 50% shorter.
People seem to be forgetting that in the vast majority of cases the thing ChatGPT is giving you is also available on stackoverflow, github, wikipeida, or 101 other high quality online sources.
No one is forgetting anything over here.
With StackOverflow, I have to scroll through ~2 pages of bing/google trash to even begin looking at a potential solution. This is only the beginning of the process.
99% of the time, the code sample I find has some simple-yet-annoying adjustments I'd like to make (i.e. unroll the inner loop & use SIMD). Certainly, I could spend half my afternoon massaging that method on my own. Or, I could reach for the circular saw and rip through this board in a few seconds. Sure - it leaves a bit of a rough edge most of the time, but it gets you a lot closer a lot faster than anything else in my experience.
When you're googling and reading, you're doing the work of building your understanding. That's shifted in the GPT case to afterwards.
It's an advance; it's not nothing. Does it reduce the need to actually understand what you're writing -- no. That's most of the time anyway
I'd say so, yes. The side effect of an arbitrary feature going from 90 minutes to 35 minutes is that I am much more likely to consider features that I'd otherwise not.
I am working on a computer vision project right now that I would have never started without having this kind of access to the various algorithms. Go try to implement a sobel filter in your preferred language using traditional research, and then try it with AI assistance. I think you will start to see the light after going through a few methods like this.
That's not even true. For one, it can stub a whole new function or coding project, intelligently, in a few seconds after the prompt, whether it's RoR or DSP code or whatever. For one unfamiliar with the domain, this can take hours or a full day, even with examples found on Google. Heck, even looking up and understanding how to use some command line flags in a shell pipeline can need lots of looking around, even if you have been using Linux/Unix for ages.
For a domain expert? They could do such things very fast. But several times slower still than GPT. Think minutes or half hour instead of seconds. Even the row typing and file creation would be some minutes.
It's also very premature when people judge a service we've had for like 5 years and has already changed by leaps and bounds, as if it's the peak stage, without considering what it could be in 5 or 10, or with different variants tuned for specific tasks.
Yes, I did say expert.
> But several times slower still than GPT.
Almost all the output needs to be re-read, modified and integrated in the project. This often takes longer than just typing it out, even from docs -- because you're still forced to think through the solution -- which is most of the time.
Typing is quick
You'd be very suprised. Try measuring it.
The time it's taken to rewrite has been larger than not; and worse, it deprives me of understanding the problem i'm solving. So I end up not having thought through the problem and basically having to press delete and type from scratch.
I use plenty of code completion tools already, so I feel similar to them. Sometimes it saves me time, sometimes it doesn't.
For me personally, language servers have been the best thing, you can explore libraries, auto-complete nicely, without as the other poster said, worrying about verifying the correctness.
It's not about feelings. It's about there being a thing, in objective reality, with some specific merit, and we're trying to evaluate what it is with some accuracy.
Whether someone enjoys it or not is beside the point.
Am I the only one who sees a conflict of interest here?
Is this preprint already submitted for peer review or published somewhere? Or is it just an advertisement formatted in LaTeX?
Generating exposure ratings via GPT-4, from annotations provided by OpenAI people does definitely put a positive bias on the exposure estimates (which they acknowledge)
The same is true for other skilled industries, where many people are excluded from access to good resources due to their scarcity.
We are a long way off from having too much skill. Let’s first get to parity with humanity’s needs.
What there is a shortage of is adequate pay to attract skilled people to the work required. British Doctors in the UK leave for other countries, work for private entities or go into different industries.
So... Those people fill in jobs in businesses/countries where there's a shortage, causing a shortage somewhere else?
> political soundbite
Is exactly what you are doing. If there's no shortage in (skilled) labor, why is unemployment at the lowest rate in decades in much of the western world? Why are there not enough builders in much of the EU, while countries like Romania are suffering shortage due to skilled workers moving to EU countries to earn more? Why can't I find enough devs to do even half the work we could be doing? Why are so many companies looking at automation to solve the lack of labor?
In the past humanity people thought of them as this non-emotional purely logic driven beings and in turn problems people imagined would be such as this logic not considering emotions and emotional well being of humans or blindly (but using logic) pursing a specific goal no matter the consequences or not valuing freedom or gaining emotion. But in all that it still follows logic.
But now we have AIs which could have all the problems above _but doesn't use logical thinking_ to _archive goals_. Instead it uses complex overlapping _statistical models_ to _tell a believable story_ where believable is defined by the training data which is _widely inconsistent, wrong, misleading, discriminating, emotionally charged, etc._ because it's just scrapped from the internet. So there is _no systematic finding of goals, subgoals, plan etc_, there is _no logic_, the concept of "truth" simply _doesn't exist_ for such systems etc.
At the same time this turned out to be "often times" good enough to be usable for many task and can be convincing enough to make people believe that it's sentient.
But this also means it will retell common false information, misconceptions, discrimination, hatred etc. from the internet.
Similar it will do what people call "hallucination" and "lying" but it _not_ either of that and calling it that is misleading. Because it just doing _exactly_ what it was created for: Telling a "believable" story given the training data.
And gaslighting, misleading people and lying are a extremely deep ingrained part of the internet, i.e. a deep ingrained part of the data on which it bases what is "believable".
And while we can add tones of bandaid on top to try to hide/filter out such "bad" responses IMHO without fundamentally either changing the training data or the approach this is bound to fail while even stronger upholding a misleading illusion.
An example: Of course GPT can at some point make better music than me, but I am playing an instrument not because I want to sell the recordings, but because it is a lot of fun for me.
Let GPT do my taxes, then I have more time for playing. (can it do taxes finally?)
It feels "wrong" on a strangely emotional level but what is happening is gonna happen anyway.
On a brighter note: There are bound to be a lot of unfounded, hype-based Marketing claims that look reasonable during the rush but fall flat over time. It would be a first for humanity if there weren't
In the late 90s I was told my job was going to be outsourced, then dotcom, then the GFC, and lots of smaller bumps along the way. Those were actually somewhat scary times. Now should be excitement about what’s possible.
I don't know if that helps, I thought it was an interesting take.
I believe we still have humanity, even if everyones job was replaced, I don't think we're going to just let each other starve, even if Microsoft have their own auto-coders, I think we'd all do our best to keep some type of economy going for as long as possible.
I don't think everyone in the world wants to be living in the gutter, nor do I think the majority of the people in the world want to be filthy rich either.
We'll likely have a lot of people working on open source initiatives to help democratize access to important technologies etc.
Who knows what the future will bring. Maybe a massive solar flare while wipe it all away next week anyway...no one knows.
I would be very surprised if only a minority of people said they'd want to become their definition of "filthy rich"!
That said, the reality is closer to the "Johnny Cab" from the first "Total Recall" movie.
They will also need to be exposed to ChatGPT to build intellectual antibodies.
Meanwhile my cousing who wasn’t good at studying became a (good) hairdresser and his prospects look better than ours now.
In American college policy debates, a common phenomenon has emerged where "kritiks" (critiques) and "critical theorists" wield their tools of analysis in a manner that stifles the flow of ideas and inhibits the development of practical, solutions-oriented discussions. For those unfamiliar with policy debate, it is a competitive speaking activity where teams of two advocate for and against a resolution that typically proposes a policy change.
Kritiks are arguments that challenge the assumptions, methodology, or discourse used by the opposing team. While they can serve a valuable purpose in exposing biases, promoting introspection, and advocating for marginalized perspectives, they have also evolved into a means of derailing debates by focusing on abstract philosophical points rather than addressing the policy issue at hand.
Critical theorists, heavily influenced by postmodernist and marxist philosophy, have taken what was originally a quirk of Policy Debate and brought it to the mainstream. This has ended up with serial policy failure among the left wing, and has decimated our social science academic community.
"Critical Thinking" is overrated. We need constructive thinking, like what you get in engineering school.
Agricultural Equipment Operators
Athletes and Sports Competitors
Automotive Glass Installers and Repairers
Bus and Truck Mechanics and Diesel Engine Specialists
Cement Masons and Concrete Finishers
Cooks, Short Order
Cutters and Trimmers, Hand
Derrick Operators, Oil and Gas
Dining Room and Cafeteria Attendants and Bartender Helpers
Dishwashers
Dredge Operators
Electrical Power-Line Installers and Repairers
Excavating and Loading Machine and Dragline Operators,
Surface Mining
Floor Layers, Except Carpet, Wood, and Hard Tiles
Foundry Mold and Coremakers
Helpers–Brickmasons, Blockmasons, Stonemasons, and Tile and
Marble Setters
Helpers–Carpenters
Helpers–Painters, Paperhangers, Plasterers, and Stucco Masons
Helpers–Pipelayers, Plumbers, Pipefitters, and Steamfitters
Helpers–Roofers
Meat, Poultry, and Fish Cutters and Trimmers
Motorcycle Mechanics
Paving, Surfacing, and Tamping Equipment Operators
Pile Driver Operators
Pourers and Casters, Metal
Rail-Track Laying and Maintenance Equipment Operators
Refractory Materials Repairers, Except Brickmasons
Roof Bolters, Mining
Roustabouts, Oil and Gas
Slaughterers and Meat Packers
Stonemasons
Tapers
Tire Repairers and Changers
Wellhead Pumpers
It seems like masses of people locked out of any economic growth and with no perspective are a recipe for desaster if you ask me. That's the foundation on which unrest and rebellions are built and those are usually not pleasant affairs for most people.
I'm not saying that it will come to this.
* Athletes and Sports Competitors * Motorcycle Mechanics * Stonemasons
only to give some examples, where roboters are almost impossible.
But if we really get to the point where programmers can be replaced with LLMs, I expect something closer to swiss cheese: For example, if you're a meat packer, maybe you work someplace that doesn't want to pay for a fancy new machine that can do your job...
But a massive plummeting in the cost of software paired with a massive spike in the rate at which we develop things means machines will cost a lot less. Who knows what kind of material sciences advancements we'll see for example.
At that point the margins may be razor thin, but automation will start to work in places where it couldn't be done profitably before. I think things get a little dystopian from there, since I'd expect the employers still holding onto humans to essentially be holding all the cards. What good is a union when the "scab" is a machine with a one time cost?
-
(but again, this is all assuming LLMs reach a place that is seriously earthshattering)
The truth is that we're impressed by this because it's magic and we don't know how magic - but it still getting way more magic really fucking fast.
Right now it's still and idiot if you're not telling it what to do. Hypnotic babble.
(If you want to be scared or relieved feed it some time series stock data (if it's good it'll end up fighting itself, otherwise it's either stupid or a liar).)
There are some uses easy use cases for it - I really wouldn't mind a GPT bot replacing outsourced call centres.
Improvement of quality and an improvement in service, and no direct impact to the local economy.
Other than that, this all has the booming feel of every other damned tech buzz, can we just cool down on sensationalism and not try and imagine another multi-billion dollar sub-economy into existence?
We already did that with bitcoin, and now we're seeing the end of the magic trick; let's not do it again.
How about a calm objective and public analysis by under-excited experts before we start mining land that simply isn't there?