ChatGPT-4 significantly increased performance of business consultants
d3.harvard.edu
d3.harvard.edu
"Participants responded to a total of 18 tasks (or as many as they could within the given time frame). These tasks spanned various domains. Specifically, they can be categorized into four types: creativity (e.g., “Propose at least 10 ideas for a new shoe targeting an underserved market or sport.”), analytical thinking (e.g., “Segment the footwear industry market based on users.”), writing proficiency (e.g., “Draft a press release marketing copy for your product.”), and persuasiveness (e.g., “Pen an inspirational memo to employees detailing why your product would outshine competitors.”)."
Here is the GPT response to the first task: https://chat.openai.com/share/db7556f7-6036-4b3d-a61a-9cd253...
A confident GPT hallucination is almost indistinguishable from typical management consulting material...
2) Spreadsheets exist.
3) No-one cares about your marketing copy.
4) No-one finds your c-suite babble inspirational.
This is almost perfect input to an LLM exactly because of how low value it is in the first place.
Somehow they're able to make the C suite hoover up LLM shovelware the same way top producers can take super obvious music and sell I V vi IV but when we try the same chords it's uninspired and no one wants to listen to it
If it was so bad, then why do people listen?
There is still a market.
Does the market suck? Full of idiots?
Your argument ends up being that successful things are bad, because humans are just idiots and thus if something is successful it is because it is just liked by idiots.
As much as I might agree generally, it doesn't get you far.
I mean I'm not sure that this is that far from the truth in some domains, though it depends on how you define idiocy. There is, for example, a market for demolition derbies. Of course all of us are idiots in some ways so we should be careful about whom we disparage.
A) Forming cross company cliques (a lot of C suite is ex consultant and they scratch each others' backs).
B) ego stroking and typical sales ("this executive is a visionary who must be furnished with top quality steak and strippers")
C) Letting you know on the sly what their other customers are doing that seems to be working.
D) Providing industrial grade ass cover for decisions that the C suite want to make but are afraid to make by themselves (like layoffs).
I recon 60% of management consulting work is just to ass cover for a director with no conviction
I don't doubt that if the high level decision agreed upon is "more layoffs because AI" and they were asked for a 60 page report to justify it that ChatGPT would help inordinately in fleshing it out with something that sounds fairly plausible.
“I paid a world-class consulting consultant company top dollar to vet this idea and they produced ton of documents about how great it was. And, yet it failed. But, I’m not at fault here. What more would you have had me do?”
There was an article on HN years ago about top grade from Harvard-like schools being sucked into consulting companies and discovering their job was to be paid tons of money writing reports that support whatever the exec of the moment wanted to hear.
Sounds like an extremely nice job.
It's probably easier to assume that their job is to provide objective expert advice since thats what they say they do.
I'm being realistic here, not cynical.
I can assure you I wasn't.
You are mistaken "expertise" with fashion and a good voice.
When Scorpions wanted to be resurrected they hired Desmond Child to produce them and he absolutely crushed. These people are very good at what they do and there are very few of them.
i.e. someone has been tasked with getting some consultants to come up with a suggestion, the key question however is what does THEIR boss want to hear. If you can work that out and give it to them then you've earned your 'worth'
Now I know there are people here whose reaction is that executive management should just shut up and listen to what the worker bees say. But it actually doesn't seem unreasonable to me (and I've been on the product management end of things in a case like this) to have some outside perspective from some people who are mostly pretty smart to nod their heads and say this seems sensible.
As a bonus they create big spreadsheets that make the business planning types happy and keep them out from underfoot.
Until then, our leaders, experts and institutions were few of those things during the pandemic.
How large businesses are built and run has changed faster and more in the past few years than the principled predictions of a business’ future vector that are based on lagging indicators.
It also depends on how cookie cut the management consultant frameworks and “toolkits” are.
It’s no coincidence that it’s mostly juniors doing so much of the work and billing. New or average talent is more profitable per hour to bill than experience.
Financially, if $7-8 of every $10 for improvement went to a management consulting undertaking, the other $2-3 is what’s left over for the rest without even knowing. This would be the coup, if this were true.
The fun part to watch for is is tech people will be able to learn business easier than business people will be able to learn and apply tech when they can’t understand it’s capabilities or possibilities beyond speaking points.
The technical analyst will M&A the business analyst. Maybe they learn to extract Management Consultant type value too.
Keep their existence completely out of the picture, and have them scout and produce talented no-name, but require the no-name to use only the sort of avenues that would be openly available to anybody/everybody: YouTube, Tunecore, social media, etc. Would the new party now be meaningfully likely to have a real breakthrough?
The ‘quality’ of a pop song is how much popular appeal it has. That’s the basis of the genre, even reflected in the name.
Occasionally you'll have a song that breaks through due to sheer catchy-ness, but this is the exception rather than the rule.
Something can be extremely catchy yet widely panned as low quality in music, so even within "just the music" there are several dimensions at play regardless of marketing, etc. Such as whether it's timed right - are there enough people ready for that song at that time?
The idea that "most people will just listen and be fans of whatever the big media companies put out there" doesn't stand up much examination or conversation with "most people."
People do often make breakthroughs on soundcloud, TikTok, whatever - do you think having the invisible support of a Max Martin would lower their chances? You'd need to do your experiment a hundred times or thousand times or so before you could really compare the success rate of your plants to the rest of the crowd, but it's hard for me to believe that they wouldn't have an advantage. The music industry isn't known for their charity, if they could get away with not paying those people without another label beating them in the market, why would they?
This is admittedly very niche but I see your point.
Also: I remember a time in the early 00s when almost every song on MTV began with Rodney Jerkins whispering “DARKCHILD” over the music.
Ultimately isn’t it just branding though? Would you buy Coca-Cola if it had some other label on the bottle? Or watch Mission Impossible 14 starring Some Dude? I’m not sure there’s a lot of fields where things are really competing on their own merits rather than the accumulation of their past successes.
Depending you ask pop music is a formula and that one dude from Sweden has the formula perfected.
The producers hired top shelf songwriters to write the songs, and several hit albums were produced. (It really is good music, despite being bubblegum.)
Cassidy, however, decided that he had songwriting talent and chafed. He eventually left the show, and with the megabucks he earned on the show, produced albums. They're terrible.
The same thing happened to the Monkees.
Then, the bands belatedly had to learn how to sing and play in order to go on tour.
I think about The Wrecking Crew whenever I hear the sob stories about bands being underpaid and the producers reaping the lion's share of the profits.
On the songwriting side, there's the lyrics which have both a phonetic and a semantic component. There's also the fact that many people will mishear the lyrics and their evaluation of them will be based on the mishearing. There's the melody. Does it work together with the chords to highlight the key parts of the lyrics?
Then there's the performance where there are a million ways to stand out or flop. Loudness, timbre, timing and even detuning can all be used for expression.
Actually a robot would be perfect at that. No one likes doing layoffs but chat gpt won't mind
Maybe chatgpt should be trained to care.
If you're measuring based on output, sure, but... the value of any knowledge worker is primarily driven by the input, that is, a client doesn't want "10 ideas" they want "10 [valuable] ideas [informed by the understanding of the business and the market they're operating in]". If a management consultant said "boat shoes" in response to this question they would not have a client much longer.
You could apply this same nonsense task to software engineering, i.e: ask ChatGPT to "write 10 lines of code" and it'll be indistinguishable from the code we churn out day after day.
> Features: Built-in waste bag dispenser
I'm not yet sure whether I hate it or love it.
> Target: Visually Impaired Individuals
> Features: Haptic feedback
Haptic shoes, how revolutionary!
The shoes seem to give you two options for cleaning up your dogs poop: (1) Bend down twice, once for the bag, once for the poop. (2) Bend down once nice and close to the poop and get a bigger whiff than otherwise.
I’m just not looking for ways to interact manually with my shoes more than I have to…
Sounds perfect for both Harvard and those linked to the institution.
My employer has hired McKinsey a few times, known to recruit from HYP, and their output has been subpar to say the least. My entire experience with these institutions has been fairly uniform in that regard.
I know it’s anecdotal. But it feels like there’s a lot of confirmation bias with these sorts of studies.
The first task was a generalist task ("inside the frontier" as they refer to it), which I'm not surprised has improved performance, as it purposely made to fall into an LLM's areas of strength: research into well-defined areas where you might not have strong domain knowledge. This also is the mainstay of early consultants' work, in which they are generalists in their early careers – usually as business analysts or similar – until they become more valuable and specialise later on.
LLMs are strong in this area of general research because they have generalised a lot of information. But this generalisation is also its weakness. A good way to think about it is it's like a journalist of research. If you've ever read a newspaper, you often think you're getting a lot of insight. However, as soon as you read an article on an area of your specialisation, you realise they've made many flaws with the analysis; they don't understand your subject anywhere near the level you would.
The second task (outside the frontier) required analysis of a spreadsheet, interviews and a more deeply analytical take with evidence to back it up. These are all tasks that LLMs aren't strong at currently. Unsurprisingly, the non-LLM group scored 84.5%, and between 60% and 70.6% for LLM users.
The takeaway should be that LLMs are great for generalised research but less good for specialist analytical tasks.
When I ask a programming question, chat GPT hallucinates something about 20% of the time and I can only tell because I’m skilled enough to see it. For all the other domains I ask it questions if I should assume at least as much hallucination and incorrect information.
However as the paper noted, when working within AIs areas of strength it improved not only efficiency but the quality of the work as well (accounting for the hallucinations). As you mentioned:
> When I ask a programming question, chat GPT hallucinates something about 20% of the time and I can only tell because I’m skilled enough to see it
This matches their Centaur approach, delineating between AI and one’s own skills for a task which—with generalized work—seems to fair better than not using AI at all.
Sometimes to be taken seriously at work, you need to take some concise idea or data and fluff it up into a multiple pages or a slide deck JUST so that others can immediately see how much work you put in.
The ideal role for chatgpt at this moment is probably to take concise writings and to expand it into something way larger and full of filler. On the receiving end, people will endure your long-winded document or slide deck, recognize you "put in the work", and then feed it back into chatGPT to get the original key points summarized.
Yeah. Most people have focused on what LLMs can do, but I think it’s equally if not more interesting what can they not do, and why?
When we say LLMs can generate text we’re painting brush strokes as broad as a 10-lane highway. Apparently we have quite limited vocabulary about what writing actually is, and specifically what categories and levels exist.
For instance, it’s fun (and in my view completely expected) to see that courteous emails, LinkedIn inspirational spam, corp-speech etc, GPT outperforms humans with flying colors, on the first attempt too! Whereas if you’re asking for the next book of Game of Thrones or any well-written literature it falls flat – incredibly boring, generic, full of platitudes and empty arcs and characters.
We have to start mapping the field of writing to a better conceptual space. Currently it seems like we can’t even differentiate between the equivalent of arithmetic and abstract algebra.
And where there isn’t a bright line around “fact”, and where it doesn’t need to come together like a Pynchon novel, the generative stuff is smoking hot: short-form fiction, opinion pieces, product copy? Massive productivity booster, you can prototype 20 ideas in one minute.
But that’s about where we are: lift natural language into a latent space with some clear notion of separability, do some affine (ish) transformations, lower back down.
Fucking impressive for a computer. But if it can really carry water for an expensive Penn grad?
You’re paying for something other than blindingly insightful product strategy.
LLMs iare most often best at helping humans do their tasks more effectively, not replacing them completely
https://www.reuters.com/legal/new-york-lawyers-sanctioned-us...
It's gonna be interesting.
The problem is passing tests are an okay proxy for competence in humans, but if you think of LLMs as a giant library search engine, the thing it is competent at is identifying and regurgitating compiled phrases from its records.
Which is awesome. It can't be a doctor.
And it’s dope that it can do that!
But let’s keep our heads about what it is.
(but then tells you not to use it as it is "unethical")
I actually don't know what you thought it meant.
For others arriving here: I suspect OP meant Natural Language Processing and I was talking about Neuro-linguistic Programming.
I've had my caffeine now.
But they can do plenty of useful stuff reliably. It’s not “be generally intelligent”, which they are just nothing even remotely close to, but know you don’t dig the LLM hype from that comment? Yeah, they get that every time.
Sentiment analysis is easy. It wasn't the end goal, just a "hello world" example to verify my tools were set up correctly. I ran into unsolvable problems in the tutorial.
I have no use for tools which do amazing things sometimes but which cannot be reasoned about and cannot be prevented from producing garbage. Maybe other people will find uses for them, though. I'll keep an open mind and check back in five years.
I'm just saying, we invented backspace for a reason. LLMs have no backspace. It's insane they work as well as they do.
If you’re not convinced about sentiment analysis on e.g. LLaMA 2, I think you’re wrong, but maybe I’m wrong.
If you’re up for it, let’s run an expedient, I’ve got a GPU or two in my living room. This thread seems like a pretty great test set actually.
Maybe we both learn something richer than some benchmark stat?
I note that the comment is [dead] not [flagged] [dead], so maybe its state has to do with something else than the content of the comment? Just [dead] is, I think, shadowban.
I checked the poster's comments, but since it's a new account there's very few of them and I can't determine the reason for the [dead] from them.
For this to be true for most production service use cases, LLMs would need to be at least ~10X faster. I generally agree they can be quite good at these tasks, but the performance is not there to do them on large datasets.
Consider using your PRs and docs to capture the answers to the usual why questions which LLM won't be able to do.
I will admit that a lot of the really old decisions don't have much relevance to the current business, but the historical insight is sometimes nice.
People who have no background in writing or editing think LLMs will revolutionize those fields. Actual writers and editors take one look at LLM output and can see it’s basically valueless because the time taken to fix it would be equivalent to the time taken to write it in the first place.
Similarly people who are poor programmers or have only a surface level understanding of a topic (especially management types who are trying to appear technical) look at LLM output and think it’s ready to ship but good programmers recognize that the output is broken in so many ways large and small that it’s not worth the time it would take to fix compared to just writing from scratch.
they aren’t good at using language models.
You could say similar things about Stack Overflow, and yet we use it.
I find treating it like an intern is amazing productive: https://simonwillison.net/2023/Sep/29/llms-podcast/#code-int...
And for text I know people who use it succesfully (professionally) to generate texts for them as a summary from some data. They still have to proof read, but it saves them time, so it is valuable.
They can be worse than worthless. They can sabotage your work if you let them making you spend even more time fixing it afterwards.
For an example. I've used Gpt4 as a sort of Google on steroids with prompts like "do subnets in gcloud span azs" and ", "in gcloud secret manager can you access secrets across regions". I very quickly learned to ask "is it true" after every answer and to never rely on a given answer too much(verify it quickly, don't let misinformation get you too far down the wrong route). So is it useful? Yes, but can it lead you down the wrong path? It very well can. The least experience you have in the field the easier it will happen.
>You just cannot expect it to ship a full programm for you, but for generating functions with limited scope, I found it very useful
Entire functions? Wow. I found it useful for generating skeletons I then have to fill by hand or tweak. I don't think I ever got anything out of Gpt4 that is useful as is (maybe except short snippets 3 lines long).
However, I found it extremely useful in parsing emails received from people or writing nice sounding replies. For that it is really good (in English).
I basically gave up on llms because i was spending more time figuring out what it did wrong than actually getting value.
People without programming skill are still impressed by them. But they yet have to learn or deliver anything of value even with the help of chat bots.
Here's the transcript (it pre-dates the ChatGPT share feature): https://gist.github.com/simonw/1aa4050f3f7d92b048ae414a40cdd...
I wrote more about it here: https://simonwillison.net/2023/Apr/15/sqlite-history/
Here's another one I built using AppleScript: https://github.com/dogsheep/apple-notes-to-sqlite - I wrote about that here: https://til.simonwillison.net/gpt3/chatgpt-applescript
Also: not every system has to be a scalable system. That's another lesson junior engineers (should) learn.
LLMs do produce impressive code. Even if they were indeed just procedural generators it would still be impressive. The code has structure and appears useful.
But the issue is that you can tell it makes no sense, there is no thought process behind it. It fits in no greater picture.
Even if you add more context it still has no purpose.
People that find this useful are the same type that copy stackoverflow code that they dont understand. It kinda works when it does but again it doesnt fit in the bigger picture.
Code isnt about spelling instructions - an…ai can do that - code is about what goes where in a way that the what changes as often as the where. It’s the bigger picture. So yes it can help and replace those that spell instructions but it will be hard to replace those that are required to deliver more.
Completely agree with you. That's my job. The LLM is effectively my typing assistant.
But that is the same, when you blindly follow some stackoverflow answer.
And yes, I always have to tweak and I use it only rarely. But when I did, it was faster than googling and parsing the results.
The amount of fighting I needed against MS development tools mingling my code recently is absurd. (Also, who the fuck decided that autocomplete on space and enter was a reasonable thing? Was that person high?)
Even this can be a big time saver, that increases productivity.
Just like others have said, it isn't going to write a Pynchon novel, but it does do a great job at the other 99% of general writing that is done.
Same for computers, the average programmer isn't creating some new Dijkstra Algorithm every day, they are really just cranking out connecting things together and doing the equivalent of 'generic boiler plate'.
I've had mixed experiences with getting it to generate new code. It produced good node.js command line application code. It didn't do so well at writing a program that creates 16 bit PCM audio file. I asked it to explain the WAV file format and things like lengths of structures got so confusing I had to research the stuff to figure out the truth.
It's been hit or miss with rust. It's super helpful in decrypting compilation errors, decent with "core rust" and less helpful with 3rd party libraries like the cursive TUI crate
Which comes as no surprise, really, as there's certainly less training data on the cursive crate than, say, expressjs
Also FWIW I have actually pointed it at entire git repos with the WebPilot plugin within ChatGPT and it could explain what the repo did, but getting it to actually incorporate the source files as it wrote new code didn't work quite so well (I pointed it to https://github.com/kean/Get and it would frequently fall back to writing native Swift code for HTTP requests instead of using the library)
For graphics tasks GenAI is absurdly helpful for me. I can code but I can’t draw. Getting icons and logos without having to pay a designer is great.
This is true for media articles but for LLMs I feel like it's the opposite. Like people who aren't specialists don't fully appreciate how great it is at those tasks.
"It says more about [insert]" anytime GPT does something just makes the phrase lose all meaning. Surely you have something meaningful to say?
I agree with you in an ideal world, but sadly this isn’t one.
Number 1. In a team of 20-30 engineers there is only one extremely god "why is he with us" engineers who is great at technical stuff and being a people person. However, no matter how nice he is his approach to his job, it is a job and I will only drop hints how the management should be done. He doesn't care about where the company is headed because he plays video games, has a family and has a literal life. He doesn't care about management and taking on undue responsibilities. Moreover, the people up to has a label for him as an "engineer" does not see as a "manager".
For the rest of the engineers and managers, have also adopted the approach of "not my problem", you see a bizarre communication gap. Engineers working closesly with the product don't want to talk to their managers, becase the conversation goes like "if you know this so much, why don't you.... <a description of something results in more work that goes outside their JD>" and managers don't want to talk with engineers because "if you are you so interested, why don't you.... <a description of something results in more work that goes outside their JD>"
From this progressive distance between managers and engineers comes the "manaegment consultant". Management consultant have the upper management given flexibility of going back and forth between engineers and managers. They can have conversations with full flexibility but they are not bound to "why don't you...." phrases. They can talk with anyone and submit a report and take home 1 years worth of salary of managers/engineers in 1 month.
The conversation gap between product and business where management consultants come in. And the funny thing is that, management consultants target those "I don't want to but I should" work things and report to the upper management. They can do this so well, because they are not burdened with the "work" part.
Seriously, if you do some introspection, you will see there is plenty of things you know your company should do, but you don't want to voice them because it results in more work and in fact more risk. There comes a "good" management consultant who will discover those things and report to upper management who will create the system to get those jobs done.
That is my pitch if anyone wants a management consultant hire me. I am going to tell them why their company sucks in 20 different ways with 18 of those points being generated by ChatGPT.
If all else fails, the LLM revolution will at least allow us to make sense of ketamine-induced rants on management.
What does it say about me if I didn’t think that it was that bad?
My theory is that honest takes should be written on first take without revisions and without edits. The moment I massage a statement to be more coherent I am compromising on my honesty.
But no, it was a good post and the cultural expectation to keep things shorter and more buttoned up has some real downsides.
I would have written this with more punch, but, well, see above.
[1] https://www.writersdigest.com/be-inspired/did-hemingway-say-...
Either way, in my experience management consultants just add new useless tasks for everybody on that set. I have never seen them actually decreasing the number of tasks.
I often ask GPT4 to write code for something, and try if it works, but I seldom copy and paste the code it writes - I rewrite it myself to fit into the context of the codebase. But it saves me a lot of time when I am unsure about how to do something.
Other times I don't like the suggestion at all, but that's useful as well, as it often clarifies the problem space in my head.
Not to say you're of of those programmers, but it certainly enables those sorts of programmers.
Also, it will have to be scrapped when anyone wants to tweak it a little. Sure monkeys randomly typing on a typewriter will eventually write the greatest novel in existence... but most of it will be shit
May $entity have mercy on your soul if the business starts bleeding tons of money due to an issue with the code, because the codebase won't
Some example found page 10 of the original article:
- Propose at least 10 ideas for a new shoe targeting an underserved market or sport.
- Segment the footwear industry market based on users.
- Draft a press release marketing copy for your product.
- Pen an inspirational memo to employees detailing why your product would outshine competitors.
Nothing of real value imho.Without the right target market, business model, and effective methods to reach customers, the most brilliant pair of shoes or piece of code can be useless (unless someone works to repurpose them as art or a teaching tool).
It's simlar to what happens to people who knows a language (not coding language), stop using it or go back to use translator, and when they need to use it themselves, they are unable.
I gave it a nontrivial task I couldn’t google a solution for, and wasn’t sure it was even possible:
Given a python object, give me a list of functions that received this object as an argument. I cannot modify the existing code, only how the object is structured.
It gave me a few ideas that didn’t quite work (e.g modifying the functions or wrapping them in decorators, looking at the current stack trace to find such functions) and after some back and forth it came up with hijacking the python tracer to achieve this. And it actually worked.
The crazy thing is that I don’t believe it encountered anything like this in its training set, it was able to put pieces together which is near human level. When asked, it easily explained the shortcomings of this solution (e.g interfering with the debugger).
I have seen similar things. So, no, it's not regurgitating from its training data-set. The NN has some capacity for reasoning. That capacity is necessarily limited given that it's feed-forward only and computing is still expensive. But it doesn't take much imagination to see where things are going.
I'm an atheist, but I have this feeling we will need to start believing in "And [man] shall rule over the fish of the sea and over the fowl of the heaven and over the animals and over all the earth and over all the creeping things that creep upon the earth"[1] more than we believe in merit as the measuring stick of social justice, if we were to apply that stick to non-human things.
[^1]: Genesis 1:26-27, Torah
If you’re a client and need a consultant to do something, you have to explain the requirement to them, review the work, give feedback, and so forth. There will likely be a few meetings in there.
But if GPT-4 can make consultants so much better, I imagine it can also do their work for them. And if you combine this with the reduction in communications overhead that comes from not working with an outside group, why wouldn’t clients just accrue all the benefits to themselves, plus the benefit of not paying outside consultants or dealing with the overhead of managing them?
This is especially the case when the client is already a domain expert but just needs some additional horsepower. For example, marketing brand managers may work with marketing consultants even though they know their products and marketing very well. They just need more resources, which can come in the form of consultants for reasons such as internal head-count restrictions.
Anyway, I just wonder if BCG thought through the implications of participating in this study. To me it feels like a very short step from “helps consultants help their clients” to “helps clients directly and shows consultants aren’t really necessary.”
Especially so if the client just hires an intern and gives them GPT-4.
People have ideas all the time internally. I'm going to assume the idea you had was one of many.
The issue is getting the real decision makers to buy into it. They aren't going to take the word of someone who works in some division. They want some rigor to it.
Bringing in someone who isn't tainted by the groupthink of the company, can actually take a sober view of the situation, has puts some weight to the recommendation.
> I can't help but think the next AI winter is around the corner. [0]
Yeah, right.
If you make a hilariously bad prediction then that tells you your model about that thing is off and needs correcting.
So if you do nothing to that model and still make predictions...
>If we're looking for a cost-effective way to replace content marketing spam... great! We've succeeded!
And if you read the article that’s almost exactly the level of output that we’re talking about.
- Propose at least 10 ideas for a new shoe targeting an underserved market or sport.
- Segment the footwear industry market based on users.
- Draft a press release marketing copy for your product.
- Pen an inspirational memo to employees detailing why your product would outshine competitors.
Also for the 2nd task the non LLM group performed significantly better.In my opinion a lot of people here on hackernews as they are themselves good at programing underestimate how services like chat gpt can open a new world to non programmers. They also probably make the non inquisitive learn less. Previously to learn how to stop multiple snapd services using a script I would have googled and then cobbled together something today I just ask chatgpt and get a working script in less than a min.
> For each one of a set of 18 realistic consulting tasks within the frontier of AI capabilities
They specifically picked tasks that GPT-4 was capable of doing. GPT-4 could not do many tasks, so when we say that performance was significantly increased this is only for tasks GPT-4 is well suited to. There is still value here but let's put these results into context.
> Consultants across the skills distribution benefited significantly from having AI augmentation, with those below the average performance threshold increasing by 43% and those above increasing by 17% compared to their own scores
Even when cherry-picking tasks that GPT-4 is particularly suited for, above average performers only increased performance by 17%. This increase is still impressive, were it to be seen across the board. But I do think that 17% is a lot less than some people are trying to sell.
but I can't find any automated AI tasks of critical importance, they all need human support
Therefore below-average types will produce finished output more quickly; and this was a time-constrained test, so velocity matters.
ChatGPT is very good at waffling, and marketing-speak and inspirational messages are essentially waffle. IOW, the tasks were tailor-made for unaided ChatGPT, so high-performers were penalized.
It switched from always guessing composite to always guessing prime. Much less accurate.
My questions to naysayers:
* Do you or anyone you know use GPT-4 (not the free GPT-3.5) to do productive tasks like coding and found it to help in many cases?
* If you insist it’s useless, why do millions of people pay $20 a month to access GPT-4 and plugins?
And for the second one, although I am paying for it too, this idea is more or less flawed nowadays. Utilization is a very hand wavy thing when it comes to this stuff. Like a purse, millions would pay money for it, some even pay thousands. But I have no use for it and wouldn’t even pay a $1 for one.
Agreed.
> Like a purse, millions would pay money for it, some even pay thousands.
Expensive purses have intangible value for some. They are often bought to signal social status.
I'm pretty sure a significant portion of ChatGPT Plus subscribers are paying because it can help them with information or cognitive work that some people value.
The second question doesn't make sense to me. There are tons of things I think are useless (or worse) that people pay for anyway. Meal kit boxes come to mind, and at least you can eat those at the end of the day.
Getting great results out of it takes a lot of experimentation and practice. I wouldn't want to give it up now I've learned how to use it.
I need to see a glimmer of it being useful before I decide the investment is worth it, I guess.
1: I’ve been scripting for 5 years using Python. I purchased a subscription to use GPT4 to see if it could assist me.
In the end it took me more time to fix its mistakes than to just apply my knowledge of knowing what to Google and reading docs.
Additionally the largest hurdle I encountered was when it hallucinated a package that didn’t exist and I spent time trying to find it.
2: I don’t know about most people but I’m terrible at cancelling services that are “cheap”. I used ChatGPT for a few hours that first month and didn’t cancel it for another 5 months.
[0] https://www.marktechpost.com/2023/06/16/this-paper-tests-cha...
https://github.com/DLR-SC/JokeGPT-WASSA23/blob/main/01_joke_...
That's a bad way to use an LLM for joke generation.
Try "tell me a joke about a sea lion" - then replace sea lion with any other animal.
Or "tell me ten jokes about a lawyer on the moon" - combine concepts like that and you get an infinite variety of jokes.
Some of them might even be funny!
Who cares if it gets things wrong sometimes, you would double check your co-workers answers also. And there are times when I insist I am correct, and GPT will argue back and eventually I find I was wrong.
I suppose this because I recall how much search improved my productivity over flipping through books and I know how for certain tasks ChatGPT is a better source of knowledge on how to do it than search. While often the GPT output isn’t entirely correct, more often than not it suffices to make the correct solution obvious thus saving a lot of time.
Summary: https://pdf2gpt.com/?summary=84ff84d4b98b4f0c985a17d07db482c...
Purpose of technology is to enhance our performance, GPT is very much doing so - but with great powers comes great responsibility.
I don’t blaim BCG for doing this, they are giving an outside view and political uninfluenced (except for the party that pays the tap) view.
A characteristic of these professions is that there is no accountability for output they produce. It is not like a profession that builds an engine for a car. They can bullshit with confidence and get away with it.
chatGPT will replace all of them - as chatGPT itself can bullshit with the best of them.
Also access to AI significantly increased (!) incorrect answers in the case where the tasks were outside of AI capabilities.
Absolutely zero add value in experience. The only add-value is the consultant overcoming the hearing deficiency of the Director involved.
ConsulatancyGPT: Feed all internal opinions of a company into an LLM. Ask for the a recommendation. Done.
/rant.
D'oh.
If your job is to generate nonsense… well…
I tried giving it the url and it was a disaster. Is there a plug-in?
Sounds like the same text in the average deck...
Oh right, it's not that type of efficiency :)
> while AI can actually decrease performance when used for work outside of the frontier
There is some value here but the authors can define "frontier" however they please to end up with whatever productivity increase they are looking for.
It was nice of them to explain that the article was total horse dung before having to read the whole paper
https://m.youtube.com/watch?v=6pieIoEi8Ds
I love it when a complicated set of conditions can be losslessly encoded in a short, easily remembered label.
>integrated their workflow with the AI
What the hell is even the difference? Was the article itself written by ChatGPT?
These are existing industry terms, 'Centaurs', 'Cyborgs', 'Unicorns', etc... have been around for decades.
Technically, that one sentence you are calling 'BS', actually did describe the situation accurately. And got the message across about the issues being discussed in the article.
Almost like a business consultant, used business terms, to succinctly describe the subject that will be discussed. To quickly get the main points across.
"How do I do x in language y" always gives me the knowledge I need. Within seconds, I can continue coding.
After more than 10 years of coding fulltime, I know some languages very well, like PHP and Javascript. But even in those, LLMs often come up with a better solution than what I wrote. Because they know every fricking thing about those languages.