Field experimental evidence of AI on knowledge worker productivity and quality
oneusefulthing.org
oneusefulthing.org
https://www.oneusefulthing.org/p/centaurs-and-cyborgs-on-the...
There are two sides to this thing. The first author of this paper has also written another paper titles "Falling Asleep at the Wheel: Human/AI Collaboration in a Field Experiment". Abstract:
"As AI quality increases, humans have fewer incentives to exert effort and remain attentive, allowing the AI to substitute, rather than augment their performance... I found that subjects with higher quality AI were less accurate in their assessments of job applications than subjects with lower quality AI. On average, recruiters receiving lower quality AI exerted more effort and spent more time evaluating the resumes, and were less likely to automatically select the AI-recommended candidate. The recruiters collaborating with low-quality AI learned to interact better with their assigned AI and improved their performance. Crucially, these effects were driven by more experienced recruiters. Overall, the results show that maximizing human/AI performance may require lower quality AI, depending on the effort, learning, and skillset of the humans involved."
https://static1.squarespace.com/static/604b23e38c22a96e9c788...
I think you also need to consider whether you can afford not to have the skill; what would happen if the AI were taken away, or malfunctioned? Airplane pilots are an example of this. In that case you simply have to learn the skill as if the AI will not be there.
The jury is still out on how "temporary" of a problem that will be. Improvements have been made but nobody really knows how to get an LLM, for example, to just say "I don't know".
Oooh, I really like that. Very succinct and topical statement of Goodhart's Law.
Obviously there jobs out there that do have to worry about these things (mostly at companies like TSMC and Cisco), but in practice I doubt regular engineers at even the largest scale software companies have to worry about stuff like that.
Say hello to undefined behavior :)
Appropriately enough, much of the reason undefined behavior is a consistent problem is exactly the same—it can do something reasonable 99% of the time and then totally screw you over out of nowhere.
I used to be a more conventional engineering guy. We do a ton of calculus to get our degree. My first job - you just didn't need calculus (although all the associated courses in school used it to develop the theory). As a result, the one offs (less than once a year), when a problem did come up where calculus knowledge was helpful, they couldn't do the task and would go to that one guy (me) who still remembered calculus.
BTW, I guarantee all the employees there got A's on their calculus courses.
Not at all, and your argument isn't the same thing.
The problem with current AI, which is very different from past forms of automation, is that, essentially, its failure mode is undefined. Past forms of automation, for example, did not "hallucinate" seemingly correct but false answers. Even taking your own example, it was obvious to your colleagues that they didn't know how to get the answer. So they did the correct thing - they went to someone who knew how to do it. What they did not do, and which is what most current AI solutions will do, is just "wing it" with a solution that is wrong but looks correct to those without more expertise.
- C programmer
Undefined behavior in C is actually strictly defined: if you want to avoid undefined behavior, don't do stuff the spec specifically says results in undefined behavior, like deference a null pointer.
Writing up a set of rules to guard against undefined behavior (which is a primary job of a C developer) is simply not possible with an LLM.
This isn't new, like since LLMs were born, either. In the aviation community, and probably many other domains, there has been a running debate for decades about the pros and cons of introducing more automation into the cockpit and what still should or shouldn't be taught and practiced in training (to be able to effectively recognize and handle the edge cases where things fail).
And in military aviation, this isn't just a question of safety; it's about productivity just like it is for knowledge workers. If an aircraft can fly itself most of the time, the crew can do more tactical decision-making and use more sensors and weapons. Instead of just flying their own aircraft, the crew can also control a number of unmanned aircraft with which they're teamed up.
There's a brief example of a type of task that each approach excels at, but I'd definitely like more. The crux of this piece is that our human judgment of when/how to use AI tools is the most important factor in work quality when we're on the edge of the frontier. But there's no analysis of whether centaurs or cyborgs did better at those outside-the-frontier tasks; it's not clear why those categories are even mentioned since they appear to have no relevance to the preceding research results. And as the article mentions, the frontier is "invisible" (or at least difficult to see); learning to detect tasks that are in, on, or outside of the frontier seems like an immensely important skill. (I also realize it may become completely obsolete in <5 years as the frontier expands exponentially.)
I understand the goal of this research was not to find these "edges" and to determine how we can improve our judgment about when & how to use AI. But after reading these strong results, that's definitely the only thing I'm interested in. I use AI almost not at all in my work, web development. It has been most useful to me in getting me unstuck from a thorny under-documented problem. Over a year after ChatGPT released, I still don't know if modern LLMs can actually be a force multiplier for my work or if I'm correctly judging that they're not appropriate for the majority of my tasks. The latter seems increasingly unlikely as capabilities advance.
(In cases like this, it's nice to have both the readable summary and the link to the paper.)
The discussions I hope are happening is the prospect of intentional capability limits of AI in critical industries/areas where humans must absolutely and always intervene. It seems analogous to having a superpower without knowing how to control it, creating a potentially deadly situation at best.
Then perhaps more long term at the generational scale, will humans cognitively devolve to a point where AI is making all of our decisions for us? Essentially we make ourselves obsolete and eventually we’re reclassified as “regular/inferior” species of animals. Profit and return on investments here cannot get in the way of handling these details with great oversight and responsibly.
It either needs to be 100% or (bad).
Anywhere in the grey zone, and you retard human performance while still occasionally needing it.
[1] https://ckrybus.com/static/papers/Bainbridge_1983_Automatica...
Some were more creative in nature, and some were more analytical. For example, ideas on marketing sneakers versus an analysis of store performance across a retailer's portfolio.
In general they found that GPT4 was helpful for creative tasks, but didn't help much (and in fact reduced quality) for analytical questions.
I think these kinds of studies are of limited use. I don't believe raw GPT4 is that helpful in the enterprise. Whether it is useful or not comes down to whether engineers can harness it within a pipeline of content and tools.
For example, when engineers create a system to summarize issues a customer has had from the CRM, that can help a customer service person be more informed. Structured semantic search on a knowledge base can help them also find the right solution to common customer problems.
McKinsey made a retrieval augmented generation system that searched all their research reports and created summaries of content they could use, so a consultant could quickly find prior work that would be relevant on a client project. If they built that correctly, I imagine that is pretty useful.
GPT4 alone will especially not be that useful for analytical work. However, developers can make it a semi-capable analyst for some use cases by connecting it to data lakes, describing schemas, and giving it tools to do analysis. Usually this is not a generalist solution, and needs to be built for each company's application.
Many of the studies so far only look at vanilla GPT4 via ChatGPT, and it seems unlikely that, if LLMs do transform the workplace, that a standalone ChatGPT is what it will look like.
Unfortunately not so true. A simple way to check this is to look at what the brain is doing. Coding barely triggers the language components of the brain (your usual suspects like the Vernicke's and Broca's areas + other miscelleneous things). In fact what you'll find is that coding triggers the IP1, IP2 and AVI pathways. These are part of the multiple demand network. We call them programming "languages" because we can represent them in a token-based game we call "language" but the act of programming is far from mere language manipulation.
This makes sense. Translation to precise code would include more of the logic/planning bits of the brain, since that's the context of the output. I suspect a person fluent in Lojban[1] would show similar areas activating during translation from English to Lojban. I also assume translating code to words, with documentation, would include more language, since that's the context of the result.
That doesn't mean it's not translation. It's just translation to a precise description of intent.
I've been thinking about this a lot lately. I'm a data analyst and all these "give GPT you data warehouse schema, let it generate SQL to answer user's queries" products completely miss the point. An analyst has value as a curator of organizational knowledge, not translator from "business" to SQL. Things like knowing that when a product manager asks us for revenue/gmv, we exclude canceled orders, but include purchases made with bonus currency or promo codes.
Things like this are not documented and are decided during meetings and in Slack chats. So my idea is that in order to make LLMs truly useful, we'll create "hybrid" programming languages that are half-written by humans, and the part that's written by LLMs when translating from "human" language is simple enough for LLMs to do it reliably. I even made some weekend-prototypes with pretty interesting results.
I'm not so sure about this. A lot of work in and for especially large organizations is writing text that often doesn't even get read, or is at most just glanced at. LLMs are great at writing useless text.
The effect would be even worse the first time the author missed a mistake from the AI and had to walk it back publicly. Nobody's going to trust anything they put in writing for a while after that.
For a start, most of my work is about using internal APIs. But nevermind. Sometimes I come across a generic programming problem. ChatGPT get me to 80% of a solution.
Getting it from there to a 100% solution takes as long as doing it from scratch.
Just my $0.02
Some people could be not able to figure out that the code GPT produced is not good because they lack the skill level to review it efficiently.
Conversely, could you share an example of a nontrivial, practical programming problem ChatGPT can solve, but only if you use it right?
I'm a senior dev at a python shop and I almost never find the solutions GPT4 gives in Python to be very good. However, I've been learning some C# lately in my free time and I find those solutions to be pretty good. I have a feeling I find the C# solutions higher quality, because I'm not experienced enough to see the problems.
I would really like to know what I'm doing wrong (or if everyone else is in the same boat).
Even more common is getting shell/UNIX commands to do what I want. I don't have all the standard UNIX tools (cut, etc) in my head, and I definitely don't keep all the command line options in my head. GPT4 tells me what I need, and it's much quicker to confirm it via the --help or man pages than craft it myself by reading those pages.
Not something I relished spending days on. Instead I spent thirty seconds and our customers are getting a major update shipped a hell of a lot faster and I can get back to bigger things.
I had it write me a handful of sticky functions. The thought of using a search engine for all of that brought tears to my eyes.
Just searching simple things anymore is gut wrenching.
So AI makes consultants faster but worse. As a consultant, this should make my expertise even more valuable.
That was just for tasks not suited to GPT-4 (and I assume, LLMs in general).
> "For a task selected to be outside the frontier, however, consultants using AI were 19 percentage points less likely to produce correct solutions compared to those without AI."
For tasks suited to GPT-4, consultants produced 40% higher quality results by leveraging GPT-4.
My take on that is that people who have at least some understanding of how LLMs work will be at a significant advantage, while people who think LLMs "think" will misuse LLMs and get predictably worse results.
I've found that now that I see LLM's less as thinking and more as stochiastic parrot I get a lot less use from them because I give 3-5 turns at most versus 10 turns.
You know, everything in moderation and what not.
The problem for me is it used to be an addition to my energy and confidence, but when you "take the magic out of it" it doesn't help out there so much.
In antirez's blog he mentioned how it made things "not worth it, worth it" and I agreed, but that happens much less often if you think an LLM helping you is a fluke.
43% better it made below average people, 17% better it made above average people.
Same paragraph.
Looking into the full paper, it looks like the task identified inside the frontier is much more broken down, with 18 numbered steps. The outside the frontier task is two bullet points.
So for clearly scoped tasks, ChatGPT is hands down good. For less well scoped work it reduces quality.
Interestingly, the authors also looked at the amount of text retained from ChatGPT, and found that people who received ChatGPT training ended up retaining a larger portion of the ChatGPT output in their final answers and performed better in the quality assessment.
Consultant logic at its finest.
Usually the goal is to produce less garbage outputs slower.
I will say GPT-4 is mostly better than referencing stack-overflow or library documentation. Maybe once that gets fast/cheap enough to use as co-pilot we'll see some of these mythical productivity gains.
The reason being that the source of material for GPT-4 is rotting underneath. As the internet gets ruined by AI garbage, the quality of output of AI will get worse, not better.
edit: This is a modern day tower of Babel event. We had a magnificent resource and then polluted it with garbage and now it's quickly becoming useless.
1) The data scraped from the web is filtered, deduped, ranked, and cleaned in a variety of other ways before being used for training. While quantity is necessary for a model to learn the structure of language, quality is even more important once you have a model that can produce coherent output so a lot of work goes into grooming the data.
2) There have been a bunch of papers and training runs that show synthetic data created specifically for training a model is as good or better than scraped human produced data. The importance of scraped web content is quickly declining because you can now generate infinite higher quality examples using existing trained models. The only relevance it has now is for knowledge about new developments and that is a much easier stream to filter since most of the important stuff comes from official sources and you don't need as many variations since you can just generate your own using one or more examples.
It's really good at spitting out code in common, familiar patterns, but the failure rate is very high for anything a little unusual.
They wake up every morning hell bent on maintaining existing customer accounts, growing existing accounts, and sometimes adding a new account. They maintain a professional demeanor, with just the right kind of smalltalk. Finally, they produce a report that justifies whatever was the decision that the CEO already made before he engaged the consultants
Thats pretty much the whole job
What I do is pretty similar to what I did before in my CTO roles, just a bit more strategy and less execution. Clients come to me with questions and problems and I help them solve them based on my experience. If you've got multiple C levels, middle management and a board, you're in a similarly consulting position also as a full time CTO from what I've seen.
I don't go anywhere _near_ LLMs for this, I'd find it somewhere between disrespectful and unethical. They're paying for _my_ expertise. I can imagine it could make a decent alternative to Google and Wikipedia for research purposes, but I'd have to double check it all anyway. I don't see how it'd make my work any easier.
I bet someone said this 30 years ago. As it stands today, LLMs can't wholesale do the job for you, they merely augment it. And they are finicky and fallible, so expertise is needed to actually produce useful results and validate them. There are many reasons why an LLM isn't the right tool for a certain task, wanting to drive on some sort of moral ethical high road however is imo _not_ a great reason.
Continue on that road and people who have no qualms about using whatever tool at their disposal will eat your lunch. Maybe not today or tomorrow, but eventually.
For support tasks (like research or even supporting a writeup) I'm not ethically opposed to using LLMs. It's just that I try every couple of months and conclude it's not a net improvement. Your mileage may vary of course.
On HN, they get a bad rap from the "i'm smart, I know everything, worker bee crowd" who think business is just slinging JS code and product updates. Consultants are often brought into those people orgs to deal with / or around, valuable know-it-alls, fiefdoms, poor vision / strategy, or politics that are holding the company back. They're often cited as "just doing reports", when in fact consultants sometimes stay out of implementation to make sure the owning team owns & implements the solution.
Consultants often provide the hiring manager a shield on dicey projects or risky outcomes. Where if it goes wrong, the manager can say "it was those consultants".
Teams are brought in as trained, skilled, up to date staff the company cannot hire fire, or shouldn't due to duration of need, or skillset
Sometimes they're brought in because the politics of the internal company lead to stalemates, poor strategy, etc.
Often at the higher levels, they're brought in due to focusing on a speciality, market, or vertical to have large experience that isn't possible to get in house.
One's experience with consultants frames their opinion. I've only worked with very high quality teams in the past that provide a healthy balance of vision, strategy, implementation, etc
I guess they have the advantage of experience doing these things over and over again for similar organizations. It's also an outside perspective, which has both advantages and disadvantages. In the conversations that I was invited into, we spent most of our time just explaining things that any mid-level member of the organization would know, trying to get the RAND people up-to-speed, so that seemed wasteful. But the military is a huge bureaucracy dominated by people who've climbed up the ranks for 30-40 years, so there isn't a lot of fresh outside thinking, and it seems like it could be valuable to inject some.
It is a problem though, because frankly, although you might think that moving to a more modern way of working is a simple fix, we have found that your employees are actually your enemies and are plotting to kill you and your family. You might think that you depend on them to get work done, but we can fix that for you and get our much more skilled and low paid workers to do the job instead.
You will find out that if you allow us to take the required steps to prevent your family from rape and murder your shareholders will get a bump as well! And of course you will as well, you can expect your bonus to at least triple, and you are worth it for sure.
We have already set up the culling stations filled with whirling blades to deal with the scum downstairs, if you say nothing at all we will act and slaughter them instantly. There is no guilt, only reward.
We love you.
Note: After reading the paper, there are serious methodology issues. ChatGPT without any human involvement (simply reading the task as instructions and producing an answer) produced a better "quality" result than any human involvement in many tasks.
My read: ChatGPT produces what we pre-view as "correct".
They even kind of state this: "when human subjects use ChatGPT there is a reduction in the variation in the eventual ideas they produce. This result is perhaps surprising one would assume that ChatGPT, with its expansive knowledge base, would instead be able to produce many very distinct ideas, compared to human subjects alone."
My read: The humans converge towards what is already viewed as "correct". Its like being in a business and your boss already knows exactly what your boss actually wants, and any variation you produce is automatically "bad" anyways.
There is also a drastic difference in the treatment of the in/out/on "frontier" subject. Lots of talk about how great adding AI is to "inside frontier" tasks, no similar graphs/charts for "outside frontier" tasks. Finally, if these people are consultants and business profs, they produce some crazily bad charts. Figure 3 in the appendix is so difficult to read.