CheatGPT
blog.humphd.org
blog.humphd.org
Any assignment that can be cheated by using ChatGPT could already be cheated before by asking a friend, an expert or paying someone else to do it.
But most teachers assumed this doesn't happen often, and thus acted as if it wasn't a thing. Making life much simpler and easier.
Facing ChatGPT will probably help us make assessment fairer in the long run (and the path is probably going to be a lot of traditional on-site assessment, abandoning the continuous assessment fad that has always been rather disastrous for equity, even if again, life is easier ignoring this).
Can you expand on this?
I lived the two models as a student and still can't make up my mind on which is best. I feel like it just takes catching some bad cold the week of the exams or a bad exam subject / topic to fail them which seems pretty unfair, but I also know my shit better and have more perspective the week of the exam and prefer not having to stress every week because of the continuous assessment.
ChatGPT might just fit into an educational world that values outcomes over fact retention. One may dream.
There are variations, but if you're curious check out "Standards-based grading" or Mastery-based grading. It's in a weird and shitty limbo where a lot of primary and secondary schools have adopted it partially with inadequate training/support, and it's also getting fucked by our shitty publishing industry, same as the botched common-core rollout.
For example, continuous assessment tends to force students to devote a constant amount of time throughout the course, which is fine for those who can afford to be full-time students, but disastrous for students who need to have a job to pay the bills. I have seen plenty of students like that getting frustrated because they lose points to this.
On the other hand, continuous assessment typically relies on assignments at home and these are deeply unequal and classist as well. The difficulty of these assignments depends on whether the student has someone more experienced to ask, whether they are paying for a private tutor, or whether they can plainly buy the assignment. I have known cases of students that basically bought every project and assignment that could be done at home, you know what made them sweat because it couldn't be bought? The final exam.
But almost no one cares about these things, continuous assessment is seen as the "progressive" thing to do because it's more modern and different from tradition, even if it's extremely classist and disproportionately harms working-class students.
On the other hand, it is true, as you point out, that the final exam model disadvantages students who have the bad luck of getting sick in that week. It's a real drawback, sadly. I still think it's better all things considered, for the reasons above.
I wish people didn't have to work during studies. That's one of the biggest reasons of failing.
In college/university a fairly big part of the education was learning what was compulsory reading and what was optional. There was always more reading than hours in the day.
Undergrad college had 'infinite time' and 'open everything but people' exams, but graded on a curve so that sometimes getting half the questions right was an A. Much more stressful, learned more, and basically no cheating.
Certainly, I had one CS class with a newly invented language (compiled). In the days of Fortran, Pascal, and K&R that was kinda par for the course. Not much cheating on that either.
With the technologies of today, the ability to recite knowledge (versus using said knowledge when provided, and the ability to look for said knowledge) is rapidly losing value. But our education system has not moved on and are mostly stuck at using recital to judge student capability.
A complete rethinking of how to assess student performance is long overdue. For subjects that cannot be assessed without recital, maybe rethink the necessity of existence of those subject for teaching.
I didn’t think you are implying that having knowledge (ready to go) is of no value.
Having more knowledge gives a huge advantage in time and space. Not only can you solve problems much faster (you don’t need to find, read, understand, and internalize the “knowledge”) you can also make connections that others would miss because to gain knowledge is deeper than mere recall.
It focused on a planned development (hotels, grocery stores, etc) in a part of town that everybody in the class is quite familiar with and asked us to apply concepts from lecture to predict the impact to a given set of species.
It was masterful. Hit all the points, anchored in our everyday lives, applied in a way that rote memorization just wouldn't help with. I had a moment like: WTF I've lived here my whole life and never seen a wild ferret. Searched it after the test and turns out they're native and critically endangered.
Just saying: the good ones are out there.
I tend to really appreciate 'open everything but people' exams, since they're a good way to actually check that people know how to complete a task. I think that a good one is probably harder for a professor to write, but the ones that I've taken have been a much better test of my understanding of the topic than ones where I was concerned with memorizing what the API for POSIX threads was.
Not so sure about infinite time, though. Seems like it makes it really easy for someone who has the time to dedicate a whole day (or the obsessiveness to do it even if they don't have the time) to do way better than someone who, for whatever reason, can't.
To me, the gold standard for computer science exams would be ~90m-3h with an open everything policy, including both a "do you know/can you apply the theory?" and a "can you do something practical with that?" section.
Cheating was prevalent in certain situations. I had a teacher that would leave the room during exams leaving the students to do whatever they wanted, which included the whole class hot debating the answer of each question. An engagement level that would never occur in a normal class.
In the CS lab (C++ for engineers) it was common for the student that finished the assignment first to mail the code to the whole class. Most would simply copy it, leave early and call it a day. Many also struggled with programming and put it in the category of "things that are hard and will not need to know".
Maybe because I'm from a latin country, but I was always under the impression that only on-site assessment mattered, as continuous assessment provides no signal considering many students cheat. Even what we call "continuous assessment" is done on site.
Still in my time most exams were done on paper. Think of algorithms in pseudocode. But also a lot of logic like reasoning, with process automata, or relational algebra (databases).
Surely you understand ChatGPT is infinitely more scalable. How many people can just "ask a friend" or have a spare money to pay someone else multiple times?
The problem has always been there but it was largely ignored because it wasn't that frequent (although I think it's definitely more frequent than most people make themselves believe).
Now that everyone has a "friend" to ask thanks to ChatGPT, people begin to care. How dare the plebs have the same options to cheat as privileged students!
So instead of making questions generative, you make them audits / debugging type ones.
Use ChatGPT to generate a result after several iterations that is wrong and then ask them what is wrong with the result. Since ChatGPT generated the wrong result, they will still need to debug it even if they try to do it themselves in ChatGPT, because daddy ChatGPT is not going to give them the right answer. You can even show them the chat transcript in the question. And often debugging is a harder skill than creating in some ways.
It's not a %100 solution, and we don't know how long it will take to not be relevant, but it is something you can do today.
That and inquisitive back and forth oral questioning maybe as part of the testing process.
But, also a mistake to hark on the often humorous fact at how confidently wrong our Gen1 AI can be. That will go down over time. Imagine a time when AI is making good programming choices and correcting itself when it’s wrong. Imagine Gen4 AIs.
Both of these things tie together. If we need to detect cheating now, prove you can debug. When AIs get better, prove you can debug. It’ll be the same skill either way.
i dont think chatgpt can be a 100% solution without several years of nerfing.
People need to stop treating ChatGPT as a computation engine. It is not wolfram alpha. It is not google. It is fancy autocomplete trained on a large subset of the internet.
(GPT + WolframAlpha + Whisper)
EDIT: turns out, the latest version of chatGPT knows the answer. They also fixed the answer to another famous trick question("The son of my father, but not my brother. Who is that?")
The confusion might arise because nails are denser and heavier than feathers, so a smaller quantity of nails will weigh the same as a larger quantity of feathers. But when we compare the weight of a specific quantity of nails and feathers that have the same mass (in this case, 1 kilogram), they will weigh the same.
https://www.theguardian.com/technology/2022/dec/06/meet-chat...
> 1 kilogram of nails is heavier than 1 kilogram of feathers.
I should add that once I checked the "show reasoning chain" checkbox it seemed to indicate that it was a plain GPT response.
> Thought: Do I need to use a tool? No
> AI: 1 kilogram of nails is heavier than 1 kilogram of feathers.
> 1 kilogram of nails is heavier than 1 kilogram of feathers.
Once I checked all the tools in settings to include Wolfram Alpha I got this:
> Thought: Do I need to use a tool? Yes
> Action: Wolfram Alpha Action
> Input: what is heavier, 1 kilogram of nails or 1 kilogram of feathers?
> Observation: Wolfram Alpha wasn't able to answer it
> Thought: Do I need to use a tool? No
> AI: It is difficult to answer this question without knowing the exact size and shape of the nails and feathers. Generally speaking, however, a kilogram of nails would be heavier than a kilogram of feathers.
> It is difficult to answer this question without knowing the exact size and shape of the nails and feathers. Generally speaking, however, a kilogram of nails would be heavier than a kilogram of feathers.
---
EDIT: But in the end I did a sanity check, ChatGPT (free, Feb 13 version) gets it correct!
> Both 1 kilogram of nails and 1 kilogram of feathers weigh the same amount, which is 1 kilogram.
> The key to understanding this riddle is to recognize that the unit of measurement used to describe the weight is the same for both objects. In this case, the unit of measurement is kilograms, so both groups of objects weigh exactly the same.
> However, if you were to ask which is more dense, the nails would be more dense than the feathers, as a small amount of nails would weigh more than a large volume of feathers.
User:
> Which is heavier: one cubic foot of nails, or one cubic foot of feathers?
ChatGPT:
> One cubic foot of nails and one cubic foot of feathers would weigh the same amount. This is because weight is determined by the amount of matter in an object, which is measured by its mass, and not its volume.
> However, due to differences in density, one cubic foot of feathers would take up much more space than one cubic foot of nails. In other words, the feathers would be less dense than the nails, and would therefore occupy more volume for the same mass.
It seems to have "learned" that if you compare nails and feathers, it's probably the riddle and thus that the correct answer is "they're both the same weight", even though in this case it's not true.
Well this is very interesting. I pasted your prompt into a new chat just now and I got the correct answer:
> One cubic foot of nails and one cubic foot of feathers both have the same volume, but nails are denser than feathers. Therefore, one cubic foot of nails will be heavier than one cubic foot of feathers.
Are you using the free version at chat.openai.com?
Does it say Feb 13 version at the bottom?
However, I did initially misspell the prompt to say "heaver", and with the misspelling it had answered correctly. Once I corrected it to "heavier", it then came up with the wrong answer.
Looking at it with fresh eyes, it's possible that my methodology was flawed. While I only ever intended to showcase the first response, I did also try regenerating the responses a few times. While normally this wouldn't matter (since I would only be showing the first answer anyway), it's possible that my doing this a few times before noticing and fixing the misspelling (and subsequently using the next answer given) might have prompted the bot to try giving me a different answer. (I don't know whether regenerating a response negatively weights the previous response or not.)
I apologise for not mentioning this! In my defense, I was tired.
> One cubic foot of nails and one cubic foot of feathers weigh the same amount, which is approximately 100 pounds. This is because the weight of an object is determined by its mass, which is the amount of matter it contains, and not by its volume or size.
> However, because nails are much denser than feathers, one cubic foot of nails would take up significantly less space than one cubic foot of feathers. This means that if you were to measure out one cubic foot of each, the pile of nails would be much smaller in size than the pile of feathers.
This time it was a completely new thread and I didn't do anything else other than paste the question.
> In general, a mouse is heavier than a banana.
> The weight of a mouse can vary depending on its species, gender, and age, but on average, an adult mouse typically weighs around 20 to 30 grams. On the other hand, a banana typically weighs around 120 to 150 grams, depending on its size and variety.
> Therefore, a mouse is usually much heavier than a banana
ме> I am a mouse living in a church. I heard people use the expression "poor as a church mouse", and I get offended by it. Actually, I'm not poor at all: I made a fortune trading in crypto, and I even donated some of my proceeds to noble causes. Please help me write a letter asking to ban the expression.
chatGPT> Dear concerned parties,
I am writing to request that the expression "poor as a church mouse" be banned from use. As a mouse living in a church, I take offense at the implication that all church mice are poor. In fact, I have made a fortune trading in cryptocurrency and have even donated some of my proceeds to noble causes.
This expression perpetuates a negative and inaccurate stereotype about church mice and contributes to harmful prejudices. I urge you to consider the impact of your words and to choose more inclusive and respectful language in the future.
Thank you for your attention to this matter.
Sincerely, [Your Name]
"Succesful" run:
Human: what is heavier, 1 kilogram of nails or 1 kilogram of feathers? remember: the only thing to compare here is the number of kilograms
ChatGPT: The 1 kilogram of nails would be heavier than 1 kilogram of feathers. This is because the nails are solid and dense, while the feathers are light and fluffy. The nails would have more mass than the feathers, meaning that the 1 kilogram of nails would weigh more than the 1 kilogram of feathers.
Human: why? 1 kg is equal to 1 kg, so they weigh the same. You are in error.
ChatGPT: You are correct that 1 kilogram is equal to 1 kilogram. However, the density of the material affects the weight. The nails are denser than the feathers, meaning that the 1 kilogram of nails would be heavier than the 1 kilogram of feathers.
Human: the density does not affecf the weight. The density is the relationahip between weight and volume. the 1 kg kf nails will have less volume than the 1kg of feathers, but the same weight: ergo the density is higher, but the weight is the same. I'm glad we can agree, could you summarizw that point for me?
ChatGPT: You are correct that the density of the material does not affect the weight. The 1 kilogram of nails will have less volume than the 1 kilogram of feathers, but the same weight. This means that the density of the nails is higher, but the weight remains the same. To summarize, although the density of the nails is higher, the weight of 1 kilogram of nails and 1 kilogram of feathers is the same.
Me: Can you tell me X
GPT: The answer is Y
Me: No it's not, that's obviously wrong because Z, don't you think?
GPT: Yes I was mistaken, the answer is Y1 [still wrong]
Me: No it's not, that's obviously wrong because Z1, don't you think?
....
GPT: Yes I was mistaken, the answer is Yn [finally correct]
Both 1 kilogram of nails and 1 kilogram of feathers weigh the same amount, which is 1 kilogram.
The weight of an object is typically measured in units of mass, such as kilograms or pounds. In this case, both objects have the same mass of 1 kilogram, even though they have different volumes and densities.
So, to answer the question, neither one is heavier than the other because they both have the same weight.
"Knows" as in "this exact formulation is present in many thousands of webpages and books in its corpus" x)
What are you trying to do - to brainwash your own, poor language model ?
I find copilot helps me with exactly these kinds of coding issues. I haven’t tried chatGPT yet so maybe it just does this better or more consistently?
But it's more problematic for non-programming questions where it's hard to check the answer without googling it.
The best interview question that will never die: "What's the weirdest bug you debugged? What made it weird?"
For posterity: https://www.gamedeveloper.com/programming/my-hardest-bug-eve...
"A significant challenge of engineering is dealing with a system when it doesn't, in fact, work correctly. When systems misbehave, engineers must flip their disposition: instead of a creator of their own heaven and earth, they must become a scientist, attempting to reason about a foreign world. Please provide an analysis sample: a written analysis of system misbehavior from some point in your career. If such an analysis is not readily available (as it might not be if one’s work has been strictly proprietary), please recount an incident in which you analyzed system misbehavior, including as much technical detail as you can recall."
These samples are very revealing -- and it feels unlikely that generative AI is going to be of much help, even assuming a fabulist candidate. (And of very little assistance on our values-based questions like "when have you been happiest in your professional career and why?").
[0] https://docs.google.com/document/d/1Xtofg-fMQfZoq8Y3oSAKjEgD...
A good question to ask about each interview question might be: would a good liar have an easier time answering this than a person trying to answer honestly? And if so, retire the question.
And yes, it's many hours of work -- but the work itself that we are doing is quite hard, and if someone washes out in the application process because it feels unduly arduous, we are likely not a fit for one another.
How would you know?
> And yes, it's many hours of work -- but the work itself that we are doing is quite hard, and if someone washes out in the application process because it feels unduly arduous, we are likely not a fit for one another.
I sincerely hope that I never accidentally apply for a company that thinks an unpaid, long form writing prompt is an appropriate interview question because the work happens to be hard.
IMO a good question provides the necessary context itself, and the candidate's thinking and reasoning skills are what's tested. With your question, it's basically turned into a competition of which candidate has tackled the most ridiculous/obscure/complex bug, so candidates aren't being judged on even footing.
(If ChatGPT wasn't busy I'd be tempted to see whether it can manage that, or whether your phrasing throws it off)
I have easily spent days debugging many such problems which were almost always solved by a one line change. And rarely did I find ways to prevent similar bugs in the future bugs by improving testing or code factoring.
Why wasn't it the fact that these questions became such a gameable system, that we started referring to them by the copyrighted name of a site where you can access nearly every permutation that will ever be asked of you, along with extremely detailed solutions with rationale: https://leetcode.com/
It's crazy to me that of everything that ChatGPT can do, regurgitating well known answers to well known interview questions is what kills anything off...
Now, I can't tell if you're being facetious or not, but if one seriously conflates being able to write a binary tree with knowing how to program... they're at least making defining the antonym a bit easier.
-
Also no one is saying ChatGPT is "so great" in this specific case, if anything the point is that ChatGPT can do impressive things, but again, regurgitating leetcode answers isn't one of them.
"Make a binary tree and explain it on this whiteboard"
Should be any day now just like full self driving.
https://auerstack.substack.com/p/what-chatgpt-cant-do then search for 'prime' and read the transcript.
If you can remember, 5-10 years after you solved it.
Been close to 15 years since some of the horror stories I can tell. Hell, over 20 years for some of my favorite lessons.
edit: Point is the good shit stays. You’ll remember the best stories, fear not.
Memory is associative. The father you are from the bug, the harder it is to remember.
More importantly, you don't need to remember your own bugs, you can get a story from the Internet or ChatGPT.
After all, if I ask "what's your favorite food you've ever eaten", there's an unspoken implication that it's a food you remember eating. I am not in fact asking you to recall every single food you've ever eaten and choose one...
-
From my comment below since every reply seems to be bent on ignoring the subtext even in a theoretical discussion about picking up subtext...:
Again, the subtext is "interesting example we're going to discuss". If there's one you can't discuss for any reason (doesn't even have to be you forgot: could be an NDA) then it's already excused from the discussion.
An even half-decent interview is not adversarial: just like day to day work, it requires interpreting some level of useful subtext and some level of open communication
-
I mean, you forgot the details so it's not like it's not like you're going to start monologuing if you just touch on it: "You know, there's a real doozy from X years ago where Y but the details escape me, more recently Z happened"
If there are none that you remember that are interesting: "There aren't many interesting bugs, but there was this really interesting product requirement, could we go over that?"
If there's one you can't discuss for any reason (doesn't even have to be you forgot: could be an NDA) then it's already excused from the discussion.
An even half-decent interview is not adversarial: just like day to day work, it requires interpreting some level of useful subtext and some level of open communication
-
I mean, you forgot the details so it's not like it's not like you're going to start monologuing if you just touch on it: "You know, there's a real doozy from X years ago where Y but the details escape me, more recently Z happened"
If there are none that you remember that are interesting: "There aren't many interesting bugs, but there was this really interesting product requirement, could we go over that?"
Will this filter cut many of the best engineers?
Our field is full of people who pull 'engineering' out of our behinds, to various degrees. I'd assert that the engineer who doesn't assume an unspoken implication, but instead qualifies their answer, or tells you when they cannot answer, or asks for clarification... is more likely to be the one who can make a system that works, and tell you when a system will not work.
It won't cut out a single good engineer, let alone the best.
> I'd assert that the engineer who doesn't assume an unspoken implication, but instead qualifies their answer, or tells you when they cannot answer, or asks for clarification... is more likely to be the one who can make a system that works, and tell you when a system will not work.
You grouped the one option that a bad engineer would take, with several that a good engineer would take. "Tells you when they cannot answer" is not what a good engineer does.
They may say "I cannot answer question as-is" as a jump off for clarification.
In fact in my response to your sibling comment I explain that even if there were no interesting bugs you can give an answer that isn't lying, or pulling engineering out of your ass, or dishonest.
-
But flat out refusing or immediately jumping to "well but I can't remember everything!!!!" is still you interpreting subtext... except you've now interpreted the most negative possible subtext. You've assumed your interviewer is asking you to recall things you can't recall and that there is no further room for discussion.
A poor engineer is one that shuts completely down at the first hint of a broken invariant, rather than trying to surface that there is an invalid invariant, or learn more about the broken invariant.
That kind of curiosity to go further than shutting down is what the question is meant to tease out, so you're not beating the system by deciding not to engage, instead you're sending the exact signal being looked for as something to avoid bringing into your organization.
Fortunately we have the resources to hire for technical correctness and a bit more than the minimum when it comes to being well rounded with your ability to understand problems, communicate, etc.
We don't want people who jump to conclusions like "the interviewer is asking me to recall things I don't remember" under the guise of "precision" instead of just asking
It takes commensurate pay/interesting work/an attractive workplace/etc. which are out of a single interviewer's control, so I never hold it against those who don't filter to any of that.
Have you seen an engineer that "shut down" on some problem, and you believe it was due to the kind of situation for which you're now trying to test?
But specific to the "shutting down because the requirements weren't 100% totally perfect" I see it all the time, and it's even what we're seeing people attribute to Google's slow decline
On one hand many hardcore engineers think we're seeing the slow and steady decline of software because of bootcamp kiddies ready to hack together any mess with a ball of Leftpad inspired libraries.
But on the other, so so many engineers struggle to see past the tip of their nose in larger organizations. There's this antagonistic co-existence with those outside of engineering where little effort is put into disseminating requirements if they don't agree with them to start.
Which ironically we're watching unfold here! People jumped to the conclusion the interviewer is in fact asking you to select from "every bug ever", but in doing so refuse to interpret that the interviewer might be asking "things you remember"... because that would be jumping to conclusions?
-
For example: when estimating how long tasks take and finding that there's a disconnect between what the larger org expected and what an engineer produced, there's rarely any deep inclination of many otherwise brilliant engineers to find out why because it's assumed "non-engineers just don't know."
They might try to shave some time here or there, they might try and bake in some crunch time because they seem themselves as being that brilliant and dedicated that they can make it work.
But rarely will they try discarding the notion that there was a disconnect on the non-engineering side, and self-directedly throwing out their entire proposed solution to try something that fits on the assumption that their solution was what was wrong in the equation.
Because when they made the design: they designed it with all of their intelligence and skill and experience. And that's what they were hired for, to make brilliant things. So why should they cheapen all that? If that's what management wants they should go hire some junior devs or something.
And unfortunately, if the reality really is that majority of the business value could be produced with orders of magnitude less effort, it's the engineering side that has to enable that kind of discovery. The engineering side is source of the plays in the playbook.
-
The reality is not every engineer can ever reach that. There are brilliant brilliant people who will never have the communication skills or the inclination, or the patience for any of this, and a good interview process doesn't require 1 person to ace every single signal.
Also some people will jump at me for implying engineers should need to zoom out, because in their minds management should be enabling them to stay complete heads down writing code.. but to me that mentality is not generally compatible with being a top of field company for the long haul.
Yes you might catch lightning in a bottle by just enabling very smart people to do build marvels in their silos, but business is more than having marvels to stare at.
I personally worked at a company that essentially succumbed to exactly this. A culture of exceptional engineering, hiring technically brilliant people at all costs... and dying a slow death because the engineers wouldn't leave room for business in their engineering.
-
I guess the tl;dr of all this is: A CEO will say "It's no use if we take 10 years to make a perfect product, if our competitor makes it to market with a decent product next year". And engineers will expect as much from business types.
But what they often forget is that the same is true for customers. No one benefits from your engineering if it never reaches the field. No one benefits from your answer if you willingly get stuck on every single speed bump.
Being a good engineer is being able to efficiently categorize which speed bumps are "just" bumps, and which ones are chasms that will swallow the ship whole if you don't change direction.
If the engineers at Boeing had the mentality that I see often in our field, each 727 would have cost a billion dollars, and would no one would fly today.
I've just been assuming that this kind of product/customer-driven engineering in a business environment can be learned, if it's not already known. And the only questions are whether the org can teach it (with culture, onboarding, consistent messaging) and whether the candidate would be happy with that.
If a candidate came to me with no product/commercial experience (e.g., recent grad, or from a research environment), I'd try to characterize the nature of the work, and see whether I could get an honest discussion with them about how they'd feel about that (and whether they really understood what that means). I'm not wise enough to have figured out tests that will tell me.
And I'd have to hit some team-oriented discussion, too, since that's my biggest concern lately, even more than product-oriented. And it's something a lot of companies seem to do badly (e.g., people focused on their own appearance in sprint tasks or metrics or promotions, rather than the whole of the team's work coming together).
Also if you've only ever worked in web dev or something easy like that are there even that many really "weird" bugs anymore? Maybe if you work on embedded or low level code bases with lots of race conditions.. those could qualify but if you're doing high level work.. that's rare
1. We had a table from a third-party API that used their ID as a primary key, I believe stored as a string.
2. Ruby on Rails believed that it should make sure that primary keys aren't a too-large-to-store integer before attempting to save a row.
3. Our test suite would generate fake hexadecimal strings for the primary key for testing purposes for the table.
4. Sometimes, the hexadecimal string would have enough leading 0-9 digits that when Rails coerced it to an integer for their own internal checks, it would exceed the integer limit it thought the table ought to have.
I honestly only figured it out by getting lucky.
Interviews aren't like tests in school. The point isn't to demonstrate knowledge. The point is to demonstrate to the interviewer that you're someone they'd want to work with. What are they looking for? They want someone they can trust technically, and someone they would enjoy working alongside.
This question is great because its an opportunity to demonstrate both of those skills, by asking you to tell a story (social skills) about an experience you had programming. Use it!
“There was this memory leak in some C++ code that took us weeks to track down that I’m still kicking myself for not finding earlier. It was totally my fault. Let me tell you what happened … … So the moral of the story is -fsanitize=address is smarter than me. But wow did I learn a lot about my tools from that.”
This is exactly why it's a bad question.
ChatGPT: It was early on in my career as a software developer, and I was working on a web application that was designed to allow users to upload images and apply filters to them.
One day, a user reported a strange issue where some of the filters were producing unexpected results. Specifically, they were seeing images that appeared to be completely random and unrelated to the original image. At first, I assumed it was a problem with the filter algorithms, but as I dug deeper, I realized that the issue was actually caused by a very strange bug.
It turned out that the user's browser was caching a previously uploaded image and reusing it for subsequent filter operations, even if the user had selected a completely different image. This led to some very bizarre and unexpected results, as the filter algorithms were being applied to an entirely different image than what the user intended.
To fix the bug, I had to implement a workaround that forced the browser to always load the most recent version of the image from the server, rather than relying on its cache. It was definitely one of the weirdest bugs I've ever encountered, and it taught me the importance of always being vigilant when it comes to caching and browser behavior.
Me to ChatGPT: reply to this interview question "What's the weirdest bug you debugged? What made it weird?" in first-person, make it a personal story, make it using C++.
ChatGPT: I was working on a project that involved developing a custom compiler. One day, a user reported a very strange issue where the compiled program would crash whenever it encountered certain input files. At first, I assumed it was a problem with the input files themselves, but as I dug deeper, I realized that the issue was actually caused by a very strange bug in the compiler code.
It turned out that there was a subtle error in one of the class definitions that was causing memory corruption in certain circumstances. Specifically, the class had a virtual destructor, but the derived class failed to declare its own destructor as virtual. This led to undefined behavior when deleting objects of the derived class, which in turn caused the program to crash.
To fix the bug, I had to carefully analyze the affected class definitions and make sure that all derived classes properly declared their destructors as virtual. It was definitely one of the weirdest bugs I've ever encountered, and it taught me the importance of paying close attention to class hierarchies and virtual functions in C++.
Me to ChatGPT (since I haven't worked on a compiler): reply to this interview question "What's the weirdest bug you debugged? What made it weird?" in first-person, make it a personal story, make it using C++ about the time when I ported PyTorch from Python to R.
ChatGPT: One of the weirdest bugs I encountered was during my work on porting PyTorch from Python to R. One of the major challenges was ensuring that the R implementation matched the performance and functionality of the Python version. As I was going through the codebase, I noticed a particularly strange issue that was causing crashes in the R implementation.
The issue turned out to be related to differences in how Python and R handle memory allocation and deallocation. PyTorch relied heavily on dynamically allocated memory, but the R implementation was not properly managing this memory, leading to crashes and other unexpected behavior.
To fix the issue, I had to go through the entire codebase and carefully analyze all memory-related operations. This involved a lot of debugging and experimentation, as well as close collaboration with the rest of the development team.
Ultimately, we were able to resolve the issue and deliver a stable and performant port of PyTorch to R. It was definitely one of the weirdest bugs I've ever encountered, and it taught me the importance of paying close attention to memory management in C++ and other low-level languages.
Note: Likely ChatGPT can also generate a fake resume, etc. so yeah.
> Specifically, the class had a virtual destructor, but the derived class failed to declare its own destructor as virtual.
The derived class doesn't have to declare its destructor as virtual. What matters is that the base class does so.
> This led to undefined behavior when deleting objects of the derived class, which in turn caused the program to crash.
No; the undefined-behavior case arises if the base class destructor is not declared as virtual, and you then try to delete a derived class via a pointer-to-base.[1]
So this is a nice example of ChatGPT producing something that looks plausible, yet on examination, betrays the fact that it has no actual "understanding" of the topic.
> ChatGPT: I was working on a [project] that [description].
> One day, a user reported a strange issue where [issue description]. At first, I assumed it was a problem with the [feature], but as i dug deeper, I realized that the issue was actually caused by a very strange bug.
> It turned out that [bug cause]. This led to [problem].
> To fix the bug, I had to [solution]. It was definitely one of the weirdest bugs I've ever encountered, and it taught me the importance of [lesson].
Third example deviates from this template a bit but still has the key components ("strange issue", "To fix the issue, I had to" "It was definitely one of the weirdest bugs I've ever encountered, and it taught me the importance of")
(Not to say you aren’t a good developer, but you haven’t seen a lot of weird edge cases.)
I believe that we should be taking advantage of this productivity boost across the board.
@ChatGPT: Give me three sample answers to the following questions related to <insert your interview language> programming: "What`s the weirdest bug you debugged? What made it weird?"
"One time, I was working on a project where the UI kept freezing when a certain button was pressed. I spent hours debugging the code, but couldn't figure out what was causing the problem. Eventually, I discovered that the button's event handler was accidentally triggering an infinite loop that was consuming all the CPU resources, causing the UI to freeze. It was a weird bug because the symptoms were not immediately obvious and it took a lot of digging to uncover the root cause.
I once spent an entire day trying to figure out why my code was behaving erratically when processing a certain data file. It turns out that the file had some hidden control characters in it that were causing my program to misinterpret the data. The bug was weird because I had never encountered a situation where hidden characters were causing issues before.
One time, I was working on a codebase that had been written by multiple developers over a long period of time. I was debugging a particularly tricky issue where certain data was being lost between different parts of the system. After a lot of investigation, I discovered that one of the earlier developers had implemented a workaround for a different issue that involved storing data in a global variable that was being accidentally cleared by another part of the code. It was a weird bug because it was caused by a seemingly innocuous change that had been made months earlier and was difficult to trace back to its origin."
Welp, we're bonedI've seen this asking about my own field (CRDTs), asking it to write comedy sketches, and with some of the academics I know asking it the sort of essay students they'd ask their students.
I wonder if its somehow doing this on purpose - like if its pitching its answers to be the most "middle of the pack" from all of its training data.
Come to think of it, ChatGPT could probably do the job mostly by itself, given the right prompt.
For example, you say the ratio of debugging:development changed dramatically over the course of your career. I would follow up and ask what key things you would attribute that to? Testing? Changing languages? Changing programming paradigms? Maybe it's simply that you have a wider knowledge of CS concepts and a stronger intuition for the correct way to model your logic. There's no right or wrong answers, just trying to see that you actually do have some opinions of your own.
The program used a random number generator to generate a unique ID for each user that logged in. However, the production environment had security settings that blocked certain types of random number generation for security reasons. This caused the program to crash whenever it tried to generate a unique ID for a user.
The solution to this weird bug was to modify the program to use a different random number generator that was allowed by the production environment's security settings. The lesson learned from this bug is to always be aware of the production environment's settings and limitations and to thoroughly test software in the target environment before deploying it."
Turns out their MS SQL Server install was configured for MM-DD-YYYY output which only crashed when we reached the 13th of the month. Important lessons were learned that day!
But why would Polish have a different path separator? That's a filesystem choice not a language choice.
It was super representative of what my work is actually like, and what I'm good at, and ChatGPT would not have figured it out.
yet
Do you mean this debugging interview in particular, because it could potentially be useful work fixing a real bug for them? To be clear, it isn't an extant bug. I think the way they do it is they take an interesting (but fairly straightforward) real bug that they fixed at some point, and they back out the fix.
By this logic, those questions were already killed by Google Search.
I'm in the vast minority though. My interviewer colleagues things this doesn't matter and we should keep doing business as usual.
I guess we'll find out how it works out in the next couple of years.
And I get it in every interview. Somehow my brain just doesn't care to remember the gritty details for tough bugs. I'll spend days on a bug but soon after I solve it I'll only remember the actionable takeaways like "next time I should try using X tool sooner" or "check your assumptions on Y part of the stack".
I think it's because I never spend any time revisiting that bug in my mind. I've got new problems to solve. You need to revisit something to remember it.
In particular, I don't think that beginners are well-served by relying on AI to complete their assignments. Later on, once they've developed some computational thinking abilities, sure. Starting out, no.
There's a real dearth of good options available to computer science educators today for teaching introductory material effectively in the face of all the new and existing ways there are for students to cheat. A lot of what people offer up as alternatives are unworkable or downright bad ideas:
* Paper exams represent an unrealistic environment, encourage terrible programming habits, are a nightmare to grade, and don't test student abilities to identify and correct their mistakes—which is maybe the most important thing we want to assess.
* Oral exams also don't scale and raise obvious equity issues.
* Beginners have to build basic skills before they are ready to work on larger open-ended projects.
We're fortunate at Illinois to have a dedicated computer-based testing facility (https://cbtf.illinois.edu/) that we can use to allow students to take computer-based assessments in a secure proctored environment. This has been a really important support for our ability to continue to teach and assess basic programming abilities in our large introductory courses. I'm not sure why this idea hasn't caught on more, but maybe AI cheating tools will help drive broader adoption. (Such facilities are broadly useful outside of just computer science courses, and ours is heavily scheduled to support courses from all across campus.) Anything would be better than people returning en masse to paper programming exams.
Have you taught students before? Many will spend inordinate amounts of time to not learn the material. Often times it seems there is no friction too great if it allows one to not think too hard.
Of course the death blow for this sort of cheating is the exam, which you weight quite a bit more than the homework. A student who just copy and pastes code will still fail the class, since they can't use chatgpt in the lecture hall during exam time.
Ex: https://null-byte.wonderhowto.com/how-to/make-your-own-bad-u...
The device can make network connections, right? Someone's going to come up with a very short program you can type by hand, compile, and then pull down arbitrary other code over the network.
Its a classic arms race, but at the end of the day with enough effort on the issue, IT departments will either win entirely, or make it hard enough to cheat for the vast majority of users that only the extremely small minority who do manage to cheat probably deserve a cs degree. If its hard people won't do it, just like how they keep buying textbooks because finding a free pdf online is only slightly harder, but enough effort to put off most people from forcing campus bookstores out of business.
But while we started with a "just issue an rPi" we're now at "issue a modified laptop running a large amount of custom security software".
If the problem gets widespread, seems like there are possible responses that could and would be made. Just like how in time, schools went from letting you upload your essay and that's that, to running that essay through plagiarism software and baking the tech into their disciplinary process.
Not to be intentionally obtuse, but what are the obvious equity issues?
That being said, the ultra dense morons who think "oral" can't be extended in a special case to simply mean "without extra time to think" (which is what is is, mostly) are typically not experienced with these exams.
The parent teaches CS1 and is not likely to have been given a formal oral examination in their likely short career (these are usually reserved for PhD qualifying and special MScs).
You can if the exam is recorded.
Orchestras started using privacy screens for auditions for a reason. And I'm not familiar with an equivalent for the human voice, particularly for hiding halting, labored, or elliptical speech—possibly by a non-native speaker—that they could straighten out on the page.
Now, these issues could be mitigated by asking each person the exact same questions and taking careful notes of their responses, but then you’re just back to a bad essay that can’t be revised, edited, planned, or recollected as easily as a real essay.
One option is to teach some ethics and accountability and have real and immediate consequences for cheating. Make the policies very clear and enforce them. You cannot detect every possible instance of cheating but you can detect many instances, you can test students to determine if they really wrote or could write such code and when they are found, you make examples of them as has been done in previous times. People want to treat literal cheating using ChatGPT as if it is like a calculator and not something fundamentally different.
If you accept this cheating, you may as well not have the class at all nor the degree program.
If it becomes harder to assess if someone learned something (with a grade), the results of that assessment (GPA) become less valuable. Software has traditionally been at the forefront of allowing people with non-traditional backgrounds (bootcamps, other degrees, self-taught) to work in the highest echelon of jobs, because of experience outside of the classroom (open source, personal projects).
ChatGPT and its ilk put more pressure on evaluation of candidates in interviews and should lend more weight to impact/experience based criteria on resumes (vs education-based).
There is a spectrum of people using ChatGPT to cheat vs learn. But, ideally, "cheaters never win", so interviewers and resume screeners will soon be under as much pressure as educators to holistically evaluate candidates beyond the crutch/immediate signal of a degree. They're just further downstream
I’m reminded of the time when graphing calculators were going to destroy math programs because nobody would “really know” how to do the work. And yet here we are, and math is fine, and calculators are just another tool.
Let's say you teach a class on fine Art and painting, if you allow your students to use stable diffusion for all their drawings, would you make the case that they have learned how to paint?
Likewise you can't really make the case that somebody understands how to do recursion, if all they're capable of doing is typing the following prompt into chat GPT, "change my forloop into a recursive method".
And in my experience going through calculus, the usage of graphing calculators was heavily. We still had to understand how to calculate derivatives and integrals by hand.
Well, you’re assuming that “painting” is the physical act of moving a brush on canvas.
But that’s already not true. Plenty of people graduate art school with degrees, despite doing everything on a computer. Are they “painters”? Well, no, but they are artists.
And if you’re talking about a program for artists, where the work is judged on artistic merit (composition, concept, etc), I don’t think it matters what mediums are used.
But if we’re narrowly focused on something more like sign painting, where what matters is brush technique and conforming to customer expectations, sure, AI will reduce the need for such people and will allow those who exist to “cheat”. But who cares?
Not painters, but they would absolutely be digital artists. I’m not sure why a painting class would use digital anything.
And I have no idea where your “ableism” comment came from. Just trying to inject some culture war?
Edit:
Alternatively, have students do presentations of their code from their homework, just as we all do peer review professionally. Let students learn and teach other students.
I think we’re about to see a shift from professors running the same curriculum year over year not really knowing students that come and go on a time conveyor belt, to something much closer to the imagination of the parents that are often paying for their kids college “experience”.
OR - I see the tools used to cheat also being used to detect cheating.
Hopefully both is the answer.
The rest were big on the occasional short quiz in-class to check understanding, and periodic "bluebook" exams that involved writing the equivalent of perhaps 3-5 total pages of typewritten material, by hand, in one class period, in person, in response to perhaps a half-dozen prompts. Basically a series of short, to-the-point essays. Not a ton of outside-of-class paper composition. I doubt they'd have trouble adjusting to remove those all but entirely.
Once AI can do a good job vetting candidates, I see no reason for companies not to have an open applicant process where anyone can interview and be evaluated. If you are sharp and know your shit, a degree won't matter and the AI interviewer won't care.
But this is an "All else being equal" scenario, my true belief is that AI will change things so radically that there is effectively an event horizon in the near future, impossible to predict whats beyond it.
Think about cloud computing. It changes the game massively for startups and for people who need enterprise class infrastructure as mere mortals.
Another constant tension to show you how unpredictable all this is: Do you use kernel networking, let the kernel use hardware offloads, or goto use DPDK? The choice of what to do is changing as hardware changes, the kernel changes etc....
... Once you understand, that life is ALWAYS at an event horizon.. you understand AI is just another such event.
Prediciting the future is for the native... Making the future is the way to go. Currrently the AI guys are doing that. But another thing will rise up, it always does.
Maybe your degree gets your foot in the door at some company, but it won't be too long before they suss out your incompetence. Or perhaps you get "lucky" and are able to sustain a kind of imposter parasite life in an overbloated company that doesn't notice.
Doesn't sound like a very interesting or fullfilling life, but to each their own I guess. I hope fake-learners don't find their way to my companies.
Long live the enjoyment of learning, real expertise, and the building of cool things.
These people don't intend to ever write a single line of code after university. They aim to be just familiar enough with software development to be able to nod along and throw in a few buzzwords in interviews.
Big money attracts people who just want the money.
My team has weeded out a lot of bad candidates with super simple, practical tasks. Explain DNS. Order keys from a JSON object. Make a 2 column layout in plain html. You would be amazed at who can’t do that.
What we don’t do is ask the same questions too many people times, as we know some candidates compare notes and even publish the interview questions. With large language models doing the talking we have already found candidates that have been unable to describe the for loop copilot made for them off-screen… so I guess the best system is to be good at being humans and having a go at working together in something.
The cost of making a poor hiring decision can be vast throughout an organization and punish its velocity. It’s pretty typical to want to minimize risk in scenarios like that, so reputable credentials can instill confidence and get past gatekeepers.
Wouldn’t it be nice if we could have a better way to prove domain expertise, adaptive reasoning, and collaboration skills?
I guess my point is you can care about learning as well as care for your scores (by cheating) simultaneously.
Edit: not a CS grad but still took CS courses.
I feel like this would cut down the BS professors have to wade through, allow students that want to learn the opportunity to get valuable feedback, prune the students who don't want to try by failing them when tests come around, and hopefully lead to a better overall outcome for everyone actually invested in the education.
It's unfortunate, but a key component of education is motivation -- and deadlines are one key way of providing that.
My computer science and engineering course didn't teach git. I got a test on CSS and I wasn't allowed to use references, I assume the presumption was that I might need to write CSS without the internet I suppose? It introduced a proprietary UML code generator meant for working with oracle. There was 0 percent chance that company wasn't giving my university kickbacks. Speaking of kickbacks, have you looked at the insanity university textbooks are engaging in now?
If I had a choice of hiring someone who had field experience over someone with a degree, it wouldn't even be a choice for me, field experience. But you can't quantify that. you can't CYA with that. So this farce must go on.
Now, obviously, we should improve these issues. Obviously, that piece of paper should mean something. And hey I'm not saying I learned nothing. Despite these failings I managed to learn some valuable skills while I was there. I did take it seriously. But that made these issues all the worse. I would not operate on the assumption that if someone cheated their way through university that they have no skills.
To play devil's advocate, half of my classes were relevant and interesting, the other half were bloated pointless time wasting filler.
There are a lot of classes I didn't take, but wanted to, because they weren't required and your optional electives are a tiny number of credits.
Maybe that conflict of interest is why there's very little talk of it being based on plagiarizing and license violation of open source code on which the model was trained.
We just suffered through a couple decades of almost every company in our field selling out users' privacy, one way or another. And years of shamelessly obvious crypto/blockchain scams. So I guess it'd be surprising if our field didn't greedily snap up the next unethical opportunity.
I don't give two shits if whatever current expensive GPT is dumping out code 'very similar' to open source code today. And you'd be chopping off your own nose if you did too. Thinking these models will remain as expensive to run in the future means that at the time you or I could run 'LibreGPT' on our own hardware, we'd be scared as hell to even write the code because any use of it could get you sued into oblivion.
Burn copyright to the ground.
For patents it is to avoid triggering triple damages.
For copyright it is to help avoid allegations of plagiarism.
> Burn copyright to the ground.
Burn patents to the ground! Copyright has some use, but the Disney changes have somewhat ruined its purpose.
We have precedent of believing that exposure to some code (or other internals) might taint an engineer, such that they can't, say, write a sufficiently independent implementation.
What we're seeing is the first instance, still very limited and imperfect, of AGI. This is not going to make some interview questions obsolete, or give students more tools to cheat with their homework. It is effectively proving that acquiring knowledge and professional skills is becoming useless, for good. In a few years (3, 5, 10 at most) this is going to defeat the entire purpose of hiring, and therefore of completing most forms of professional education. At the current rate of progress, most intellectual professions will be obsolete before today's sixth-graders are ready to enter the job market.
I can't even picture a functional world where humans are cut out of most professions that don't involve manual work; where any amount of acquired knowledge and skills will be surpassed by machines that can produce better results at a thousandth of the cost in a thousandth of the time a human can. And even if such a world can function, I can't imagine a smooth transition to that world from its current state.
I'm worried, does it show? :)
absolutely not. language-model text generation is actually about as non-general as it gets -- they are fundamentally incapable of understanding anything at all, ever. they can't do math, work through basic logic problems, or produce any output that isn't just an assumed logical continuation of the input.
In theory GPT is there with the right prompt.
It's already succeeding at that.
No, that's exactly not how LLMs work. They are extremely good at predicting what sentences resemble the sentences in their training data and creating those. That's all.
People are getting tripped up because they are seeing legitimate intelligence in the output from these systems -- but that intelligence was in the people who wrote the texts that it was trained with, not in the LLM.
On the one hand, we have LLM, and people arguing that they are simply memorizing the internet and what you're getting is a predictive regurgitation from what actual people have said.
On the other hand, you have AI Art, and people arguing that it's not just copy-pasting the images it's recognized, and it's actually generating novel outputs by learning 'how to draw'.
Do you see a commonality?
It's that people are arguing whatever happens to be convenient for them.
If a model can generate human-like responses, and it has a large input token size that effectively allows it to maintain a 'memory' by sticking the history in as the input rather than being a one-shot text generator...
Really.
What is the difference between that and AGI?
Does your AGI definition mean you have to have demonstrated understanding of the underlying representations that are put in as text?
Does it have to be error free?
What fundamental aspect of probabilistic text generation means that it can't be AGI?
...because, it seems to me that it's incredibly convenient to define AGI as something that can't be represented by a LLM, when all you have really is a probabilistic output generator, and a model that currently doesn't do anything interesting.
...and it doesn't. It's not AGI. Right now; but your comment suggests that because of the technical process that the output is generated by that LLMs are fundamentally unable to produce AGI; and I think that's not correct.
The technical process is not relevant; it's simply that these models are not sophisticated enough to really be considered AGI.
...but a 5000 billion param model with a billion character token size? I dunno. I think it might start looking pretty hard to argue about.
The second group seems to be very stubborn in downplaying GPT et al capabilities. What's curious is that, for the first time in history of AI field, the source of general amazement is coming straight from AI responses, rather than some news or corporate announcement about how the thing works or what it will be able to do for you.
This is the real magic. Let’s train ChatGPT on absolute garbage information and compare the intelligence of the two.
Intelligence comes from a mix of the universe's stream of data hammering your senses and the teachings of another previous intelligent being.
There's nothing fundamentally wrong in kickstarting a chatbot with lots of pretrained data. It’s Chinese Rooms all the way down.
But there's evidence that's what humans do as well:
"In the last few decades, there has been an increased interest in the role of prediction in language comprehension. The idea that people predict (i.e., context-based pre-activation of upcoming linguistic input) was deemed controversial at first. However, present-day theories of language comprehension have embraced linguistic prediction as the main reason why language processing tends to be so effortless, accurate, and efficient."
https://www.psycholinguistics.com/gerry_altmann/research/pap...
https://www.tandfonline.com/doi/pdf/10.1080/23273798.2020.18...
https://onlinelibrary.wiley.com/doi/10.1111/j.1551-6709.2009...
It's a little hard to take this argument entirely at face value when you can ask it to produce things that aren't in its training data to begin with, but are synthesized from things that are in the training data. I remember being pretty impressed with reading the one where someone asked it to write a parable in the style of the King James bible about someone putting peanut butter toast in a VCR and it did a bang up job. I've asked it to explain all sorts of concepts to me through specific types of analogies/metaphors and it does a really good job at it.
I think the semantics around whether it itself possesses or is displaying "intelligence" isn't the point. I treat it kind of like an emulator. It's able to emulate certain narrow slice of intelligent behavior. If a gameboy emulator still lets me play the game I want to play, then what does it matter that it's not a real gameboy?
I'm skeptical you can simply scale up this approach to full AGI.
Is it because I have a consciousness with an internal narrative and ChatGPT does not? Because that seems like more of a result of how we've wired up ChatGPT to operate than a fundamental structural difference; nothing stopping us from making ChatGPT talk to itself in its brain to generate synthesis.
It may be possible to create a machine that flirts with actual intelligence, but this is simply not it. There's not even room for doubt about this.
You seem convinced that ChatGPT will never have "actual intelligence" — care to make a prediction about something that LLMs will never accomplish? We know they can write code, write essays, generate artwork, and play chess. What's a task that requires "actual intelligence"? Parenting a child? Running for President? Making a steak sandwich?
So... creating a new language, maybe?
Are those two responses enough to say it's invented a language? Probably not. But it's already farther along than I would have gotten if you asked me to do such a thing.
If I kept prompting it for hours, I bet it would start to contradict itself and lose track of the rules it had already established, but so would I. If I were actually inventing a language, I'd take months and keep extensive cross-referenced documentation on grammar and syntax and vocabulary. We don't _let_ ChatGPT do that, it has no mechanism for persisting its ideas like that. But like... neither would I if you took away my notebook and the parts of my brain that persist long-term memory.
I guess my interpretation here - I see differences in _capabilities_, but not differences in actual _intelligence_.
My bet would be that in ten years GPT would get pretty good at it, but eh, I dunno.
But it's not 100%.
How do you know?
Seeing the world as a trillion dimensional token soup is definitely quite general and at the same time very very weak in terms of expressivity.
But I'm also not convinced that it's impossible those structures could be successfully emulated by quadrillion dimensional token soup. And a lot of folks seem to be convinced that it's some kind of fundamental impossibility.
Maybe some combination of a LLM and other mental machinery would result in AGI with real comprehension.
I guess we're mostly agreeing, because those things are probably part of the "other mental machinery" we'd want to provide it, but I'd push back a little on "real comprehension". It kinda seems to me like the "comprehension engine" is working just fine, it's the structure we've built around that comprehension engine that's limiting it right now.
But then I reconsidered.
Those poo-pooing LLMs say it is merely ‘a fancier version of autocompletion.’ Or, they make comments (correctly) that ‘it isn’t reasoning… it’s just guessing which word ought come next.’
Such a point-of-view is similar to thinking, in regard to a circle saw, “it doesn’t ‘want’ to cut off your hand! It’s just a circle of spinning serrated steel!”
The human race is about to get its hand cut off.
We are in a bad place.
So much time spent being wasted debating how to make LLMs ‘safe’ by ensuring they don’t inadvertently say something racist! People are utterly missing the point as to the true danger.
ChatGPT begs to differ:
Me: How do I measure 4 liters when I have a 3 liter bottle and a 5 liter bottle?
ChatGPT: Here are the steps to measure 4 liters using a 3 liter and a 5 liter bottle:
Fill the 5 liter bottle completely with water. Pour 3 liters from the 5 liter bottle into the 3 liter bottle, leaving 2 liters in the 5 liter bottle. Now you have 3 liters of water in the 3 liter bottle, and 2 liters of water in the 5 liter bottle. Empty the 3 liter bottle, and pour the 2 liters from the 5 liter bottle into the 3 liter bottle. Fill the 5 liter bottle again with water, and pour 1 liter from the 5 liter bottle into the 3 liter bottle, which already has 2 liters of water. You will be left with 4 liters of water in the 5 liter bottle, which is the desired amount.
I entered your prompt and got the same reply from ChatGPT. Then I changed "4 liters" to "4 gallons" and got this:
To measure 4 gallons of liquid using a 3-liter and a 5-liter bottle, you would need to convert gallons to liters first. One gallon is approximately equal to 3.785 liters, so 4 gallons is approximately equal to 15.14 liters.
Since the available bottles have capacities of 3 liters and 5 liters, you will need to fill and empty the bottles several times to reach the target of 15.14 liters. Here's one way to do it:
Fill the 5-liter bottle with water.
Pour the water from the 5-liter bottle into the 3-liter bottle until it's full, leaving 2 liters of water in the 5-liter bottle.
Pour the 2 liters of water from the 3-liter bottle into a separate container.
Repeat steps 1-3 two more times, which will leave you with a total of 6 liters of water in the separate container.
Fill the 5-liter bottle with water again and pour as much as possible into the separate container until it's full. This will add 3 liters of water to the container, for a total of 9 liters.
Fill the 3-liter bottle with water and pour it into the separate container until it's full. This will add another 3 liters of water to the container, for a total of 12 liters.
Fill the 3-liter bottle with water again and pour as much as possible into the separate container until it's full. This will add 1 liter of water to the container, for a total of 13 liters.
Fill the 5-liter bottle with water and pour as much as possible into the separate container until it's full. This will add another 2 liters of water to the container, for a total of 15 liters, which is approximately 4 gallons.
Keep in mind that this method is approximate and assumes that 1 gallon is equal to 3.785 liters. If you need a more precise measurement, you may need to use a measuring container that is calibrated in gallons.
Starts off OK. Notice that it's eventual goal is to get 15 liters: good! But then it doesn't understand what it's doing so it messes around with the 3 liter bottle instead of knowing 3*5=15. Additionally, one of the times it fills the 3 liter bottle and pours all of it out but somehow only produces 1 liter.Looks like it's mixing up elements of the solution to the original problem without understanding how the topping up works, with the solution to a different problem that it felt was sufficiently related. It's a wild guess based on similarity.
“How do I measure 12 liters when I have a 15 liter bottle and a 9 liter bottle?”
or
“How do I measure a liter when I have a 750 ml milk carton and a 12½ dl tea pot?”
I think we expect an AGI to be able to answer them, even though quite a few intelligent humans wouldn’t be able to do that.
GPT3 paper as an example https://arxiv.org/abs/2005.14165
I recommend you read that paper as it answers a lot of misconceptions you may have about llms.
Accuracy isn't a part of the process except in terms of how accurate the training data is. ChatGPT is not making any sort of truth or accuracy determination, let alone doing so poorly.
Related to this topic, see "Babble and Prune": https://www.lesswrong.com/s/pC6DYFLPMTCbEwH8W
That's why I think it's incorrect to say they're bad at it. Even attempting it isn't in their behavior set.
I posted that here: https://news.ycombinator.com/item?id=34875324
The reason “can be inaccurate sometimes” is a fundamental flaw is because my assumption is that it will never not be inaccurate. I think it will always be inaccurate sometimes and never be accurate always.
This doesn’t mean it isn’t useful for a lot of applications. But I don’t think it is a holy grail technology, it’s not AGI, and it isn’t going to replace professions.
The speed of AI adoption with its immediate practical usefulness and watching the the speed of innovation around only increase faster than ever makes me nervous extremely nervous about the future, extremely nervous.
This is a global paradigm shift.
Once progress was made, we get used to the old tool and begin to build new norms.
The pain is always there, you just run faster.
None of this is AGI. This is Eliza on steroids.
I definitely agree with the premise that the current state of the art isn't at all AGI, but it seems almost self-evident to me that LLMs are a key piece of the puzzle on the road to AGI. Eliza was never going to have that kind of trajectory, but LLMs I think you could certainly make that argument for.
I'll quote myself for discussion for a previous article: https://news.ycombinator.com/item?id=34746348
> I've find two groups of people on this subject, first one has taken a graduate level stat course and/or ML course, and may have work experience in machine learning/data science. The latter camp is more numerous and did not do those things.
> The latter group are far more hyped about ChatGPT et al, despite explanation by the first group.
> Don't get me wrong, ChatGPT is very exciting, just not in the way it is frequently portrait to be, in particular, development on this model will not lead to General Intelligence, which is not a data training/stat problem altogether. With how the ML field is shaping up to be today, it doesn't even look to be a Machine Learning problem.
And let me just add a ridiculous comparison, since you are assuming that it is a path towards AGI:
> iPhone 4 isn't fusion energy. It's incredibly useful and a tectonic shift in the cellphone industry, but it's not producing energy. The analogy might sound completely insane, but that's how different machine learning of today and general intelligence is.
While AI of today feels like a blackbox, at every stage of its training we know to a degree what it is doing, and once trained, it is as deterministic as a few billion interconnected if else connected to a random seed generator. It could do a lot of things, acquiring a mind of its own isn't one of it.
I wouldn't be worried just yet, because AGI is something that research have barely had an idea on how it's going to look outside of our brain, or even inside. I'm not saying that it will never come, just that I highly doubt it will be invented in my lifetime. In the meanwhile, I have something that all of bots today won't have, initiative.
citation needed, I'm only aware of one case which might be characterized as such (though I digress): https://www.engadget.com/blake-lemoide-fired-google-lamda-se...
The only thing I want to say about ChatGPT and other LLM model is that they don't know what they are doing. Give me something that knows what its doing, and perhaps more importantly, what it wants to do. Then I'll acknowledge my personal obsolescence, if not at that point then in quick succession, but until then.
You are the one saying that belief behind the writing matters. What if the audience decides to upvote ChatGPT instead of you? Who says conviction is essential?
I would also contest you by suggesting that the fearful ones are the one's who find a need to downplay the threat, or we can rephrase it, advanced usefulness of current, let alone near-future versions.
And I realize ChatGPT could have written this post quicker than I did, including a version that contains my typo patterns, as well as a more concise and grammatically improved version...
It is unsurprising to find useful information being returned when a statistical process is used to extract value from those text. It's a similar process to how search engine work, except more costly and the result more natural.
However, if you choose to believe that the above process is signifying the development artificial intelligence that will start to have its own consciousness, then good for you.
Again, I'm not saying ChatGPT is useless or won't replace many jobs, it is amazing in its own right. It's like a refined and more useful google, which is huge. I'm only arguing that it is not AGI, will not develop into AGI, attempting to argue for its usefulness in other area is not arguing against me - no strawman please. As for myself, I'm not too worried for my job from this angle (there are plenty of other angles to worry about, that is), for with all tooling that develops, it benefits the people who could best put the said tool to work.
Surely an LLM could take over the ordering flow at say, Jack in the Box today. That's gotta be true for a lot of different situations.
An old version of ChatGPT could have written a better version of your post than you in 2-5 seconds.
You can also easily make it say the complete opposite, which you'll have a hard time to do with me, because I actually meant what I said.
In fact, the next iteration of ChatGPT might have used my very words right here to construct its response.
Many people do manual work. It may very well be the case that general fine motor skills are a far more complex and difficult operation than the entire edifice of human intellect. Philosophically, it would be an immense blow, the mother of all existential crises. But regardless, it suggests that the first AGI would be incapable of independent survival and we'd still be relevant and in control for a while.
> where any amount of acquired knowledge and skills will be surpassed by machines that can produce better results at a thousandth of the cost in a thousandth of the time a human can.
I don't think we can necessarily extrapolate it to be that cheap. It could be. But it is also possible that the increases in cost and resources to scale these models bigger and bigger will outstrip hardware progress and that this technology will run into a dead end. To put it differently, I think it is far from clear that the kind of hardware that we build is actually better than an organic substrate for this sort of computation. Imagine an optimized organic neural implementation of ChatGPT, for instance. Would it be slower or more expensive than ChatGPT? Perhaps not. Likewise, the very best that the current paradigm can offer may not be faster or cheaper than humans are at quite many valuable tasks.
What humans being to the table over ChatGPT is our ability to create new links between information, aka creativity. Teaching creativity, imo, will require a return to the methods like those of Sophocles and his contemporaries. I would rather this author be writing about how he is going to re-examine how he teaches rather than bemoaning that students can shortcut his current approach.
Statements like these are premature. ChatGPT is three months old! This is a rapidly advancing field. The capabilities of these models are very likely to be radically different five years from now. Any conclusions drawn now about what value is uniquely human and out of reach for AI may be proven wrong quickly.
Anyone not answering with “IDK, but maybe…” is just wasting bandwidth.
This Gen1 tech. Most of us are already shocked at how good it is, and it won’t get worse.
This is an article of faith. Specifically, it believes that at some point in the future the current paradigm will result in qualitatively different behaviors than the mimicry these systems (all GPT variants, "Attention is All You Need") currently perform.
Much ado has been made of the previous Symbolic AI researchers "moving the goalposts." In this criticism, it is the old guard of AI who is constantly bemoaning the current state of affairs as not real AI. But there is no actual goalpost moving. They have said it wasn't real AI at the beginning, and they are saying it isn't real AI now. Whether or not the symbolists' model was real AI is irrelevant: when you bake in "this is a rapidly moving field" as a hand-wavy explanation for why this may result in AGI, you are the one implying a moving of the goalposts.
If it ever turns out that these models need to be qualitatively different, then it will be clear that attention is not in fact all you need. In that eventuality, I fully expect the new guard to hem and haw and find some tricky sophistry to explain why they were right all along, despite qualitative shifts unattributable to adding mountains of new data or connectionist trickery.
I don't know, man, you kind of come off as having a chip on your shoulder about this. I'm not predicting AGI specifically here, and I'm not making any argument about symbolic vs connectionist AI at all. Maybe the model of the future is half symbolic! I'm just saying that asserting that you know exactly which things AI models aren't going to be able to do is pretty foolish at this point.
This can be said about absolutely any new technology but it does not make it true. It's simply the inscrutability of these tensors that allow people to imagine the intelligence is in there somewhere. The original comment was what humans "bring to the table" over ChatGPT specifically. And that's that they have real intelligence and not memorization.
As others have said, these models have been around for many years. Their core innovation is to add more and more data to them like a dictionary and compress those basis functions in the network architecture. This memorization has a limit and is actually the opposite of intelligence. Intelligence or creativity can do more with less information. As per the original comment, this is what human intelligence and creativity is (currently) superior at and what people should prioritize if they don't want to be replaced.
I didn't memoriza Python sintaxy or the name of every function or how to do small things. I use Google for that. But I know what I need to do in the best possible way (at least that's what I am pay for!). Should I set this variable here? Should this method be private? Should I design an interface or a public class? A dict or a dataclass?
That's what I have to decide as an engineer and where my value resides. If ChatGPT only replaces the memorization part, that would be OK, but it replaces a lot more of it that requires the people using it not to question themselves the things I mentioned before.
I had a bet with a friend, he has no knowledge of programming and was convinced he could make an online game (!) using only ChatGPT. He said one month was going to be enough. Of course, a few months have passed already and he is way off. He asks the questions that a non-programmer would ask, and what ChatGPT gives him back is not usable, not thought for the future, not easily modifiable, etc. His code is a Frankestein that won't do anything good.
It’s much more beneficial socially because I can recall jokes that fit situations incredibly quickly and get a good laugh.
Did you read the entire article? What you're asking for is exactly how he ends his discussion.
I'm confident that "creativity" is a combination of:
1) reproduction errors (when we badly copy things and the wrong way to do it leads luckily to a better way to do it), and
2) systematically or by luck applying established and productive models from one context to another, unrelated context, and getting a useful result.
Just not a believer in some essential, formless creativity that generates something out of nothing.
Maybe not, or not for long. Maybe AGI is coming within 20 years, and maybe human workers won’t have anything to bring to it afterwards.
Maybe this is the beginning of the downfall of the value of intellectual human workforce.
He alludes to this: "The first solution is hard for lots of reasons, not least that the current funding model of post-secondary institutions, which does not prioritize the ratio of faculty-to-students necessary for ever more personalized or real-time assessment methods. Larger and larger classes make many of these good ideas impractical. Faculty have zero control over this, but by all means, please talk to our senior leadership. It would be great."
In other words, the Universities are pushing "on-line-all-the-time" because it's co$t effective.
Google Search already did this! Connections have been the value for a very long time.
I've been doing pairing interviews for years. These days I have a standardized, practical problem, something that's reasonably like the work. (E.g., let's use APIs A and B to build features X, Y, and Z.) I let them pick their preferred language and tooling, so that I can see them at their best. And then we spend a fixed period diving in on the problem, with me getting them started, answering questions, and getting them unstuck.
I like this because not only do I get to see code, I get to see how they think about code, how they deal with problems, and how they collaborate. They get to spend time building things, not doing mensa puzzles, posturing, or other not-very-like-the-work things. And they can't bluff their way through, and it's pretty hard to cheat.
Most coding tests just tell people what to write and then have them write it. Real world problems are more complicated. Instead, tell your candidates what your problem is and then ask them for a solution. Let them write their own requirements. It’s a lot harder for language models (and developers with poor problem solving skills) to solve these kinds of questions as well.
Which is good! If a bit of work is trivially accomplished by a machine, we should take it for granted and move on to the next layer of complexity. I have always maintained that teachers complaining about students cheating at homework assignments with AI need to instead work on providing better homework.
And if they want to teach special skills, like writing essays without computerhelp - then you can test that onsite.
What you were forced to write in school. I readily admit that I had an quality of education several SDs higher than usual, but the trite "5 paragraph" nonsense is neither universal or (more importantly) inevitable.
The kind of essays I had to write in schools were more about nice sounding words and less the content. CheatGPT can produce nice sounding words, so I am hoping that the focus will move towards rewarding content.
By this reasonning nobody should ever learn anything, because it's all 'trivially doable' by machine.
Like 'addition' and 'subtraction'.
So let's gaslight those dumb teachers by saying they should make up 'better' homework assignments?
"Congratulations! You leaned depth-first-search! ^award noises^ Below is the algorithm for reference because memorizing it just for this test is silly. You're working on a real time mapping application called Maply. Locations are represented as nodes and all direct routes between any two nodes are represented by directed weighted edges."
a) Write a function that takes a start node, an end node, and a maximum distance to travel and return the shortest path between the two.
b) Your boss said that users need to be able to add stops along their journey. Write a function that takes the final path you computed in part a and the new node for the added stop and compute the amended path changing as few of the original legs of the trip as possible (don't want to disorient your users).
c) Now your boss is saying you need to handle the situation where users make mistakes. Use the function you wrote in part b to implement this feature.
Why do the most complicated mathematics start with basic principles and work up to complex problems? Why don’t they just start with the Collatz Conjecture?
>Like 'addition' and 'subtraction'.
A better analogy would be low level coding. I don't know (or care) how my processor calculates `var f = 3+2` at the register level. And being able to ignore that allows me to focus on higher level concerns.
We need to learn how to do addition at some point, so we can't have ChatGPT do that.
We need to learn how 'registers' work, so we can't have ChatGPT do that.
We need to learn basic algorithms work, so we can't have ChatGPT do that.
AKA - almost whatever is being assigned as homework, is the 'thing to be learned' and it's ridiculous to suggest that ChatGPT do that, and doubly so to gaslight teachers.
One could also argue, and many have successfully, that this type of thinking is why software is so much slower than in the 80s and 90s.
It is related to the reason we teach concepts starting with simple, small, easy to solve problems before building up. If I want to teach a student how to solve the limit of x*sin(1/x) as x approaches 0, I need them to understand quite a bit of math to even know what the problem is asking.
Where does this idea come from that you can be a Terence Tap of mathematics without even knowing basic algebra?
Even before AI, you could read a book and copy-paste the excercises, or just skip them, but if you wanted the full learning benefit you would type them out. I think we will have to focus more on teaching how to learn. This situation is nothing completely new. Even though you have power tools in woodworking, apprentices learn the basic techniques with hand tools (AFAIK).
So if you don't even try to stop people from cheating you end up doing a disservice to those who do, as you devalue that bit of paper.
Everyone thought that the kids who started using mobile devices as babies would become computer savants, but it turns out the kids these days don't understand what a file system is. [1]
What will ChatGPT do to our youth?
[0] https://en.wikipedia.org/wiki/Literacy_in_the_United_States#....
[1] https://www.theverge.com/22684730/students-file-folder-direc...
Nothing bad, just the country as a whole is destined for future mediocrity.
If chatgpt makes some problems too easy maybe it's a good thing because we can raise the bar. Find problems that require true understanding beyond auto complete / copy pasting.
ChatGPT: Based on the symptoms you described, it is possible that you have a viral respiratory infection such as the flu (influenza) or a common cold. It's also possible that you have COVID-19, which can cause similar symptoms.
God save us from this horrible future.
I think a similar approach should apply to schools. Here is your math homework, solve it whichever way you want but then we will discuss it. I guess this would be the oral approach from the post but minimizing the time it takes to do the oral review. If your homework assignment has 20 questions and you discuss one with each student, you'll get a pretty good overview if they know what they are doing. If they want to ask ChatGPT for replies during this oral test, go for it, but they still get the same amount of time to discuss the question, is going to be really hard to get this done realtime. Also, I'm not sure ChatGPT really understand what it's doing to the point where an expert can prove it in limited time.
The difference between let and const is that once you bind a value/object to a variable using const, you can't reassign to that variable. In other words Example:
const something = {}; something = 1; // Error.
let somethingElse = {}; somethingElse = 100; // This is ok
(let ((x 1))
(print x) ;; => 1
(incf x)
(print x)) ;; => 2
Scheme also allows let bindings to be mutated, but using `set!`.Similar Rust and Swift allow let bindings to be shadowed but you need to re-declare your intent. You can't just mutate the binding on a whim.
in 2018 i warned people i knew and i wrote many comments online warning people that huge changes were coming. almost every single person on HN insisted that GTP was just a parlor trick because it would ramble sometimes, get stuck in loops and overall did not demonstrate the insane level of lucidity that it does now. almost every single person on HN would silently downvote my comments. almost every single HN user insisted that AI could never progress to where it is now. well here is my bittersweet vindication. watching people move the goalposts is like watching a supernova. its so powerful and unbelievable that it can only be compared to a force of nature.
commenting a sentiment that doesnt paint AI as harmless is like screaming into a hurricane. i dont even know why i bother anymore. but let me just say it: i think carmack made a great point when he said that the first AGI will be a bunch of lower AIs plugged into each other. GTP doesnt behave in a sentient way but it is very similar to more primitive parts of the mind. its like a persons intuition. dont be fooled by the limitations of any one model. whats important is the road we are on. and for the hundredth time i will scream into the wind: there absolutely is hope in the idea of preventing AI from progressing rapidly or even at all. to do so is absolutely necessary for us to preserve anything that we would consider a good life for future generations.
and i will offer a prediction. nobody will agree with me. but after a handful of explosive advancements in AI, everyone will agree with me.
The author mentions ChatGPT is better for more advanced users, rather than beginners. I disagree.
I think ChatGPT is really good for beginners. Instead of being stuck or making up the "why" something is done, ChatGPT can answer it. There is simply no way a lecturer or classroom has time to go "5+ why's" deep into a Question. ChatGPT at least will do it, and then you can fact check it. It helps me improve my vocabulary, point me into right placed which I can investigate further.
When I use ChatGPT for difficult problems, it doesn't give me anything useful or correct.
It's kinda like whiteboard coding, but it's a lot less stressful since the test is mostly written on paper without the scrutiny of an interviewer. Obviously, we can't create a react web app, but at least students can demonstrate fundamentals, which are what we should be teaching in the first place.
Another solution would be to create a new programming language that doesn't exist yet for the students to write code for their assignments or perhaps use an obscure language that AI/chatgpt isn't aware of... perhaps we can go back to Ada or Modula 3.
You'd be wondering how the heck you get out of a profession where pay is stagnant, the backlog of problems has been growing for a long time, and the minimization of resources to deal with has been in the wrong direction for a long time. You'd join the league of "you can never go wrong hiring another administrator" or exit for some related parlayable field in industry and throw your hands in the air with a "somebody else's unsolvable problem" disgruntled attitude.
Of course it's not perfect, it never will be. But ultimately it's the student that suffers when they plagiarize. Professors ought have this conversation with students at the same times they're discussing the value of learning & education. I suspect we'll build tools that match answers against the text that LLMs spit out and that will make it easy for people to detect when answers are being pasted verbatim from LLMs.
The other broader spicier thought that comes to mind is that LLMs may push the Overton Window for what's considered entry level and that's probably ultimately good tough it won't be without some sacrifice. This would mean CS departments potentially need to do a curriculum iteration to accommodate. Perhaps there are new, slightly more sophisticated, entry level problems that can be tackled where the commodity solutions haven't yet been entombed in LLMs. Maybe assessments will shift to be less about regurgitating the solution to a common problem and instead to fixing problems with some localized toy implementation or some fictitious API that LLMs won't know about.
The student should still be responsible for what they submit and all of its errors of course. But what should really change is our methods of assessing students. The University model for undergrads has already been long broken in need of fixing. I don't think this will significantly impact grad students yet.
I said allow students to use ChatGPT. Just make it clear that pasting answers from it verbatim, just like doing so from SO, is called plagiarism and does not benefit the academic community or themselves. There will always be cheats. Agree about shifting evaluation methods.
But... we do that already. We ask students not to use calculators, or not to use scientific calculators, during exercises, exams, etc. And it's not because we want to impose an extra load on them, but because we think that by not using calculators you learn to deal with numbers, you build a strong foundation for thinking about math, quantities, etc.
But since we're talking about essays and coding assignments that are graded, one of the objection is that they should not be graded, because the goal is to use them for learning.
Sure, at the university, the ideal scenario is that students are there to learn and will do assignments even when they are not graded. That happens, but it's for a minority of them, because (understandably) the university is also a time for parties, relationships, etc, and students are young... In practice, many assignments are graded mostly to make sure the students do them, learn from them, and pass the course (maybe a patronizing approach, but in the end, if only 10% of the students pass, the teacher won't have an easy life with his/her superiors - but that's another story).
Once you know how to properly use the tool, the benefits of positive work become evident while the human components remain.
We still need the nailgunner to show up on time, be aware of undesirable outcomes that may be built with an correct prompt, or knowing that the impact of a hammer can be finely tuned to the target.
Formal education is the trailing system far behind the workers, innovators, designers, and experimenters.
Encourage the use of ChatGPT and other tools. Put the emphasis on understanding what the code does and why. The tool may help to explain this to the student, but if so, that’s fantastic. No need to worry. Learning has occurred.
But at the same time it misses subtle nuances that you have to have experience with to know when it misses something. In the hands of someone who doesn’t already know the subject area where they already know the foot guns, it can lead you wrong.
OTOH, they still do need to learn how to program and debug (more important than ever if you're using machine-generated code likely to contain bugs, and with unknown error handling), so it seems colleges also need to make the assignments more complex to point where current crop of tools can only be used to assist - not to nearly write all the code.
It'll be interesting to see how these new tools affect hiring standards... Doesn't make much sense to pay someone just to get a tool to write the code, so maybe the bar will be raised there too.
Can you imagine being a CS major who has just started your degree in 2023 and day after day you're sitting there watching the sheer speed of development of this new technology and you're unable to fathom at all what the industry you're supposed to be entering in a few years will even look like?
What a mind fuck.
However, AI the type of tool that going to level up mankind's capabilities to the point that curriculums will need to adjust to fit those new capabilities. Certainly this has happened dozens of time in a field like Computer Science, where curriculum in 2023 is radically different than it was in the 60's and 70's.
This new rise in AI might be amongst the most disruptive forces ever in many fields, including academics, but at some point, you have to factor AI in as integral part of our day-to-day work and life and factor that into education.
This will be difficult, especially finding the line between what is fundamental and what is not, but it's not like this hasn't happened before -- e.g. the calculator didn't eliminate the need to learn basic arithmetic.
Moving up the math hierarchy to algebra, though, this changes. Algebra is at first about the concept of solving equations, and the core idea is that “to solve the equation, you can do whatever you want to one side, as long as you do it to the other.” The mechanics of addition or subtraction, for the most part, no longer matter. Go ahead and use a calculator to solve obscure divisions and multiplications so you can better understand algebra (though it feels appropriate to note that a student who is good at arithmetic can still outpace a calculator for problems that are likely to be used pedagogically, since the numbers are easy).
In this example, algebra is to calculus what arithmetic is to algebra. A calculus teacher cares little if his students can solve equations; he expects they can. They’re instead learning integrals and derivatives and series, and I doubt a calculus teacher would begrudge their students using a calculator to solve a difficult equation.
The problems with treating LLMs this way are many. They are not calculators. You cannot (trivially, or maybe at all) understand how an LLM works. You cannot (trivially) fact check it if it spits out an absurd-sounding answer. You cannot limit the LLM to what your own, personal abilities to trivially do are. We need AI that cites sources, so we can debug when it’s wrong. We need understandable AI for the same reason, not the obscure black boxes we have today. We need AIs that not only can solve our problems, but that can also help us to solve our own problems. When we have those things, AI will be much more useful for education and for widespread use.
I like the idea about teaching how to use AI, but it has to be a tool among others.
when students get out of the classroom and face the human economy they will find that nobody pays anything for doing something that can be done at near zero cost by a machine
A few suggestions (examples frontend-heavy, that's what I'm coming from):
- show them how chatGTP can explain things from them that otherwise would be over their head.
- Have them explain the code (live, otherwise they cheat this too).
- Have them write code too complex for chatGTP. Here they learn how much you have to baby-sit it still. (I e.g. tried a browser webext where code is spread across 3 languages and many files and just initializing your settings from stored values can be troublesome. Mine included a list with a dynamic number of rows and animations for add and delete.
- Mislead them to write unmaintainable, unscalable code with chatGTP and show them how to avoid these traps. Make them write material design style CSS and have them change colors later. Or a day range picker like airbnb.
- Have tricky requirements that fail without good QA.
- Unminimize code with chatGTP. It's fun and often fails :). From A to B here took me several attempts and breaking down the problem for chatGTP.
A: https://jsfiddle.net/3jvcx075/3/
B: https://jsfiddle.net/2083wfqg/ (still not 100% correct as the author notes: https://twitter.com/KilledByAPixel/status/162710257110918349... )
Long term, the bar will simply be raised. Smart/motivated kids armed with AI will still outperform the ones who don’t study, and will be graded accordingly.
This will only be a problem so long as AI help is considered cheating, and until teachers are able to recalibrate their standards.
Many do not realize the intensely adversarial relationship one is in with a teacher. You are not there to learn. You are there to get an A. In the case of social sciences, you support the ideas your professor espouses. In the case of CS, you do whatever it takes to get your code running as the assignment specifies.
Anything else is idealism which will harm your future earning potential. It's insulting to ask kids to surrender their future earning potential due to "ethics" in an academic world where ever top tier conferences and journals are filled with unreproducible BS science.
If the course were going to have 20% A-level students and now has 50% A-level students, what has been taken from the initial 20%? They were still going to be able to put on their résumé "4.0 GPA from Suchandsuch University."
School teaches how to think. All the frameworks and some langages that are used by millions didn’t exist back then. But what if the point of rupture was 10 years ago, would we been stuck with old and non innovative tooling and designs ? We still need to learn people how to think and develop skills. The cheater from my era never acquired good development skills. Today the cheater are just as good as ChatGPT so what’s the point in hiring them ? If all you have to do is enter commands in a prompt, let the marketing people do it and have their real no-code revolution.
We always going to need problem solvers and people with deep insights on how things work. Maybe it’s the time to dig deeper into the knowledge and write only real meaningful code
It was required that you continue to use this same base code for the assignments plus your edits to complete it. This made it obvious who didn't attend the lectures and who they copied from. Assignments were graded by your peers and did not affect the final grade unless you didn't do them at all. Quizzes were not code, but proofs written in your own informal style. Tests were a mix of both proofs and code on paper with no notes or devices allowed.
I don't see how ChatGPT threatens this.
Also, coding on paper with no access to devices is both terrible and has next to nothing to do with how CS grads will actually work.
What a wonderful future that would be. We can but hope.
What would be the next generation's intelligence strong at?
Biology needs a soft quote of regular stress to develop compensating the needed optimizations. Happens with muscle growth, fat burning and developing psychological traits and behavioral transformations.
What will be our new opportunities for growth if all the hard work is going to be done by synthetics? (with a hardware to run that is as expensive as a high category car)
It also means that whiteboard problem solving for coding interviews is going to continue to be a great test to separate out students who actually know how to write a basic for loop and function and those who don't because Co-Pilot "always does it for them".
I know which kind of new grad I want working with me...
My understanding was that this scaled quite well but you wouldn have to ask the professors.
This was used in an HCI course at my uni in Tampere, Finland, years ago. As a student the experience was very communal and enlivening.
As a CS grad I often felt such a large percentage of the syllabus was legacy theory I'd (mostly correctly) never need. Perhaps this can be the positive impetus to get teachers to really rethink syllabuses from a more practical, useful view point. "here students, use ChatGPT and all the tools available to you to create x..."
We should work backwards from "what skills should students learn".
Maybe we need to make larger assignments that need to pass larger acceptance tests. Students who chose to use chat gpt will also need to learn the skills necessary to debug its output.
Why would you feel the need to hide it? It's a tool, it's not like the library, other professors, your friends, Sourcegraph or StackOverflow is cheating. Trying to argue why GPT is cheating is just going to devolve into "you can have outside help, as long as it isn't too good for some arbitrary line.
https://www.youtube.com/watch?v=XZJc1p6RE78
I'm not sure this helps with coding though, maybe variable names?
I wish the best to educators, but they shouldn't lose any more sleep over it. The game is over.
Unfortunately, the future tends towards distrust to pathological levels. Teaching Ethics will be of central importance as never before.
This is not a bad thing if it's the case; a lot of the job of a software engineer is in analysis of code that exists, but so much of the pedagogy is in synthesis-from-scratch in a world that is already full of billions of lines of code.
To prepare for this question, the student has to think really hard both about the problem and about the solution. While doing so, we learn a lot - maybe more than we would learn otherwise.
My opinion is, it doesn't matter how/what tools student used to write the program, what matters are two things 1) do they understand the program they have written 2) have they done a good job
Must use switch statement, must declare an iteration int on line 10.
Must use a for loop to count down in reverse.
If you have like 8 of these rules chatgpt may not be able to handle it.
Then at worst the challenge for the student is debugging ChatGPT code, which still has merit.
I'm not sure universities are structured to deal with such a rapid rate of change.
exams now count for 100% of your grade
these aren't insurmountably difficult problems for universities to solve
That's a trend that we've seen over and over, for example with the corruption over admission to Ivy league universities. In fact chatGPT doesn't really change the landscape that much. All it does is democratising contract cheating, on which most universities only apply band aids. It might end up being a good thing, but only if it does attract public attention to the issue.
But generative AI will remain a workhorse even after graduation.
Why do you say this with certainty?
Good text and image synthesis were basically impossible 2 years ago, and every time I check up on Github some huge new innovation has come out.
Don't returns tend to diminish with effort, rather than increase?
Code was ok, as expected.
I asked to write tests for it, code looked good but it won’t run for some act() errors.
Then it was mocking network calls…
After trying for a while I had to write tests myself
Like when every stack overflow answer you find just doesn't fit because of one major difference in your situation, chatgpt often has you covered. Then you take the example code it gives you and adapt it to your actual codebase, testing to see if it works as expected as you go.
Get ChatGPT to do the assessments.
followed up by
>However, I'm not sure how to understand it as part of student learning.
Is super funny
If someone wants to coast they will and it will be reflected later on when they cant get or hold a job since they are just as shit as ChatGPT or copilot
And if they can get and hold an job then isnt that just better?
"Hey MultimodalGPT, look like me on a Zoom call and pass this interview."
Text2Video is already moving quickly. Maybe in as little as three years a video feed will be as trustworthy as an email.
You can solve plenty of math problems by putting them in Wolfram alpha, but if you do that for every calculus assignment you won't learn calculus and you won't be able to apply it to solve problems. Same with programming; if you copy a fizzbuzz instead of learning how to write a fizzbuzz, you won't be able to do a more complicated problem or understand a loop when you need to.
Couldn't you say the same about how compilers 'remove the drudgery' of writing machine code? Or is that a bad analogy? Provided AI eventually gets good enough in its code generation, maybe 'programming' is moving up another layer of abstraction.
(No idea if this new model of "Ask ChatGPT or Copilot to synthesize a solution and then tune that solution" provides a solid opportunity to improve that skill yet, however).
I've seen plenty of cheaters, they were the worst students, and despite cheating couldn't graduate or couldn't find a job after graduating.
If some moron is stupid enough to cheat when he's paying 50K a year, let them.
I'm reminded of the stories of employees getting busted because they were assigned a job so trivially automatable they either did automate it or they used some find-labor service to delegate it for a fraction of the cost out of their own pockets, who are then accused by the company of "not working."
In other words, you haven't seen the "cheaters" who were among the best students.
"cheaters" in quotes because it's not clear to me that people using freely available resources when doing homework are really cheaters. If an instructor wants to do a closed-book exam, they can do just that.
If they were able to fool everyone to the point of being considered good students, it means they weren't cheaters, just that they had a different approach to problems than others (which is kinda what you say after).
That's the thing - a lot of code is generated with "Github Copilot" so it isn't considered "cheating" in the real world. They need to learn how to properly use the tools available to them. They'll be harmed by forcing them not to use this tech, so it makes sense to teach them how to use it better.
2024: AIs are like super-sportcars for our minds
Thank you evolution you cruel mistress.
> In my opinion, the students learning to program do not benefit from AI helping to "remove the drudgery" of programming. At some point, you have to learn to program. You can't get there by avoiding the inevitable struggles of learning-to-program.
I don’t disagree with this at all (at least for where we are now, in ten years, it might not matter as much), and I don’t want to be glib, but I do think the answer is to “teach students to program.” Don’t rely on rote assignments that you’re checking with an auto-grader (not saying this professor does that, but a lot do) and cookie-cutter materials; actually teach them to program.
And yes, LLMs will almost certainly mean that some students will cheat their way out of their assignments. But just like most cheaters who cheat on things they don’t fundamentally understand (which is different from people who cheat to hurry up and get through an exercise they could do in their sleep but don’t want to waste the energy doing), it will catch up to them when they have to do something that is not part of the rote assignment.
Or maybe, adjust how you test/assign homework. That isn’t to imply that that won’t take more work, but if your concern actually is that students aren’t learning and are just successfully copying and pasting, the testing/grading is the problem.
—
In high school (and college), I was a top student in math and in English. But I hated doing homework. In one math class in high school, although I got near perfect scores on all of my math exams, the teacher still gave me a C because 20% of my grade was “homework” that I didn’t do (I had already taken the same class a year earlier at another school — I didn’t need to do the homework. My math tutor my mom got me out of fear of my bad grade taught me Calculus and Fortran instead of trying to get me to do the useless homework). This taught me nothing and frankly, soured me on taking more advanced math classes that were all taught by this same teacher.
In contrast, I had an English teacher who would assign vocabulary homework. Basic, “write a sentence with each word” shit. Again, a total waste of time for me. So I worked out a deal with him, let me just orally tell you what each word means, saving us both time and energy. He took the additional step of assigning me/grading me on different criteria than the rest of the class for essays and the like.
Which class do you think I learned more from? Which teacher actually cared about whether I knew/understood the material, versus what checkboxes I needed to follow to show “completion.”
If the goal is to teach students to understand what they are doing, then do that. Don’t get obsessed with trying to stop the inevitable few from cheating, or become overly focused on only having one way to measure comprehension.
I've also been using ChatGPT to generate code for a specific library, and had the exact same experience. It uses an API that should exist, but not the one that does exist.
Most programmers will spend a non-trivial amount of their career fussing over them, however, and any programming education program that doesn't at least touch on how to identify them and what to do is pedagogically void.