I believe AI is basically an amplifier of bad and good. I’m cynical about the world and assume it will be used more for bad than good, but I don’t doubt some of the best people in every field will be using AI to amplify their work in good ways.
I believe AI is basically an amplifier of bad and good. I’m cynical about the world and assume it will be used more for bad than good, but I don’t doubt some of the best people in every field will be using AI to amplify their work in good ways.
"At the other end of the distribution, AI students who spend more than 65 minutes on their homework receive homework and exam scores similar to those of non-AI students, suggesting that these students do not use generative AI for homework assignments. However, this group consists entirely of students who adopted generative AI no more than Öve months. Six months after adoption, no AI student spends more than 65 minutes completing their homework (see Figure A5). This is consistent with the gradual process of learning how to use AI tools. It also suggests that AI crowds out the highest level of e§ort."
"Interestingly, in the range of 50-65 minutes, the median and the interquartile range of exam scores of AI and non-AI students are similar. This implies that, in the range where AI students and non-AI students have overlapping homework times, students who spend the same amount of time completing homework on average receive similar exam scores."
"This pattern shows that students who spend the same amount of time on homework learn similarly, with or without generative AI. In other words, generative AI reduces time spent learning for the majority of AI students but not learning efficiency for those who spend the same time studying as the non-AI students."
they're not designed to measure general aptitude, or function as admissions criteria, or screen for job applications, or any other numerous things they are used for.
there can be many questions of pedagogy. one of them is, what do our exams measure and how do we use them? professors who say, "My exam is designed to measure who studies, not be used for all these other purposes that they are actually used for" - I don't buy it. It's the same as late night comedians saying they are not responsible for solutions, even when spending 90% of their air time making political jokes.
THIS is the pedagogical issue, that pedagogy has NEVER caught up with the scope of responsibilities. This is acute in STEM - I mean, the humanities departments are generally pretty well run, all things considered, in this regard. Generative AI is accelerating that pre-existing crisis.
Huh? They're designed to measure how much you know. They can't see how much you study, nor would they have reason to be interested.
At the end of the day though what matters is what you know. Furthermore, if it's a serious subject, it shouldn't matter whether you learned it from this teacher or from another school and teacher, as long as your knowledge is correct. Knowing the idiosyncracies of this particular teacher should not factor into the grade. A serious subject can be learned on one continent and examined on another. Bullshit courses are all about learning pet peeves and hobby horses of a particular teacher.
let's imagine a different study. we instead compare AI-users and non-users on a Wechsler (IQ-adjacent) test.
overall, it would be surprising if AI usage impacted your Wechsler scores. someone has done this study and the impact is quite quite small. BUT. do we care? We don't use Wechsler scores for admissions, we don't use them for jobs, we don't use them for... are you getting it now? A Wechsler family test is measuring something real, just like a university exam measures something. But what do we USE them for? Wechsler and a typical university exam are, in some senses, EQUALLY vague in terms of their fitness for purpose for answering a question like, "should we hire this guy?"
Like there is an association between IQ and earnings but it is actually surprisingly small! There is an association with math education and earnings and it is also surprisingly small. And consider how many people get by just fine without using a single piece of math education once they have finished school - like what if maximizing your earnings isn't all that it is about? Are you getting it now?
The issue isn't the AI usage. I can find tests that are immune to AI usage. The issue is using tests for things that they are not designed for. We pick and choose, for some subtle but nonetheless pervasive cultural reasons, which tests we use for which purpose, and very frequently, not because they are calibrated for the chosen purpose. This is coming from someone who scores very well on all these tests, and have kids, so I have a very strong incentive to buy into the status quo, and I'm telling you: academic testing has been fucked up for a long, long time.
Two weeks of ADHD-fueled research later, I concluded that academia is actively resistant to implementing assessment reform because it would expose the utter pointlessness of most of what happens in university classrooms.
The reality is that we have no idea what most university exams measure because they are ad hoc, written by amateurs (yes, most professors are untrained in pedagogical methods) with zero psychometric validity analysis.
This is a reasonable opinion after two semesters.
IMO, with more experience you should have modified your stance.
If all classes were pointless, students would leave at the same point they entered. Since we observe that is not the case, something is happening in class. Do more of that.
The explicitly stated goal of most courses is mastery of a well-defined set of domain-specific knowledge and skills. The only way to objectively determine if the course achieved that goal for a given student is comprehensive assessment. Bad assessment design, much like bad experiment design in science, is worse than useless—it actively misleads and confuses.
A well-designed, repeatable, reliable assessment is a prerequisite before one can even consider the effects of particular pedagogical methods inside or outside of the classroom itself.
"The negative learning effects are larger for students with higher initial achievement. The differences in the estimated full (6-10 month average) effects are substantial, with a 50% gap between the most negative effect (-24 percent) for the highest tercile and the least negative (-16 percent) for the lowest tercile. "
Not top 10% as you asked, but the closest to what you asked. My working hypothesis is that top performance is highly correlated with willingness to work hard, and AI decreases the motivation to work hard.
I think if you take a physics class where the student is intelligent and intrinsically motivated through their own interest (I admit this is rare) then AI probably helps.
That last part is key. Intrinsic motivation doesn't mean you pursue it outside normal bounds. My kid loves soccer, its her second favorite thing in the world, she has an absurdly high tolerance for physical discomfort while playing, but she doesn't play it at home. There's other things she's rather do, such as play with her toys.
When you move the bar to something even less interesting to most kids like science, you're going to have a pretty huge falloff. You're basically selecting for kids who choose to do it in their spare time. I know a lot of smart kids (I run a boyscout troop, my wife a girlscout troop, both with lots of high achievers), and none of them do this.
I'm confident it's an amplifier for people who know how learning works and already do a lot of it, successfully. However the level of "learning fluency" I'm talking about isn't reached for many until late college or grad school, and sometimes not at all. So I'm not surprised by the quoted results for 12-18 year olds.
The discussion is about "AI", so common sense is out the window. These people's professional reputations depend on addict-level "AI" usage remaining socially acceptable.
How would you prove/disprove this assumption without falling into a True Scotsman fallacy?
It would not prove that "the results depend on the operator". To prove this (and figure out the traits that make a proficient operator) you would either need a massive dataset to which you'd apply some machine learning to figure out the correlations, or at least you'd need a hypothesis on the traits you want to test.
Do left-handed people perform better with LLMs? Analytical thinkers? Dyslexic people?
It's not enough to just claim "some individuals will perform better with these tools" without any notion of who does, and whether it's possible to become one of these people, and what is the expected gain in this group.
Otherwise, sure, we could sell unsecured chainsaws as a tool for making ice sculptures and observe that "some people" are indeed more proficient at not getting their face cut off ; but if the tool is being marketed to schools, such a vague statement is not helpful in making a case for it.
My only point was I think it's possible to do. I personally think it's pretty clear. And there's nothing special about AI. In rural places, you could replace it with "access to books" or "access to a good tutor" and so on.
I also think you'd see the same effect with "motivation" (if you could test it), with performance on standardized tests, and with grades. AI will boost the better (by these metrics) students more.
It seems that the students who had the most skill and motivation from the start, have the most to lose.
You might say "it's good to learn research skills" and that's true to an extent, but tutors have always made people better students. And AI is a tutor you can message at any time, day or night, for free.
The good thing about a tutor is that he or she isn't always available and knows they won't be in the future, so they instill good habits and independence in you.
Hence the whole "AI is basically an amplifier of bad and good" argument made earlier. The ones who do this because they want to understand, don't need to cultivate this at all, it naturally happens with chatbots. They don't just ask for the answer to a question, but then dig into why it's like that and what not. But the ones that don't care, now have to do even less to get the fast answer without understanding.
I wasted several days trying to have AI teach me containers. I would have been much better off just reading the docs.
I wouldn’t be surprised if you did and still had the same issue, but if you relied on the training data recall alone, you definitely shortchanged yourself.
I’ve had decent results with tasking a model to run pre-research and summarization as a checklist, then have a second instance of the model(s) review and then develop my learning plan learn, versus the times I just asked Claude or ChatGPT to explain something to me “from memory.”
which makes it not a tutor. if it is true that tutors have always made people better students, then those tutors are definitely imposing some limits on how many answers they give you and requiring you to do some thinking. I think you could have premised your same argument by, "copying a smart kid's answers has always made people better students...."
First, there's a certain amount of baseline knowledge we'd like students to possess. Without a certain prerequisite amount of underlying information committed to memory, it gets far more difficult to achieve fluency in a topic.
But more than that, there's other skills we're trying to build: frustration tolerance, processing contradictory information, disciplined problem solving. You only really develop these skills through productive struggle. If you find a way to shortcut the productive struggle, students truggle.
> but tutors have always made people better students.
Sure. Bloom showed us that students taught with a combination of tutorial and mastery methods, one-on-one, outperform students in a normal classroom by roughly 2 sigma.
The paradox has always been-- why hasn't technology unlocked these gains for students in normal classrooms? If we could boost everyone's performance by this amount, it would be huge for society-- but society can't afford to teach everyone with tutorial methods.
Since the 1970s, we've invested in edtech towards trying to make this happen, but most of it has actually had net-negative effects as best as we can measure. AI, so far, looks to be much worse.
I think part of the answer is that a big part of what makes a conventional classroom work are social pressures. So far, it looks like AI (and edtech in general) does more to dismantle conventional pedagogy and to break down the social fabric of the classroom, than it has improved differentiation or unlocked this tutorial effect more broadly.
I was assuming this was a typo and was thinking about making a joke about it, but it does appear to be slang that fits the context:
https://www.urbandictionary.com/define.php?term=Truggle
> 1. the standard of perpetual intellectual failure made by an individual.
> Truggle; the standard defenintion of a person who is a failureat everything.
Was it a typo or did you actually mean this?
The answer is obvious: you're doing homework type activity (learning is not a goal), not exam type activity (learning is the goal).
Social stigma against just juices people to stick to socially desirable answers; no doubt a whole bunch who self reported as non-AI users actually used AI
"e§ort" == effort
"Öve" == not sure. Maybe "over five"?
The "slightly higher" performance is based on statistically insignificant samples (between 4 and 20 students, depending on the context, out of the total population of 26,000): https://bsky.app/profile/benjaminjriley.bsky.social/post/3mt...
Citation needed? I have no clue where you got this from. I hadn't even heard of it as a conjecture, let alone as something anyone accepted, let alone as generally accepted...
1. They don't do any homework.
2. All the in-class time is split between the teacher babysitting and playing social worker to problem students, and lecturing, with little to no opportunity to actually practice what they've learned?
I understand that some students don't have home environments that are conductive to doing homework well. I understand that some students are enrolled in five hours a day of extracurricular university-application-padding activities. I understand that some students have incredibly poor screen discipline and impulse control.
But I don't understand that anyone has magically figured out how to teach complicated things to students, and have it stick without them spending a lot of time practicing what they are learning.
As anyone who has tried to do something hard knows, the first step to being good at something is to spend a lot of time being pretty shit at it.
A student who has written and received feedback on 500,000 written words is going to be way better at writing than that same student who wrote 50,000, just like someone who has put 5,000 hours of focused practice into playing the piano is going to be better than my dumb ass, who has only put 100 hours in.
(If you found the solution to get good at stuff without practicing it, I'd love to get good at piano without putting any homework in on it.)
I'm pretty sure there is a way to fine-tune this [0] to auto-complete away any of your hesitations or mistakes :)
Art is similar. You can look up skills as necessary to try to make the piece you want. Adult art learners often just start with the type of art they want to make.
In both cases, the finished piece might not be the quality you want, but in neither case are you necessarily doing repetitive stuff to learn (playing scales over and over or sketching the same bits over and over). You can retry a piece or move on - the skills will still carry over.
None of these reflect homework in school. Now, I know I graduated school decades ago, but math homework was repetitive with seemingly no real-world application and absolutely no help if I needed it. I took a math course online some years later and it was much better: Instant feedback if I got the problem wrong and instant help to walk through the problem if I needed it - then a different but similar problem was given for the homework. This was actual practice in ways traditional homework wasn't. Homework in the traditional sense isn't practice - they are all miniature tests that affect your grade.
I didn't have to practice writing - I just did the papers assigned and rushed through them. You can grade papers for other subjects on prose and grammar instead of just doing it for language courses. You don't have to read classics to read better if you just read a variety of things you are interested in. You'll read plenty of boring things for other subjects and get that sort of practice.
With most of this stuff, having actual homework isn't necessary as long as students are given time for practice. Practice is what makes you good at something, not homework, and practice can take many forms.
You can absolutely learn to play some piano by just noodling. Your learning will be much slower, and you'll miss important fundamentals that are present in any piano method.
That's perfectly fine for learning things that you do for enjoyment, or things that don't really matter. It's a lot less fine for learning things that... do.
What I remember being questioned is does it make sense to do those exercises as homework or would they be better in school.
Or on the flip side should school get out early like 11 or noon, like I think the german gymnasium does and have all the exercises as homework.
The US system where children get out of school at 15:30 and still have a bunch of homework seems a little lopsided someplace.
The 3 month break in-between school years is definitely questionable.
I think the answer is a complex one because it intersects with personality and neurodivergence.
Depending on who you are and your family situation any form of homework can be a real challenge. Not because of what you are studying but because of how difficult it is to sit down and do anything you aren't passionate about. Certainly anyone with an executive function disability will have that challenge.
Industrialized mass education has always suffered from a unit economics problem: the labor required to assign individually-tailored problem sets and manually grade them in the volume needed for most students to actually learn the material is prohibitively expensive.
We need to shift the incentives by adding ruinous penalties for things that are currently quite commonplace if they are done by large players. Some dude training his own AI on his own computer can scrape and train. The fine for OpenAI or Meta using a single copyrighted book without permission should be in the tens or hundreds of millions.
What we're seeing currently in our society is a "loophole inversion" where the rules have an effect mainly via their loopholes. The most profitable activity is to find loopholes and exploit them as frenetically as possible to gain as much advantage as you can before the loophole is closed, or get people hooked on the loophole so it's retroactively legalized. Entities that are big enough to do this are big enough because they have lots of money behind them. Entities doing the same kinds of things without lots of money are not really doing much harm. So the best approach is to adopt a "sliding scale" in which even tiny violations by wealthy actors result in penalties enormously greater than fairly large violations by small players.
I think we don't yet know the long term view - what is harmful or damaging in the long run. Too early to tell and pass judgements. Every new technology from writing to internet had its detractors and they all pointed out negative externalities, as they seemed to appear at the moment.
We could probably cook up thousands of examples of the same problem (YouTube DIY tutorials, GPS navigation, etc.)
If it is an amplifier of "good" vs "bad", I don't know, and I do not know what that means exactly. I do not think some students are just "bad" or "lazy" by character and that explains all. I think that their environment has a big role in shaping their actions. I don't think that AI cannot be used beneficially (and the articles talks about some attempts that maybe worked well), but it requires careful considerations, guidance and intentionality from educational institutions themselves. The same goes with all kinds of technology, and AI is nothing new in that sense, but often new technologies are introduced in education without careful considerations or evidence based policies. Which is why imo articles like this are very important.
The data point around 80 minutes seems like noise to me. Looks like there isn't enough data/students who spend that much time and also used AI.
It would be nice if AI was a force for good as well as bad, but the data here doesn't support it
Same can be said of technology in general tbh.
Suppose it's good to learn how elastic the brain is, in both directions, at a young age where it doesn't matter.
Doesn't it matter the most at a young age?
It’s alarming how near universally AI has been a cataclysm in class rooms.
On homework the short answer is is that we give kids home work exercises.
AI helps learners reduce the effort expended to exercise and get results. This is making homework moot.
Evidently the AI users think hard things are beneath them. They want to win without doing any work. Sympathy, I have not.
A huge proportion of people who use AI use it to outsource their thinking.
It can be used for good for this purpose. You can use it to find primary source information. You can use it to generate challenges for you. You can use it to critique your work when there’s nobody else available. Etc etc.
Edit: Why don't they have lower-emission power sources?
It's good to think about second order impacts but this is not useful.
Just one 20 GW AI data center which is being built with gas powered turbine generator power plant, will be the largest fossil fuel power plant in the world... How exactly is that the same as one coal mine
Raising the noise floor like this only makes it that much harder to find "Smart" people, which we were already doing terrible at.
I use Claude every single day, but this is such a bad tradeoff. Maybe it will help me standup a quick fix when that is needed. Maybe it can help me dig through documentation to find relevant bits and figure out the unstated assumptions underlying it. Maybe it helps me generate test cases.
Meanwhile, my day to day life is now noise. All social media is noise. All content is noise. Slop pours onto me from all directions. Writing more test cases isn't helping me.
Am I smart? Am I dumb? I don't care, right now I'm deafened
I'm using Claude at work myself and am impressed with the product, but notice that this is the only reason I need to use it at all. Our product pages were shit to begin with, now they're AI-generated and somehow even worse. Our procedures are incomprehensible spaghetti with enough arbitrary context switching to give a sadistic Soviet municipal administrator an erection at the thought of watching anyone try to actually follow them.
Use AI to create inefficiencies, then use AI to bypass them. Those who can't do the latter will struggle to survive.
Clarification: to value “smart” people, which we were already doing terrible at.
It does give us a new heuristic, though: people who are willing to completely cut generative AI out of their lives (cold-turkey, if you ever started using it) are a much smaller group of, predominantly thoughtful, people. You do have to give up Claude to be part of this group, but from what you say, that's no great loss, and no longer being deafened is worth it.
This has considerable advantages over conventional elitism, because the barrier-to-entry is negative in almost all cases.
The one exception I've found is assistive tech, where the state-of-the-art is so poor that vibecoded slop is genuinely an improvement over the state-of-the-art, and in many cases the tooling simply isn't available to make your own assistive tech (unless you want to bootstrap an entire networked computing environment, which isn't very helpful when you want to do your online banking and do not, in fact, work at your bank).
But there are not many principled exceptions where you could seriously argue that the trade-off is worth it. Take mathematics, for example, which we often see touted as a "good use-case" of generative AI. The primary advantage of generative AI in mathematics is being able to search though a vast corpus of ivory towers and inconsistent terminology (without proper attribution) to locate and connect ideas that can help solve problems. The deficiency this is addressing is elitism, inadequate communication, and inadequate indexing within academic mathematics. This problem is entirely created by the academic mathematicians, and has been known for nearly a century (per https://en.wikipedia.org/w/index.php?title=Nicolas_Bourbaki&...):
> Bourbaki was founded in response to the effects of the First World War which caused the death of a generation of French mathematicians; as a result, young university instructors were forced to use dated texts. While teaching at the University of Strasbourg, Henri Cartan complained to his colleague André Weil of the inadequacy of available course material, which prompted Weil to propose a meeting with others in Paris to collectively write a modern analysis textbook.
To my knowledge, this is the only organised project to clean up and improve mathematical communication. Everything else (Metamath, Mizar, AFP, Lean) is yet another ivory tower. The Wikipedia article on this topic (https://en.wikipedia.org/wiki/Mathematical_knowledge_managem...) risks deletion as non-notable, that's how little anyone's actually trying. They made their own bed, and generative AI will only provide a brief respite from having to lie in it. (I was surprised how many other "compelling" use-cases evaporated when I applied this razor to them: the sibling comment https://news.ycombinator.com/item?id=49392265 points out one such.)
Vibe-coding assistive tech which doesn't yet exist, as a temporary scaffold to improve the quality-of-life of yourself and others in a social world dominated by non-essential access barriers is, to my knowledge, the only exception to this principle that can be justified. If you treat people who make other excuses, or who don't even bother with excuses, as not worth listening to, you lose little – and doubly-so, if you make your stance clear, so that others know the "cost" of gaining your attention.
... including ultimately stuff needed to make the smart people.
LLMs make significant mistakes frequently and smart people have no way of judging those mistakes outside their domain expertise. They are also sycophantic and great at being an echo chamber which makes people feel smart even if they are not.
So I think the burden of proof is on you to prove that they somehow amplify intelligence, it seems highly unlikely.
Smart people know LLMs confabulate and tell them they’re Absolutely Right! Smart people don’t want to be embarrassed by trusting the hallucination machine and revealing their gullibility to others.
All of those sound like flaws and defects of dumb people?
Great summary of the flaws with LLM "research"/"reasoning". It's always trying to con you, and I question the literacy and intelligence of the people who can't see this.
In software particularly, what AI does do is give me back my time from the drudge work that I don't care about. Keeping my build system configured and my tests up to date and my documentation synchronized is a good use of AI because I mostly don't care about how crummy the result is as long as it "works".
This is "AI"s forte. Producing more bad work at lower cost.