New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
intextbooks.science.uu.nl
intextbooks.science.uu.nl
First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement
Second, trying to incorporate past grades into their modelling is not a substitute for a randomized trial.
Third, the headline engagement number of 90% is for "engaging with the platform, via Module Review or Lesson Quizzes, at least once". I don't know why much of that couldn't just be attributed to novelty. Or even partly a professor with all sorts of enthusiasm for the platform.
Fourth, the "full dosage" effectiveness is measured based the final exam scores. Were these exam questions produced independently from the "Phosphor" materials? (e.g. by blinding?) Were they checked for direct overlap with those materials? The 0.7 sigma shift is 3 points on a 24 point exam; if even a few of the questions on that exam were very similar to those materials it could account for almost all of it. This is not clear to me from the manuscript.
If this was the case, then it's a question less of "is AI effective" vs. "did the students look at the materials". You could still argue that the AI platform got them to read, but that is a somewhat different statement than the AI helped them learn.
That means their experiment design is partially caused by their results instead of the other way around, which is a bad situation to be in. Their statistical analysis is completely inadequate for dealing with this.
And the change in engagement suggests that there's strong selection involved. Their attempt to use midterm scores to control for selection effects is unconvincing. Why not control for whether students used the platform more when there were only multiple-choice questions? Those are the ones who self-selected out of using the AI grader.
This is definitionally impossible. Testing whether the tool benefits the students that get it is an attempt to benefit certain students at the expense of others. Get over this hangup if you want to do this sort of research.
Just because I am not personally entitled (by funding, by ethics, by career position) to collect the necessary data to demonstrate a hypothesis, does not make that hypothesis demonstrable with lesser data.
Do students not like Mastership techniques?
Also, mastery learning doesn't look all that effective on standardized tests: https://www.jstor.org/stable/1170613
I think we've all had a few teachers who seem able to teach 2-10x more effectively than normal; they all seem to do it in different ways. I read recently the Gates Foundation felt they'd found nothing statistically significant in their billion+ dollars of teaching philanthropy.
The most plausible education improvement proposal I'm aware of is individual tutoring. This truly does seem to work well, and I think it's why there's so much interest in agentic tutoring. Perhaps we need a TeachBench to get some hill climbing done by the frontier labs.
1. Quiz completion is our deliberately conservative lower bound on reading compliance, and the 0.71 figure is not a claim that those 16 students each gained that much. The estimate is from a regression carried by the per-lesson slope, fit across the whole dosage distribution, and the underlying dosage-performance relationship is essentially unchanged whether or not zero-completion students are included (R² 0.091 vs 0.096). In other words, more Phosphor use is strongly associated with better performance across the whole range of usage - not just the group who completed all content.
The numbers in Table 1 show how dosage was distributed across the course. We report that across the class, the median percentage of lessons reached on Phosphor, including both students with an account and those who never logged in, was above 90%. Among platform users in particular, it was 96%.
We'd like to emphasize that for this pilot, the platform was presented to students as an entirely optional "study aid", and our adoption rates far exceed those reported in the past for optional interventions. It will be interesting to see how things go when we attach completion to the course grade, as we're thinking of doing in the fall. Past literature from interventions in college courses predicts that this will achieve far higher levels of engagement, bringing the high-dosage effect to a large proportion of the class.
2. We explicitly note this in Limitations; it's an observational study. We were unable to do an RCT for this course since it raised an ethics consideration - neither we nor the instructors wanted to deprive students entirely of a course material that could have been helpful for them. We'd love to run a randomized trial at some point though - one way to do this is a crossover, where we offer the treatment to one of two groups, then switch it over to the other midway through the trial, so that both get even treatment. Another possibility is randomly selecting students to get access to MCQ-only vs. CRQ-enabled quizzes. That being said, this mechanism of conditioning on past performance is well-known and relatively robust for observational studies of educational interventions.
3. The platform was created independently of the instructors of the course. The instructors designed their curriculum ahead of time (as had been taught for years of past offerings of the course), lectured in a conventional style, referenced the course's official textbook (Freedman, Pisani, Purves) and suggested homework problems from the textbook only. The instructional content was authored using material that every student in the course had access to, and did not feature exam questions that students were evaluated on after-the-fact.
Phosphor was not endorsed publicly by the instructors, and was spread primarily via student word-of-mouth. In fact, one instructor in this course initially believed the project would be "a waste of time" and refused to collaborate with us for the pilot. Despite this, 97% of the students in this instructor's section used the platform!
Engagement also persisted well across the full ten-week term, and two-thirds of Review attempts involved retries spaced a day or more apart — not a pattern typically produced by novelty effects.
4. Instructors wrote exams independently with their long-running FPP-based curriculum. Even if we steelman and suppose that "the platform just got students to engage" rather than truly learn, this is refuted by our result that the MCQ-only Module 2 had similar engagement but no dosage relationship. This strongly suggests that the CRQ format was a driver of the results.
As we mention in the paper, we agree that replication, especially across contexts, is a priority. For us in particular, this means not only across other courses, but across other institutions as well. And an RCT would certainly help lock in the causal claim.
Anything promising that helps students learn and learn more effectively is obviously incredibly valuable. My concern here is that I don't read anything trying to disambiguate students who just, you know, work harder, and therefore used whatever materials they had to hand including Phosphor, from those who work at the same level and chose not to use Phosphor.
We all recall from our days in undergrad that there are students who do nothing and slide by, a few who do nothing and ace everything, and students who outwork everyone else and outperform -- I think a tool that only helps grinders who were already going to grind is likely not the contribution you're hoping to make.
(ie changing the environment can lead to short term productivity gains because either participants are aware they are being watch, or it breaks up the monotony and makes people work a bit harder. )
I'm convinced this is the future of education - models are there, we need the classroom tech to catch up. The alternative is obvious and quantified in the paper - students just use models to do their work for them and learn nothing.
Practically, I think if you want the AI system to have a live view of what the student's doing you're going to have to replace one of either the tablet or the writing instrument. A wearable camera could work as well but there are issues with that.
and after looking it up, it appears they are still available: https://www.livescribe.com/landingpage/ls3_onenote/
Spaced repetition is very effective, but it's really really clunky to use. My unpopular opinion is that we all have Stockholm syndrome when it comes to creating "cards", and people talk about how valuable creating cards is; but I think it stucks, it takes a lot of time.
If AI is already teaching me math (let's say), it would be nice to tell the AI/app "quiz me on this periodically", and then the AI makes up a fresh polynomial to factor (or whatever) and presents that to you according to a spaced repetition algorithm.
Behind the scenes, the AI should have access to what has happened the last several times a specific topic has been quized, so the AI can watch to see that certain mistakes are resolved, and the AI might also know better how to correct the user if it has context about previous quizzes of that topic.
I'm willing to grant that there is some value in choosing what to put in the cards, but most of the awkwardness around making cards is UI related. Nobody creates cards on their phone, or while they're walking (AI could do both of these) - people create cards sitting at their computer (like cavemen!) usually clicking through a clunky UI and managing thousands of cards with thousands of clicks. That sucks, and people probably wont realize it sucks until something better comes along.
Wait, when are you doing it then? No wonder you think it sucks! Adopt some modern tools, yo. Use Anki, or vibe code your own app.
Anki is great for studying, but the card creation experience sucks. To be specific: I found creating any custom card type immediately dropped me into the bowels of CSS. It felt like writing HTML by hand.
Is there any facility for re-using shared pieces?
I felt like it needed a static site generator type tool to move up a layer of abstraction and reduce the copying of chunks into my card. Is there one? Please mention if so.
Using that as the input file format standard, AI can generate what you're looking for, Android app, Webapp, iOS app, pdf.
Ah, LaTeX syndrome.
I use an Emacs based SRS tool. I have a capture template to quickly make a card.
For anything tedious, it's critical to reduce the friction!
Just to pick your brain real fast, when creating new cards (from scratch?) what would a better experience look like for you?
Spaced repetition is just reviewing the same material periodically. It doesn't have to be a complicated system.
We already know how to learn and educate: spaced repetition (periodic review), and retrieval practice (frequent testing). This is how school used to be fore centuries; it's not sexy but it's effective.
Maybe reMarkable or something like it could help bridge a student's writing with an LLM without having to fall back to a laptop or ipad.
What does “bridge a student’s writing” mean?? If this is a real argument it needs to be clearer.
What’s the functional difference between a Remarkable and an iPad? The former is less responsive, costs less, and has better battery life, right? I really don’t see how that’s significant to any kind of development of anything.
Are you talking about running a local model??
I assumed that too, back when i thought that getting an e-ink pad would be cool.
Right now Amazon is quoting me $450 on a 10.3" reMarkable, but the price goes up for the 11.8" model and/or for bells and whistles. It looks like an 11" ipad is $450 and up (based on https://www.apple.com/ipad/compare/), although you need to buy the $80-100 iPencil separately.
So yeah - I was hoping that e-ink things would cost less, but they're still pretty expensive and certainly competitive with the iPad.
remarkable keeps inking data on a per-page basis in vector format, so there's a bit of work to be done to render the inking and then visually assessing it, but it's all pretty much "done" work.
Remarkable's cloud syncing tech is .. not great, and I'd speculate you'd get a better setup here with some sort of harness where you talk / text with an agent that has access to all the files, dates, OCRs, etc first before you consider trying to generate PDFs / pages and put them back onto the tablet.
https://github.com/awwaiid/ghostwriter
Installation not for the average user.
> constructed-response questions (CRQ) are graded by Claude Sonnet 4.6 against instructor-defined, question-specific rubric criteria
> Crucially, LLMs make it feasible to grade formative CRQ against rubric criteria at scale, a capability that appears pedagogically significant rather than merely convenient.
They specifically call out that the "RAG chat assistant" part of Phosphor (the platform) wasn't used much.
I commend the effort here, but I don't think these results are particularly noteworthy. The conclusion is essentially that people who do practice quizzes will do better on exams.
What do you think tutoring is?
> and lacks randomized controls. Self-selection is the central threat: students who complete more quizzes may be more motivated or higher-performing generally
But this is still a strong result. I'm excited to see more in this space.
In the 1980s, a researcher called Benjamin Bloom claimed a z=2.0 (that is 2σ) advantage for a combination of mastery learning (don't move on the the next topic until you've mastered the current one) and 1-on-1 tutoring. Later replications show there is definitely something going on, but the effect size is much lower, for example around z=0.7 in a 2020 paper [3].
I'm still open on AI tutoring, though the Dartmouth results look impressive. Someone please try and replicate this.
There's a saying that AI helps the best students get better, and the worst ones get worse. (Anthropic sort-of agrees [4].) It'll be interesting to see how that turns out.
[1] https://www.theintrinsicperspective.com/p/why-we-stopped-mak... [2] https://www.astralcodexten.com/p/contra-hoel-on-aristocratic... [3] https://www.nber.org/papers/w27476 [4] https://www.anthropic.com/research/AI-assistance-coding-skil...
Bloom's Two Sigma Opportunity suggests that there's another SD improvement available: https://en.wikipedia.org/wiki/Bloom%27s_2_sigma_problem
The average level of capability and comprehension in fundamental disciplines for students completing primary is a direct result of some fundamental differences in the way they approach classroom organization: for one middle school and high school students do not change classrooms during a school day.
Teachers rotate rooms while students remain in place and this maintains a less disjointed, less distracted transition between classes. There are still very little to zero technological advances to the teaching methods: blackboards/whiteboards, overhead projectors, hand written paperwork, textbooks.
The reasoning is simple. The fundamentals which students in primary education are in need of learning do not change much. Last years' textbooks are perfectly useful and the cost of replacement is directly carried by each student. If they damage a book, they must replace it.
Salaries for teachers are not unusually high, but they are also not low. The building and administrative costs are kept low, for one, by there not being a significant non-instructor labor expense of janitorial and maintenance workers. Schools share the pool of municipal HVAC and other trades for serious infrastructure, but janitors? Nope. The students are required to actually clean up after themselves, and to actually clean the whole school. It makes a difference, and those avoidable labor costs can be directed to proper compensation for the teaching staff.
Anecdotal observation of the overall efficacy of this approach reaches areas not usually measured as 'educational success' but also includes one noteworthy artifact. The common clerk, shop keeper, cook, or gas tank delivery driver all know that i = sqrt(-1) and that a complex number is a pair of numbers, one of which is the coefficient of (i). Let that sink in. How many people graduate with a B.A. from western schools without that?
There is not a 'silver bullet' but there are a full compliment of approaches which when applied together, consistently, and persistently, yield excellent sustainable outcomes.
Back to the falling population problem. There are pluses and minuses from the choice to provide separate boys and girls middle and high schools; one of which is lower teenage pregnancy and it's inverse, lower adult birthrates. (It's not the only reason for the lower birthrates, but it's not having zero impact.) You cant win 'em all.
On the other, I'm sceptical of that it'll have "strong benefits" at scale; I'd be more in favor if the wording was "some"/"moderate". I reckon self-selection plays a huge part, as mentioned in the "Limitations" section of the paper.
I'd also caution against attaching the tool to grading. That means students have to put more effort into the course, which increases the chances that they will use LLMs to save time rather than make the investment.
Mind if I ask what did you learn and how you're using it?
The reason I'm asking is that I repeatedly felt excitement only to realize down the line that the explanations didn't actually translate into practical skills. I'm not sure it's even an AI problem, it's a "doing versus reading" problem. Same as with reading a pop-science article and thinking to myself that I learned something about physics or medicine or mathematics.
If it's purely a correlation, then maybe those students would be more successful than average even without the study group. They're already the most motivated kids. Maybe they just do "motivated kid stuff" and would still outperform.
This sentence is accurate, but inevitably leads to the confusion you see in these comments.
Earlier I used Claude by giving it the course material and asking it to generate me exercises (our cpurse work went way over my head) and yeah i learned to differentiate a gradient or Jacobian, but it was very shallow - I knew the formulas, but not what they meant or how to apple them correctly. After I just filled glaring holes I had in Univariate Calculus by readong and doing, I actually started to understand something.
Lon story short, in my experience Learning with LLM’s is ok with very unfamiliar material that is not too complex (there’s obvious problems of LLM’s themselves being pretty ghastly with maths sometimes), but at least it os not better than the traditional method of just putting your nose on the grimd stone.
I'm curious how well you feel this worked because the subject was Statistics (objective grading) versus something more subjective like Civics or Literature.
PS - I'd say this qualifies for Show HN, too!
Do you
It still could be better for students, but it's not obvious that it would be (or maybe not as strongly?).
Just want to say that:
>In our deployment, student-reported reading completion baselines for MATH 010 were approximately 15%, with instructors estimating 10%. Individual student reports of reading compliance ranged from "literally no one does that" to "is this being recorded?"
is hilarious
Are you planning on opening access to Phosphor?
What creeps me out about bringing LLM into early education is that it's a period where kids learn to socialize and cope with problems, and I do worry about forming substitute relationships with chatbots that are engineered for sycophancy / enablement. But I guess that's a problem either way, because almost every student will try an LLM at some point.
I think there is more potential applications possible with combining LLMs with reference/text books. Like how about an assistant that points you to the correct books/chapter/paragraph for the concept you need to understand better for a project you are working on? Or clarify any confusion you are having?
Like a human tutor but infinitely patient and non-judgy + search engine.
I'm more curious how students perform on the test with vs. without AI.
The environment is just obviously two sigma better. This just... seems obvious to me? In the same way that I will get stronger much faster if I have a physical trainer to tell me exactly what I am doing wrong when I do it? And it seems obviously unsolvable other than by getting everyone a private tutor (or AI..?).
Asking from a place of curiosity.
Bloom looked at other methods used in concert to achieve similar improvements to “solve” the problem that could be delivered at scale.
LLMs may provide a new path to the 2-Sigma improvements without the same delivery problems.
And in general ongoing from “it’s obvious to me” to a quantized effect is always an effort. For instance I don’t believe you are correct in assuming that one would see a two sigma effect in physical exercise when comparing someone who attends a regular exercise course va someone who spends the same time with a personal trainer. Two sigma is a lot, and you won’t be lifting significantly heavier weights from hav a PT vs doing starting strength training. In my opinion you would most definitely see a two sigma effect from doing steroids though. But this is all pure speculation which underlines that part of the “noice” is about the documented aspect of the two sigma problem, where we have a body of data to work against not just personal assumptions. That said if anyone has studies on personal trainers va course work in fitness I’d love to see my world image challenged by data.
Then it is "effectiveness", not "efficacy". Prefer simpler and more specific words when possible, to reduce effort for the reader.
Jk, but the skepticism is inevitable. I think we can be dubious about how AI mobilizes global capital while also appreciating tutoring as one of its best targeted use cases.
Hasn't computer assisted interactive learning already been proven for years? Why does there seem to be so much skepticism about enhancing it with AI?
Is this just something like, astoundingly slow adoption or poor execution? Being held back by paper textbook makers? Teachers unions dragging their feet?
How can interactive AI driven individually paced learning _not_ be obviously dramatically more effective?
very few are actually motivated to learn and are just there to get a job or its just next thing that they have to do in life.
Motivation is also a huge part of the problem. I'm wondering if the novelty of the AI tutoring gets more people to try it and whether it would wear off?
It's surprising to me that many students at Dartmouth don't read the textbook. You'd think college admissions would select for that?
It seems promising but, as they say, more research needed.
There ARE technologies that have improved things, but so much high-cost useless tech has been shoved into every level of education that many educators are incredibly leery of new tech.
The issue is that while the underlying technology is useful, the way it gets integrated is frequently not. An administrator cuts a deal for a product they never have to use to an ed-tech giant for a huge amount. Because the ink is dry and a huge sum of money has been spent admins pressure educators to use the technology as much as possible regardless of outcome.
In that context it makes a lot more sense why there is pushback and FUD among educators.
Having significant roles in the design, funding, and implementation of various public and private school educational projects which integrated some minimal technology I have seen this firsthand as far back as it has existed.
The schools which had the GOOGLE Classroom crap 'sold' to administrators and forced upon the teachers and students were examples of absolutely abysmal failure at every role to which they were supposed to help.
GC is s complete waste of the schools limited budgets, a waste of instructors already over-committed time, and a complete failure to deliver improved educational outcomes for the students. They may have been more 'engaged' in the screens, but they learned less. In the end the GC framework becomes a leash, used by schools to monitor and invade the out-of-school aspects of student lives, and tilts the field more steeply where disparities of home income levels exist.
My particular programs' inclusions of technology were never for 'education' with the exception of elective CS and robotics classes. Otherwise the technology was for records management and course content preparation help for teachers.
One project's tech use case stands out distinct from that limit. A charter school project - instructors and students working to create video and online course materials documenting the conversion of a school program from single grade standard 'science' classes to multi-age, collaborative, hands-on STEM science programs.
But the tech use and coursework were more for documentation and dissemination of the conversion process than for the principle educational objectives.
Actual students did use adapted legacy Samsung android devices to record observations and collaborate among peers during the observations of field studies and workshop classes.
The handhelds checked out to students' on a 'per project' basis each day, so there was no 'individual content' on any of the devices.
The repurposing of these devices wasn't accomplished through 'locking down' an overcapable platform, but by building from the kernel up and only including the specific functionality as required to collect and collaborate - so, no web, no chat, no games, no distractions. The specific workshop content and objectives were available via the handheld, and at specific places in the progression the students were required to input observations and record still or videos. Those images were not shared among students and there was no incentive for shenanigans as all content was directly attached each student's workshop performance and we were clear in the start of the classes that because this was going to be part of the conversion process documentary, no people in any film. we has a big library of already categorized required captures: Plants in various stages, equipment set-ups, insects in situ for ID, test-strip color comparisons, reagent container labels, etc. Some of the kids called them something from a scifi TV show, complete with "Beam me outta here" jokes.
So, technology can be very useful, but not when it comes with external motives and further widens socio-economic disparity. If the objective is MORE engaging content and MORE accurate records of student progress, then sure.
But, that usefulness is usually directly inverse to the cost and functionality of the devices: cheaper, limited capability tech will usually if not always outperform fancier, overcapable tech.
A $very-cheap refurbished Samsung Galaxy stripped down to whitelisted bluetooth and Wifi LAN, with an distributed storage layer and local-only browsing (Apache Cordova based app) can get alot done without adding tools not required for the specific needs of the student [or educator in many cases].
Campus server infrastructure de-minimus was a solved problem with bittorrent sync and each device only subscribing to the specific file-tree required for the specific workshop/class, and student submissions are only accessible to that student, and the course instructor(s) - eliminates the bullying and other misuses of technology.
We also were able to use a direct approach to resolve the personal phones in class problem, at least for the part f the campus participating in the STEM conversion (4th to 7th grades). Students had to turn off phones, lock them in a pouch, and hand that in to receive a handheld.
If a parent needed to reach the student in school, that's why the school front office had phones. Otherwise, student's were 'in class' and not to be disturbed.
But this was over ten years ago, and today's parents and students would probably throw a toddler sized tantrum at these rules.
Too bad the project got rug-pulled for funding in between years 1 and 2 for the documentation and distributable methods layer; but the school did get to keep teaching the hands-on STEM focus to completion of the then enrolled students - we baked in the costs for that into the first year at the insistence of the participating teachers and parents.
TLDR; don't let your school get SCROOGLED.
Text book reading in this course was 10-15% at baseline ... but this AI thing got 90% voluntary usage ungraded.
Even if its worse per-hour than a textbook, you're now teaching 6x as many students _something_ instead of teaching a small minority everything.
So really it just becomes an optimization problem at that point because most students are at least in the funnel/in the running to learn something.
The paper kind of proves this itself ... they tweaked the quize formats mid-semester and where able to iterate which you can't do on a textbook that nobody opens in the first place