I’d be interested to see how Google’s LaMDA compares on this task (currently available in private preview).
One of LaMDA’s unique (afaik) features is a fact-check system that edits the main LLM outputs, to reduce bullshitting. This seems particularly important in an educational context where impressionable young minds are talking directly to the LLM.