- It turns out good math questions have a ton of forms! Multiple choice, fill in the blank equations, word problems, diagram labeling, etc, and a good math book or online tool varies the images and forms used so students don't get bored. - AI just isn't great at generating these forms directly, so I ended up having to create these heavy question form adapters to try to get the output of AI to work.
At the time I did it, about six months ago, the LLM I could afford to use for the content just wasn't good enough, and I was having a lot of issues with accuracy, adherence to the curriculum item, and a lack of variability in the numbers and words chosen. Though writing this comment actually gives me a few new ideas...
I shelved it but I'm tempted to try again seeing this and knowing how much models have improved.
I was initially wary of doing it programmatically since the slop risk was high. But in the end, thanks to the process, making doing it programmatically made me MORE confident. E.g. for many probs deterministic generators + exact/exhaustive validation was easier to audit than thousands of manually written ones (there are over 40k problems).
So I could leverage AI mostly after defining the rules and reviewing the system (which did have a lot of manual QA upfront especially for the different types, but went quickly after I built the QA tool). And that let me treat it more like a problem I'm better at solving as a product/systems guy, rather than a math expert.
TL;DR the problems were programmatically generated, but NOT LLM-generated/trusted. Rough process in case it helps you:
1. Define the problem families explicitly. Each has bounded inputs, known mathematical rule, answer type, constraints & presentation rules.
2. A deterministic compiler generated the problems. So, given the same source definitions/versions it'd produce the same corpus. So there's no AI inventing random questions at runtime.
3. Correct answers are computed from the underlying math, not from the rendered text. Integer/rational problems use exact arithmetic. More complex symbolic cases use validation recipes, incl. offline SymPy if needed.
4. LaTeX is just presentation (tried other way, didn't work). Internally the problem is represented as structured math, then laTeX/MathJax derives from that structure, so it does NOT generate a LaTeX string and then just hope it's interpreted correctly.
5. Generation/verification are separate steps. Every generated record has to pass mathematical, domain, schema, answer format & presentation checks before it can enter the corpus. Symbolic cases can be independently verified, for example by differentiating a proposed antiderivative.
6. Whole families are tested, not just samples. Finite spaces can be exhaustively enumerated. Larger spaces get property, boundary, invariant, and regression tests. It also runs corpus-wide audits for malformed questions, duplicate IDs, invalid answers, broken rendering, unreachable answer forms, etc.
7. The shipped artifact is tied back to what was verified. Versions/content digests bind source definitions, generated problem, validation result & runtime representation together. If something changes, it has to be revalidated rather than reusing old results.
One thing I'm especially loving about AI is that it lets you convert increasingly more problems to something you're good at, which makes it easier to approach/solve in novel ways. It let me do this with Mathy, which would have been untenable before. And I've been doing the same with robotics, which is exciting as a software/product guy.
It sounds like you spent a lot more time on it than I did mine haha. Nice work.
I agree with your conclusion about AI, it has really opened up a lot of things for me that weren't possible before because of time and focus limits. Like a telegram bot that manages my strength training workouts for me and can adapt things on the fly based off written feedback right in the bot. That's something I never would have invested a lot of time into before, but I had this done in a half hour!