To answer your question, I'm a product guy, so it was easier for me to start with the user app/mobile UX I wanted + constraints rather than starting with the cards.
Then I created a problem view in the dev version of the app that let me see every type of problem to see how it rendered, how input worked, etc.
Then with that plus the system around it (details below) gave me higher confidence in generating the 40,000+ problems across all the topics. It still has room for improvement, but happy with how it came out so far.
So, for the problems, they were programmatically generated, but NOT LLM-generated/trusted, so the process would give me the confidence:
1. Define the problem families explicitly. Each has bounded inupts, known mathematical rule, answer type, constraints & presentation rules.
2. A deterministic compiler generated the problems. So, given the same source definitions/versions itd produces the same corpus. so there's no AI inventing random questions at runtime .
3. Correct answers are computed from the underlying math, not from the rendered text. Integer/rational problems use exact arithmetic. More complex symbolic cases use validation recipes, incl. offline SymPy if needed.
4. LaTeX is just presentation (tried other way, didn't work). Internally the problem is represented as structured math, then laTeX/MathJax derives from that structure, so it doesnt generate a LaTeX string and cross fingers its interpreted correctly.
5. Generation/verification are separate steps. Every generated record has to pass mathematical, domain, schema, answer format & presentation checks before it can enter the corpus. Symbolic cases can be independently verified, for example by differentiating a proposed antiderivative.
6. Whole families are tested, not just samples. Finite spaces can be exhaustively enumerated. Larger spaces get property, boundary, invariant, and regression tests. It also runs corpus-wide audits for malformed questions, duplicate IDs, invalid answers, broken rendering, unreachable answer forms, etc
7. The shipped artifact is tied back to what was verified. Versions/content digests bind source definitions, generated problem, validation result & runtime representation together. If something changes, it has to be revalidated rather than reusing old results.
So, starting out I assumed doing it programmatically would be liability. But rather confidence comes BECAUSE it's programmatic. For a large class of problems deterministic generators + exact/exhaustive validation was easier to audit than tens of thousands of hand written questions. So could leverage AI for all that, just had to define the rules/review the system.
I'm very happy with the outcome, but in terms of process it far exceeded my expectations in what I learned about approaching problems like this.
I'm a very heavy AI coding using (I burn hundreds of billions of tokens a year!) so this was a fun way to validate my approach and test some new ones to build high quality experiences, particularly on mobile which is more of a taste thing. The decade plus of product experience really came in handy in guiding the AI here, which took a lot of iteration on the experience and trying different things. It was a blast, and I was generally surprised at how fast it came together. Thank you again.