The system would prompt students with problems that incorporated dozens of sub-skills, each of which incorporated other sub-skills and so on. If the student gets the problem right, all requisite sub-skills are marked as reasonably strong. If they get it wrong then the statistical model concludes that at least one sub-skill must be weak/faulty/missing, and when selecting subsequent problems the system would incorporate problems with different subsets of the original sub-skills to identify the missing sub-skill(s), with the goal being to spend time reinforcing only those skills that the student has not mastered.
I no longer work in that domain, but I recall it being amazingly efficient -- it often only took a dozen questions to assess a student's understanding of thousands of skills.