Visualizing 40,000 student code submissions
stanford.edu
stanford.edu
A couple of years ago, I made some abstract art of the inheritance structure of a large project written in a language with (extensively used) multiple inheritance, there being around 1800 classes. No one was game to produce a 15m-wide, 30cm-high wallpaper (the traditional type) of it, so I just removed the class names, leaving classes just dots and made it my computer's wallpaper. It's got quite a few comments. Still, it was nowhere near as pretty as this.
The language used is an in-house language, developed in the late 1980s and early 1990s, and one that has aged surprisingly well (with comparatively few modifications to the language since then), though there are now better options available.
clustering is putting similar things near similar things. Tree edit distance is quite a natural measure of distance for tree like things like programs.
You can't avoid some warping when putting high dimensional manifolds on a low dimensional one. You can see a lot of their data does cluster properly but their are some long range red arcs (in the embedding space) which are side effects of warping (they are near in data space).
You can see a cluster of green which is clearly of interest ... why did so many students get the wrong answer in the same way?
I see lots of value in that picture.
Well we have a lot of ideas! One thing that we did, for example, was to apply clustering to discover the ``typical'' approaches to this problem. This allowed us to discover common failure modes in the class, but also gave us a way to find multiple correct approaches to the same problem. Stay tuned for more results from the codewebs team!
This is a great way of using metadata to search for patterns in student assingments. This could detect different "approaches" or "strategies"