Markov Chain Monte Carlo (MCMC) Sampling, Part 1: The Basics (2019)
tweag.io
tweag.io
I'm happy that the "pure this pure that" fundamentalism has faded into retirement... but shut up and calculate has problems - especially around causality. The "blessings of multiple causes" paper for example.
Classical genetics and statistics were very tightly linked in this regard due to scarcity of data at the time (pre-2000).
Now that the field has exploded with big data, such statistics are no longer so neccesary to answer the same problems, and a lot of the early mathematical elegance associated with the field has given way to brute-force filtering methods.
https://www.youtube.com/playlist?list=PLwJRxp3blEvZ8AKMXOy0f...
He works his way up to MCMC
https://uk.sagepub.com/en-gb/eur/book/student%E2%80%99s-guid...
The talks of the asynchronous part are publicly accessible on the forum: https://discourse.pymc.io/c/pymcon/2020talks/15
Tickets for the synchronous part with Keynote sessions and live Q&A are still available: https://pymc-devs.github.io/pymcon/
If you want to condition on an event? Say you want to predict the weather on Tuesday, conditioned on the event that it rained on Sunday? Run a lot of simulations, and only keep the ones where it rained on Sunday.
Similarly: If you want to compute an expectation value? Run a lot of simulations and take the average.
If the event is unlikely? Say you want to condition on the fact that it rained on Sunday, and the high temperature was precisely 15 degrees C? Then you have a difficult problem on your hands.
(Sometimes the expectation value of the quantity you care about will depend a lot on a few rare events with outcomes many standard deviations away from the mean. Then you also have a difficult problem on your hands.)
Sometimes MCMC will work on this kind of problem, and sometimes it won't. Even if it doesn't, maybe other techniques will work.
To apply MCMC to a dynamical system, one method is write down all of the history of the system as a single object, say a single vector. You write down what your system is doing at t=1, at t=2, etc, and all that information goes into the vector. The rules governing the system determine a probability distribution over the vector space that the vector lives in. (Or more generally, the object in the object space. The vector axioms aren't important here, it's just a nice familiar example.)
Generally speaking, if you know how to describe your dynamical system, you know how to compute an non-normalized probability for any given vector. Usually you won't be able to compute a normalized probability for that vector. That's fine, since MCMC works with non-normalized probabilities.
The next step is simply to run MCMC. If you want to condition on some fact, cut out all the parts of the space where that fact doesn't hold, and then run MCMC. If you want to compute an expectation value, there are other tweaks to MCMC that are possible (i.e. importance sampling).
Chandler's statistical mechanics book might be a good place to start to dig deeper
I'm having a hard time finding literature on problems like this. I think the closest thing it's like is the stable matching problem, but I'm not sure. My current algorithm doesn't allow for missing or reordered pages.
Does anyone have magic words I could google to find pertinent research/algorithms? I think if worst comes to worse I'll have to try to model the incoming page stream as something like a bayseian ... graph? It's been so long since I've done something like that (5 years), and that was in school.
The most basic thing I'd imagine (assuming text docs) is something like index each document by TF-IDF, and then the new document by TF-IDF, and then rank results by similarity
You're basically doing a google search on documents, but with a really big query (another document)
The challenge would be to get a good measure of the difference, perhaps us some cosine difference again from the transformer?
If you have specialist vocab you can train the transformers - so maybe you are doing legal matching?