While I do appreciate the role of a proper citation, I'm reluctant to do it for blog posts (as opposed to my actual papers) both because it's time consuming and because for most people, links are a more natural and readable format.
12 karma · joined January 30, 2013
While I do appreciate the role of a proper citation, I'm reluctant to do it for blog posts (as opposed to my actual papers) both because it's time consuming and because for most people, links are a more natural and readable format.
pdfplumber seems mostly ok at extracting tokens. Sometimes it seems to combine tokens that should be separate. I suspect a few percent of the error is actually problems earlier in the data pipeline, as opposed to the model proper.
I also believe that preparing and cleaning this data set, and bringing a challenging investigative journalism problem to the attention of other researchers, would be valuable even if I hadn’t done any work on this baseline solution. This is a problem that journalists currently expend a huge amount of time and money on, which reduces the effectiveness of transparency around political ad spending information.