Show HN: A tool to scrape senators' stock transactions for your own analysis
github.com
github.com
I provided a Jupyter notebook with my own simple analysis (limitations included in the README), but of course it can be modified to suit your needs. Let me know if you have any questions - happy to help!
Bravo!
Doesn’t this get looked at by the SEC? Not always, but they monitor these kinds of transactions especially when millions are involved.
I’d be curious to see if a passive index to track members of the Congress would beat the S&P.
Although it may not account for other sweet deals like this option expiry extension and accelerated vesting https://theintercept.com/2020/05/12/david-purdue-senate-card...
Also consider that extremely wealthy senators (many are) might only have a small part of their assets in stocks and could be part of a bigger financial strategy, as a hedge against other asset classes rather than an indicator of where they or their agent thinks the market will move.
I looked at performance as well, but its tough to benchmark these things. There were net outflows in the market during that period and if you sold you did relatively well in a falling market, so the typical trade did well. But the relatively small size of the trades makes me skeptical that insider information was used. Matt Levine said it best:
> Of course they could be lying, but in context the defense seems pretty plausible. (Kelly Loeffler, for instance, controversially dumped about 0.6% of her portfolio at around the same time, which sure seems like the sort of thing an investment adviser would do without any input from her? You could call your adviser and say “a disaster is coming, sell everything!,” but calling them to say “a disaster is coming, sell a tiny bit!” seems pointless.)
But this doesn't include William Barr as he is not a Senator.
[0] https://senatestockwatcher.com/
[1] https://medium.com/ml-everything/analyzing-us-senators-stock...
[2] https://www.bloomberg.com/opinion/articles/2020-05-14/senato...
But that's too creepy for me (which is probably what is being counted on).
https://www.amazon.com/Throw-Them-All-Out-Politicians-ebook/...
They have a tool that takes returns and back tests then and finds correlations they run on hedge funds and can probably find patterns quickly.
Source: I worked there ~4 years ago.
https://www.nytimes.com/2011/07/10/business/mutfund/congress...
And this in support that avg congressional portfolio is not performing well:
https://www.mit.edu/~jhainm/Paper/Eggmueller_CapitolLosses.p...
https://www.morningstar.com/news/marketwatch/2020041854/sena...
Looks like there might be a historical component around the STOCK act that different studies play with for their data window. In either case — I still say for the poor level of diversification within & around asset classes that I don’t see a lot of shining examples of how to achieve success.
Wouldn't you expect a poor level of diversification in a stock portfolio based on corruption? They aren't dollar cost averaging and diversifying. They're doing the exact opposite to maximize the trade.
[0] https://insidertrading.procon.org/view.answers.php?questionI...
I realize this doesnt help for illiquid assets, but it does help for assets where you are locked into vesting schedules.
And I think it's not just Republicans who would have to do some hard thinking here. Bloomberg is a Democrat, for example. Does he want to be president, or does he want to run Bloomberg? Both of those seem like plenty for one person to do. I think it's reasonable to say that if he wants to be our leading public servant for a number of years, he should sell off his business to people who have time to focus on it.
so you're proposing one unelected entity which basically has full control of all personal assets of the legislature?
seems like a pretty big potential for vulnerability
When Social security taxes exceed payout (which they haven't since 2009), the income is used to finance deficit spending and the SSA is given a special federal bond.
[0] https://www.nber.org/papers/w26975?utm_campaign=ntwh&utm_med...
https://github.com/nathell/skyscraper
Writeup of a sample data acquisition + analysis usecase:
http://blog.danieljanus.pl/2020/05/08/making-of-clojure-depe...
There is more processing to be done. There is a Jupyter cell that he says takes 3 hours to run.
I will make a processed data file available, hopefully this weekend. If anybody wants it, let me know.
Published at https://qri.cloud/feep/senate_financial_disclosures
I may add more fields, so check for updates.
There are other sites that offer access to what they call the "Yahoo Finance API". I wonder how that works and if those are legit offers.
Does anybody know?
I wonder if it is legal to scrape stock market data and then republish it.
Secondly, it isn't legal. They are protected by license agreement.
But stock prices is not what the value of SNL Financial and/or Capital IQ is for most people. I know a bunch of people who pay the $30K license fee for these (as well as Bloomberg etc) who barely use the stock price at all.
Transactions should not.
I think it’s entirely reasonable for politicians to assume a life of less privacy since every aspect of their lives is scrutinized by the press anyway.
Term limits fixes this from both ends. Politicians are allowed their privacy, because they're just regular people and nobody is going to get rid from a single Senate term or 3-4 House terms. It eliminates the incentive to associate one's self more with DC than your home district. And it would probably result in people having a better opinion of Congress as a whole because it's not filled with people who have never done anything else.
Any particular reason you picked using selenium+chromedriver instead of an http lib (like requests) and beautifulsoup?
Obviously you'd have to handle with correct headers and CSRF tokens, but it'll be easier than Selenium for sure.
draw: 1
columns[0][data]: 0
columns[0][name]:
columns[0][searchable]: true
columns[0][orderable]: true
columns[0][search][value]:
columns[0][search][regex]: false
columns[1][data]: 1
columns[1][name]:
columns[1][searchable]: true
columns[1][orderable]: true
columns[1][search][value]:
columns[1][search][regex]: false
columns[2][data]: 2
columns[2][name]:
columns[2][searchable]: true
columns[2][orderable]: true
columns[2][search][value]:
columns[2][search][regex]: false
columns[3][data]: 3
columns[3][name]:
columns[3][searchable]: true
columns[3][orderable]: true
columns[3][search][value]:
columns[3][search][regex]: false
columns[4][data]: 4
columns[4][name]:
columns[4][searchable]: true
columns[4][orderable]: true
columns[4][search][value]:
columns[4][search][regex]: false
order[0][column]: 1
order[0][dir]: asc
order[1][column]: 0
order[1][dir]: asc
start: 0
length: 25
search[value]:
search[regex]: false
report_types: []
filer_types: []
submitted_start_date: 01/01/2012 00:00:00
submitted_end_date:
candidate_state:
senator_state:
office_id:
first_name: a
last_name:So I think we'll still need to parse from the HTML. But it might still be cleaner than using Selenium, so I'm open to any code changes.
It would make the script faster and less error prone for sure. I'll check out the repo in depth and see if there are any low hanging fruit.